Models.
Omnesly connects to Anthropic, OpenAI, Google, Qwen, Mistral, DeepSeek, xAI, Cohere, Meta Muse, Moonshot AI, Z.ai, Tencent, and any OpenAI-compatible or self-hosted endpoint.

Claude models, tuned for careful reasoning, long documents, and coding.
| Model | Best for |
|---|---|
| Fable 5 | Long-running agents, next-gen intelligence |
| Opus 5 | Complex agentic coding, enterprise work |
| Sonnet 5 | Best balance of speed and intelligence |
| Haiku 4.5 | Fastest, near-frontier intelligence |
GPT-6 and GPT-5.6 models spanning frontier reasoning down to fast, low-cost workloads.
| Model | Best for |
|---|---|
| 6 Astra | Frontier agentic reasoning, computer use |
| 5.6 Sol | Complex professional reasoning, coding |
| 5.6 Terra | Balanced intelligence and cost |
| 5.6 Luna | Cost-sensitive, high-volume workloads |
Gemini models with native multimodal input and very long context windows.
| Model | Best for |
|---|---|
| 3.7 Flash | Complex coding, agentic workflows |
| 3.1 Pro | Advanced reasoning, problem-solving |
Alibaba's Qwen family, strong at multilingual tasks and coding.
| Model | Best for |
|---|---|
| 3.8 Max | Flagship reasoning, hardest coding tasks |
| 3.7 Plus | Balanced reasoning, multimodal input |
| 3.7 Flash | Fast, high-volume tasks, low cost |
| Qwen3-Coder | Dedicated coding, large-scale MoE |
Efficient open and commercial models, including fast, low-cost options.
| Model | Best for |
|---|---|
| Medium 3.5 | Frontier agentic and coding work |
| Large 3 | General-purpose multimodal tasks |
| Small 4 | Instruct, reasoning, coding in one |
| Codestral | Dedicated code generation |
Research-driven models built around efficient reasoning at low cost.
| Model | Best for |
|---|---|
| V4.1 Flash | Complex and everyday tasks, top-tier performance at low cost |
| V4 Pro | Complex tasks, top-tier performance |
| V4 Flash | Fast, cost-efficient everyday tasks |

Grok models built for agentic tool use, coding, and fast, direct responses.
| Model | Best for |
|---|---|
| Grok 4.6 | Agentic coding and everything else |
| Grok 4.1 Fast | Fast, low-cost, 2M-token context |
Enterprise-focused models built for retrieval and business workflows.
| Model | Best for |
|---|---|
| Command A+ | Enterprise multimodal reasoning |
| Command A | Fast, efficient text, 23 languages |
Meta's Muse Spark model, built for agentic coding and tool use over the API.
| Model | Best for |
|---|---|
| Muse Spark | Agentic coding, tool loops, over the API |
Kimi models, built for long-context work and agentic tool use.
| Model | Best for |
|---|---|
| Kimi K3 | Flagship, 1M-token context window |
| Kimi K2.6 | Affordable, general-purpose tasks |
| Kimi K2.7 Code | Dedicated coding, multimodal tasks |

GLM models from Zhipu AI, tuned for high-throughput coding and reasoning.
| Model | Best for |
|---|---|
| GLM-5.3 | Flagship coding, frontier reasoning |
| GLM-5 Turbo | Fast, high-volume agentic workloads |
| GLM-4.7 | Cost-efficient everyday workhorse |

Open-weight Hunyuan models pairing large-scale MoE reasoning with a 256K context window.
| Model | Best for |
|---|---|
| Hunyuan Hy3 | Flagship reasoning and agentic tasks |
| Hunyuan A13B | Fast, cost-efficient everyday tasks |
Custom and self-hosted endpoints
Point Omnesly at any OpenAI-compatible API - including models you're running locally on your own machine. If it speaks the standard chat-completions format, Omnesly can talk to it.
Add a custom endpoint
New providers, added as they matter.
The model landscape moves fast. We add support for new providers and model families on an ongoing basis - no app update required to start using a newly connected key.