Free tool · Table
AI Model Comparison (2026)
Every major 2026 LLM side by side: context window, API price per 1M tokens, modality, and whether the weights are open. Filter by provider or open-weights, sort by price or context. Verified and cross-checked; long-context surcharges disclosed.
All models. 42 of 42 models listed.. Cheapest input here is $0.1 per million, largest context 1050K.
Swipe for price & modality →
| Model | Context | In $/1M | Out $/1M | Weights | Modality |
|---|---|---|---|---|---|
| Qwen 3.7 MaxAlibaba · 2026-05The previous Max tier; strong tool-use + multilingual (esp. Chinese) reasoning | 1M | $2.50* | $7.50 | Proprietary | Multimodal |
| Qwen 3.8 FlashAlibaba · 2026Alibaba's Flash tier for high-volume work, on the same multimodal API as the Max rows | 1M | $0.15* | $0.47 | Proprietary | Multimodal |
| Qwen 3.8 MaxAlibaba · 2026-09-02Alibaba's current Max tier, 20% cheaper on both halves than Qwen 3.7 Max; thinking and non-thinking modes on one model ID | 1M | $2* | $6 | Proprietary | Multimodal |
| Qwen 3.8 Omni FlashAlibaba · 2026Alibaba's audio- and video-understanding tier, priced identically to Qwen 3.8 Flash | 1M | $0.15* | $0.47 | Proprietary | Multimodal |
| Claude Fable 5Anthropic · 2026-06-09Legacy Fable model; Anthropic's own model page calls Fable 5.1 the current Fable model and recommends migrating for improved performance | 1M | $10 | $50 | Proprietary | Text + image in |
| Claude Fable 5.1Anthropic · 2026-09-01Anthropic's most capable model; demanding reasoning and long-horizon agentic work | 1M | $10* | $50 | Proprietary | Text + image in |
| Claude Haiku 4.5Anthropic · 2025-10Fastest Claude with near-frontier intelligence at low cost | 200K | $1 | $5 | Proprietary | Text + image in |
| Claude Opus 4.8Anthropic · 2026Legacy Opus model; Anthropic's own model page calls Opus 5.5 the current Opus model and recommends migrating for complex reasoning and long-horizon agentic coding | 1M | $5 | $25 | Proprietary | Text + image in |
| Claude Opus 5Anthropic · 2026-07-24Complex agentic coding and enterprise work; supersedes Opus 4.8 | 1M | $5* | $25 | Proprietary | Text + image in |
| Claude Opus 5.5Anthropic · 2026-09-22Long-running agentic coding and knowledge work; supersedes Claude Opus 5. Thinking is always on and cannot be disabled | 1M | $4* | $20 | Proprietary | Text + image in |
| Claude Sonnet 4.6Anthropic · 2026Legacy Sonnet model; Anthropic's own model page calls Sonnet 5 the current Sonnet model and recommends migrating for improved performance | 1M | $3* | $15 | Proprietary | Text + image in |
| Claude Sonnet 5Anthropic · 2026-06-30The previous Sonnet, superseded by Claude Sonnet 5.5 on 2026-09-28 and still sold at the same price | 1M | $2* | $10 | Proprietary | Text + image in |
| Claude Sonnet 5.5Anthropic · 2026-09-28The best combination of speed and intelligence in the Claude 5 family; supersedes Claude Sonnet 5 | 1M | $2* | $10 | Proprietary | Text + image in |
| DeepSeek V4DeepSeek · 2026-04-24Superseded by DeepSeek V4.1 Flash on 2026-09-10, and being phased out: DeepSeek routes deepseek-v4-pro requests to Flash from 2026-09-14 until a V4.1 Pro launches. Open-weights MoE (~32–37B active) with sparse attention | 1M | $1.32* | $3.96 | Open | Text |
| DeepSeek V4.1 FlashDeepSeek · 2026-09-10552B MoE that activates only 8B parameters on input and 16B on output, with native visual understanding; DeepSeek's own benchmarks put it ahead of V4 Pro on quality, cost and speed at a quarter of the price | 1M | $0.30* | $1.20 | Open | Multimodal |
| Gemini 3.1 ProGoogle · 2026-02-19Frontier long-context multimodal reasoning; largest practical context in class | 1M | $2* | $12 | Proprietary | Multimodal (text/image/audio/video) |
| Gemini 3.5 FlashGoogle · 2026-05-20Google's legacy Flash tier: baseline speed for routine high-throughput work | 1M | $1.50* | $9 | Proprietary | Multimodal (text/image/audio/video) |
| Gemini 3.8 FlashGoogle · 2026-09-02Google's current Flash generation; even at its 2027 standard price it undercuts the Gemini 3.5 Flash it supersedes on output ($7.50 against $9.00) | 1M | $0.75* | $3.75 | Proprietary | Multimodal (text/image/audio/video) |
| Llama 4 MaverickMeta · 2025-04-05Open-weights MoE with a 1M window and native vision; a large fine-tuning ecosystem and no licence fee to self-host | 1M | $0.20* | $0.80 | Open | Multimodal (text + image) |
| Llama 4 ScoutMeta · 2025-04-05Smaller open-weights MoE (17B active, 16 experts) that fits on a single H100; the cheapest row in this ledger | 320K | $0.10* | $0.30 | Open | Multimodal (text + image) |
| Mistral Large 3Mistral AI · 2025-12-02Apache-2.0 open weights, 675B total / 41B active MoE; general-purpose work over text and image input | 256K | $0.50* | $1.50 | Open | Text + image |
| Mistral Medium 3.5Mistral AI · 2026-04-28Mistral's frontier-class model for agentic and coding work; open weights under a Modified MIT licence | 256K | $1.50* | $7.50 | Open | Multimodal |
| Mistral Small 4Mistral AI · 2026-03-16Apache-2.0 open weights, 119B total / 6B active MoE with configurable reasoning effort; the cheapest Mistral row in this ledger | 256K | $0.15* | $0.60 | Open | Text + image |
| Kimi K3Moonshot AI · 2026Open-weights 2.8T-param MoE (104B active) under the Kimi K3 License; 1M context for long-horizon coding | 1M | $3* | $15 | Open | Text (agentic) |
| GPT-5.4 miniOpenAI · 2026Cost-efficient mid-tier for high-volume tasks with solid reasoning | 400K | $0.75 | $4.50 | Proprietary | Text + image in |
| GPT-5.4 nanoOpenAI · 2026OpenAI's small, latency-focused tier for simple, high-volume calls; GPT-5.6 Luna matches its input price and undercuts its output, so it is no longer the cheapest OpenAI row | 400K | $0.20 | $1.25 | Proprietary | Text + image in |
| GPT-5.5OpenAI · 2026-04-23Superseded as OpenAI's flagship by GPT-6 Astra on 2026-09-05; first OpenAI model with a 1M-token API context window | 1.05M | $5* | $30 | Proprietary | Text + image in |
| GPT-5.5 ProOpenAI · 2026-04-23The extra-compute tier of GPT-5.5: it spends more tokens thinking before answering, and OpenAI warns a request may take several minutes and should use background mode | 1.05M | $30* | $180 | Proprietary | Text + image in |
| GPT-5.6 LunaOpenAI · 2026-07-09Fastest, lowest-cost GPT-5.6 tier for high-volume work | 1.05M | $0.20* | $1.20 | Proprietary | Text + image in |
| GPT-5.6 SolOpenAI · 2026-07-09Top GPT-5.6 tier for complex coding, research, and agentic work; below GPT-6 Astra since 2026-09-05 | 1.05M | $4* | $20 | Proprietary | Text + image in |
| GPT-5.6 TerraOpenAI · 2026-07-09Balanced GPT-5.6 tier: near-GPT-5.5 quality at 40% of GPT-5.5's price, in and out | 1.05M | $2* | $12 | Proprietary | Text + image in |
| GPT-6 AstraOpenAI · 2026-09-05OpenAI's most capable model: complex reasoning, coding, computer use, research, and document creation | 1.05M | $10* | $50 | Proprietary | Text + image in |
| GPT-6 LunaOpenAI · 2026-09-22The GPT-6 tier OpenAI positions for focused, high-volume tasks | 1.05M | $0.10* | $0.50 | Proprietary | Text + image in |
| GPT-6 SolOpenAI · 2026-09-22The GPT-6 tier OpenAI builds for complex coding and agentic workflows | 1.05M | $2* | $10 | Proprietary | Text + image in |
| GPT-6.1 SolOpenAI · 2026-09-29Near-Astra performance for complex coding, computer use and professional work at a lower cost; reasoning effort low to max | 1.05M | $2* | $10 | Proprietary | Text + image in |
| Grok 4.3xAI · 2026xAI's widest context at 1M tokens and its cheapest reasoning tier; native real-time X / web search | 1M | $1.25* | $2.50 | Proprietary | Text + image in |
| Grok 4.6xAI · 2026-08-12The flagship Grok 4.7 superseded on 2026-09-21, at the same price and the same 500K window; still generally available, reasoning effort up to xhigh | 500K | $2* | $6 | Proprietary | Text + image in |
| Grok 4.7xAI · 2026-09-21xAI's current flagship for coding, agentic tasks and knowledge work; agentic tool calling and configurable reasoning effort up to xhigh | 500K | $2* | $6 | Proprietary | Text + image in |
| GLM-5Zhipu / Z.ai · 2026-02-11Open MIT-licensed frontier model; very strong coding + agentic value | 200K | $1* | $3.20 | Open | Multimodal |
| GLM-5.3Zhipu / Z.ai · 2026Z.ai's current flagship, and what its own docs tell GLM-5 users to migrate to; 1M context, weights public on zai-org under a non-OSI licence | 1M | $1.40* | $4.40 | Open | Text |
| GLM-5.3-FlashZhipu / Z.ai · 2026MIT-licensed open weights with the same 1M context as the flagship at roughly a tenth of its rate; the open side of this ledger's 1M-context price gap | 1M | $0.15* | $0.50 | Open | Multimodal (video/image/text/file in) |
| GLM-5.3-FlashXZhipu / Z.ai · 2026The faster half of the first native multimodal pair in the GLM-5 series; API-only, with no published weights unlike GLM-5.3 and GLM-5.3-Flash | 1M | $0.37* | $1.25 | Proprietary | Multimodal (video/image/text/file in) |
Standard pay-as-you-go list prices, USD per 1M tokens, as of 3 Sep 2026, for prompts up to ~200K tokens. * = tiered/long-context or promo pricing applies (hover the input price). Open-weights API prices are representative third-party hosted rates; self-hosting is free. This space moves weekly, so confirm with each provider before relying on a figure.
Frequently asked
Which 2026 model has the cheapest API pricing?
Among frontier-class options, third-party-hosted open-weights models are cheapest: Llama 4 Maverick runs about $0.20 / $0.80 per 1M tokens (input / output). Among major proprietary APIs, OpenAI's GPT-5.4 nano ($0.20 / $1.25) and Google's smaller Flash tiers are the lowest. DeepSeek V4.1 Flash is cheaper still: it lists at $0.30 / $1.20 and DeepSeek halves both halves outside peak hours.
Which model has the largest context window?
By the verified API context windows in this table, GPT-6 Astra is the widest at ~1.05M tokens, tied there by eight other OpenAI rows. Just behind sits a ~1M cluster: Claude Fable 5.1, Gemini 3.1 Pro, Grok 4.3, and five open-weights rows including DeepSeek V4.1 Flash. Google and Meta have publicly discussed larger windows (a Gemini tier up to ~2M, a Llama 4 Scout variant up to 10M), but those aren't reflected in these standard API rows.
What are the best open-weights models in mid-2026?
This table carries ten open-weights rows the ledger still treats as current, among them DeepSeek V4.1 Flash, Llama 4 Maverick, Kimi K3, GLM-5.3, GLM-5.3-Flash. Several rival proprietary frontier models on coding and agentic tasks, and all can be self-hosted or run via many providers. Note that open weights and an open licence are not the same thing, and it varies row by row rather than by maker: Mistral Large 3 ships under Apache-2.0 and GLM-5.3-Flash under MIT, while GLM-5.3 — the current Z.ai flagship — publishes its weights under a non-OSI licence, and Kimi K3 under Moonshot's own Kimi K3 License.
Why are some prices shown with an asterisk and a footnote?
Several flagships use context-tiered pricing: the listed figure is the rate for prompts up to ~200K tokens, and the footnote shows the higher long-context rate. For example, Gemini 3.1 Pro lists $2 / $12 at the base tier but >200K billed $4/$18 (long-context tier); GPT-6 Astra charges more on long prompts (>272K input reprices the FULL request at 2x input / 1.5x output ($20 in / $75 out per 1M); cached input $1/1M and cache writes $12.50/1M (1.25x uncached input); Batch and Flex 50% of standard, Fast mode 2x; max input is 922K of the 1,050K window). Claude Fable 5, by contrast, carries no surcharge note at all.
Is Grok 5 available yet?
Not as a generally available API as of 2026-09-21, when xAI's model catalogue was last read for this table. Grok 5 has been widely discussed (rumored ~6T params, 1.5M context) but xAI has published no official release or pricing for it. The current GA flagship is Grok 4.7, released 2026-09-21 at $2 / $6 per 1M tokens with a 500K-token window. Grok 4.3 stays in this table because it is the widest Grok window at 1M and the cheapest of the three.