The short answer: what does the field look like in October 2026?
AI models change so quickly that a comparison from six months ago is almost entirely out of date today. In September 2026 alone, OpenAI announced GPT-6 Astra, GPT-6 Sol, Luna and GPT-6.1 Sol; Anthropic released Claude Fable 5.1, Opus 5.5 and Sonnet 5.5; Google shipped Gemini 3.8 Flash and the limited-access Gemini 4 Argon; xAI launched Grok 4.7; Meta released Muse Spark 1.3; and DeepSeek announced V4.1 Flash. This article takes a snapshot of that busy period as of 6 October 2026: which model came out when, what it costs, where it stands in independent measurements, and what is expected in the coming weeks.
The short version: on the independent human-preference ranking (Arena), the limited-access Gemini 4 Argon is on top, followed closely by Claude Opus 5.5 and Claude Fable 5.1. On Artificial Analysis's composite Intelligence Index, Claude Opus 5.5 scores highest (58), while GPT-6 Astra, Claude Fable 5.1 and Gemini 4 Argon meet at 53. On price, the gap is huge: GPT-6 Astra and Claude Fable 5.1 charge $50 per million output tokens, while GPT-6 Luna charges $0.50 and DeepSeek V4.1 Flash $1.20. So the question “which model is best?” has no single answer; you have to answer which task, which budget and which access conditions together.
Every number in this article is based on the primary source named under each table; I did not fill in values I could not verify with estimates, but left them as “unverified”. Scores from a provider's own announcement are marked separately as “vendor-reported”, because they don't carry the same weight as independent measurement.
Current models, provider by provider
OpenAI's autumn 2026 lineup has three tiers: GPT-6 Astra as the most capable model, GPT-6.1 Sol for general use and coding, and GPT-6 Luna for low-cost, high-volume work. All three offer a 1.05-million-token context window and a 128K-token output limit, accept text and image input and produce text. The real difference is price: Luna is one hundred times cheaper than Astra on both input and output.
| Model | Release | Context / max output | API price (1M tokens, input / output) |
|---|---|---|---|
GPT-6 Astra (gpt-6-astra) | 3 Sep 2026 | 1,05M / 128K | $10 / $50 |
GPT-6.1 Sol (gpt-6.1-sol) | 29 Sep 2026 | 1,05M / 128K | $2 / $10 |
GPT-6 Luna (gpt-6-luna) | 22 Sep 2026 | 1,05M / 128K | $0.10 / $0.50 |
Source: providers' official model and pricing pages, checked on 6 October 2026. Prices are standard rates; cache, long-context and batch discounts excluded.
In Anthropic's Claude family, the newest broadly available model is Claude Opus 5.5, released on 22 September, followed by Sonnet 5.5 on 28 September. Claude Fable 5.1, whose access is restricted by policy, is the highest-capability and most expensive tier. Haiku 4.5 is still listed, but Anthropic has announced that a new Haiku will follow “in the coming weeks”. The whole 5.5 family offers a 1-million-token context window.
| Model | Release | Context / max output | API price (1M tokens, input / output) |
|---|---|---|---|
Claude Fable 5.1 (claude-fable-5-1) | 1 Sep 2026 | 1M / 128K | $10 / $50 |
Claude Opus 5.5 (claude-opus-5-5) | 22 Sep 2026 | 1M / 128K | $4 / $20 |
Claude Sonnet 5.5 (claude-sonnet-5-5) | 28 Sep 2026 | 1M / 128K | $2 / $10 |
Claude Haiku 4.5 (claude-haiku-4-5-20251001) | 2025 | 200K / 64K | $1 / $5 |
Source: providers' official model and pricing pages, checked on 6 October 2026. Prices are standard rates; cache, long-context and batch discounts excluded.
On Google's side, the picture is a little mixed. Gemini 4 Argon, announced on 30 September, is for now available only through a programme for trusted cyber-defence teams; no public API pricing or technical specifications have been published. The newest stable model developers can use today is Gemini 3.8 Flash, released on 2 September: it handles text, image, audio and video input with a 1-million-token context and is offered at an introductory price until the end of the year.
| Model | Release | Context / max output | API price (1M tokens, input / output) |
|---|---|---|---|
| Gemini 4 Argon | 30 Sep 2026 (limited access) | not published | not published |
Gemini 3.8 Flash (gemini-3.8-flash) | 2 Sep 2026 | 1M / 64K | $0.75 / $3.75 (introductory, until 31 Dec 2026) |
| Gemini 3.1 Pro Preview | day unverified | 1M / 64K | $2 / $12 (below threshold) |
| Gemini 3.1 Flash-Lite | day unverified | 1M / 64K | $0.25 / $1.50 |
Source: providers' official model and pricing pages, checked on 6 October 2026. Prices are standard rates; cache, long-context and batch discounts excluded.
Beyond the big three there are strong options too. xAI's Grok 4.7 sits in the mid price band with a 500K-token context. Meta's hosted model Muse Spark 1.3 is not open-weight; Meta's newest downloadable weights are still the Llama 4 family from April 2025. Mistral's Medium 3.5, Large 3 and Small 4 are offered both through the API and as open weights. DeepSeek V4.1 Flash, with a 1M context and 384K output, is one of the cheapest strong options. Alibaba's Qwen 3.8 Max is priced in yuan, so it cannot be dropped straight into a dollar comparison.
| Model | Release | Context | API price (input / output) |
|---|---|---|---|
| xAI Grok 4.7 | 21 Sep 2026 | 500K | $2 / $6 (prompts under 200K) |
| Meta Muse Spark 1.3 | 2 Sep 2026 | ≈1.05M | $1.25 / $4.25 |
| Meta Llama 4 Scout / Maverick | 5 Apr 2025 (open weights) | Scout: 10M (Meta's claim) | self-hosted |
| Mistral Medium 3.5 | 2026 (day unverified) | 256K | $1.50 / $7.50 |
| Mistral Small 4 | 16 Mar 2026 | unverified | $0.15 / $0.60 |
| DeepSeek V4.1 Flash | 10 Sep 2026 | 1M / 384K | $0.30 / $1.20 (peak) |
Qwen 3.8 Max 0902 (qwen3.8-max-0902) | 2 Sep 2026 (first 3.8 Max: 2 Aug) | 1M / 131K | CNY 12 / CNY 36 |
Source: providers' official pages, 6 October 2026. The Qwen price is given in yuan (CNY) as on the official page and was not converted to dollars. Llama 4 is open-weight, so hosting cost depends on your hardware.
When did it come out? A timeline from 2024 to today
Two things stand out in the timeline. First, the gap between providers' main model families is shrinking: in 2024 major releases came months apart, while in September 2026 a frontier model was announced almost every week. Second, instead of a single “newest model”, different price tiers of the same family now arrive a few days apart; OpenAI's Astra, Sol and Luna and Anthropic's Fable, Opus and Sonnet are examples.
Looking further back also shows where today's competition came from. GPT-4o arrived in May 2024, Claude 3.5 Sonnet in June 2024 and DeepSeek-R1 in January 2025; GPT-5 came in August 2025 and Gemini 3 in November 2025. DeepSeek's open models and Mistral's open-weight series have kept up constant pressure to close the gap with closed models.
Benchmarks: what do independent measurements say?
The two independent sources most often used to compare models measure different things. Arena (formerly LMArena) is a human-preference ranking in which users blindly compare answers from two anonymous models; the score is calculated with the Elo system and shows which answer people prefer, not which is correct. The Artificial Analysis Intelligence Index runs its own test suite under the same conditions and produces a composite score; that score is not a percentage, and values can change when the index version changes.
| Model / variant | Elo | Votes |
|---|---|---|
| Gemini 4 Argon High | 1525 ± 9 | 4,932 |
| Claude Opus 5.5 High | 1504 ± 9 | 4,552 |
| Claude Fable 5.1 Max | 1501 ± 6 | 11,800 |
| Gemini 3.8 Flash High | 1495 ± 5 | 26,298 |
| Muse Spark 1.3 Max | 1494 ± 6 | 12,343 |
| GPT-6.1 Sol Max | 1483 ± 11 | 3,071 |
| Qwen 3.8 Max | 1482 ± 5 | 23,353 |
| GPT-6 Astra Max | 1477 ± 7 | 9,156 |
| DeepSeek V4.1 Flash Max | 1474 ± 7 | 8,728 |
| Grok 4.7 xhigh | 1442 ± 8 | 6,405 |
| Mistral Medium 3.5 | 1427 ± 6 | 11,757 |
Source: Arena Text Arena Overall, 6 October 2026. The ± value is the confidence interval; Arena marks the Gemini 4 Argon result as preliminary. This measures human preference, not accuracy or coding success.
Notice that the top nine models on Arena are within 51 Elo points of each other; once the confidence intervals are taken into account, many neighbouring models are not statistically separable. Also remember that the chart's axis starts at 1400, not at zero, which can make the differences look larger than they are.
| Model / setting | Index score |
|---|---|
| Claude Opus 5.5 Max | 58 |
| GPT-6 Astra xhigh | 53 |
| Claude Fable 5.1 Max (fallback) | 53 |
| Gemini 4 Argon High | 53 |
| Grok 4.7 xhigh | 46 |
| Muse Spark 1.3 xhigh | 45 |
| Qwen 3.8 Max | 45 |
| Gemini 3.8 Flash | 41 (v4.3.2; 59 on v4.2) |
| DeepSeek V4.1 Flash Max | 39 |
| Mistral Medium 3.5 | 14 (v4.3.2; 15 on v4.3) |
Source: Artificial Analysis model profiles and the v4.3 report, checked on 6 October 2026. A composite score, not a percentage. Different values were published for Gemini 3.8 Flash and Mistral Medium 3.5 under two index versions; I show both.
Scores from providers' own announcements are a separate category. OpenAI reports 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3 and 100% on ExploitBench for GPT-6 Astra, but has not published enough settings to reproduce those results independently. Anthropic reports 66.4% for Opus 5.5 and 70.6% for Sonnet 5.5 on Terminal-Bench 4.0, measured with different settings. In Artificial Analysis's own Terminal-Bench 4.0 run, GPT-6 Astra scored 59.1% and Claude Fable 5.1 52.0%. Putting vendor claims and independent runs in the same table is misleading.
For well-known tests such as SWE-bench Verified, GPQA Diamond, AIME, Humanity's Last Exam and MMLU-Pro, I could not find a reliable table comparing every provider's current flagship under the same settings. Rather than lining up numbers measured under different settings, I left these tests out of this comparison; if a test matters to you, check that test's own current leaderboard and settings.
Price and context window
Price is often more decisive than quality when choosing a model, especially for applications that send thousands of requests a day. The chart below shows the Artificial Analysis index score together with the price per million output tokens. It is not a “value for money verdict” but a screening tool. The prices used are published output rates, and some are conditional: the introductory price valid until 31 December 2026 for Gemini 3.8 Flash, the under-200K-token prompt price for Grok 4.7, and the peak (cache-miss) price for DeepSeek V4.1 Flash. Cache discounts and batch pricing are not included.
Two clusters separate in this view. GPT-6 Astra is the most expensive model shown; Claude Fable 5.1 has the same $50 output rate and an AA score of 53, but is omitted because its point would overlap Astra. Claude Opus 5.5 gets a higher index score at a lower price. DeepSeek V4.1 Flash and Gemini 3.8 Flash deliver reasonable scores in a much cheaper band. I left Qwen 3.8 Max off the chart because it is priced in yuan; no converted dollar price exists in the official source.
To make the price gap concrete, let's work through a simple example. Imagine an application that sends 10,000 requests a month, each with an average of 2,000 input tokens and 500 output tokens. That is 20 million input and 5 million output tokens a month. The cost is 20 times the input price plus 5 times the output price. I used the official list prices; no cache discounts, long-context fees or batch pricing.
| Model | Price (1M input / output) | Monthly cost |
|---|---|---|
| GPT-6 Astra | $10 / $50 | $200 + $250 = $450 |
| Claude Fable 5.1 | $10 / $50 | $200 + $250 = $450 |
| Claude Opus 5.5 | $4 / $20 | $80 + $100 = $180 |
| GPT-6.1 Sol | $2 / $10 | $40 + $50 = $90 |
| Claude Sonnet 5.5 | $2 / $10 | $40 + $50 = $90 |
| Grok 4.7 | $2 / $6 | $40 + $30 = $70 |
| Gemini 3.8 Flash | $0.75 / $3.75 | $15 + $18.75 = $33.75 |
| DeepSeek V4.1 Flash | $0.30 / $1.20 | $6 + $6 = $12 |
| GPT-6 Luna | $0.10 / $0.50 | $2 + $2.50 = $4.50 |
Calculation: 20 × input price + 5 × output price. Standard list prices (6 October 2026); the introductory price for Gemini 3.8 Flash, the peak price for DeepSeek and the under-200K prompt price for Grok were used.
For the same workload there is roughly a hundredfold difference between the most and least expensive options: a job that costs $450 a month on GPT-6 Astra drops to $4.50 on GPT-6 Luna. That does not mean the cheap model is good enough for every job; but for many teams the right architecture is a tiered setup that routes hard requests to an expensive model and routine requests to a cheap one. Such routing can cut the bill several times over without a serious drop in quality, which again you need to measure with your own data.
Most frontier models now meet at about 1 million tokens of context; Grok 4.7 stays at 500K and Mistral Medium 3.5 at 256K. But a large window does not mean the model recalls the whole text equally well, and because providers use different tokenizers, the same text can map to a different number of tokens. If you will process long documents, test with your own documents.
What's coming next? Announcements, reports and rumours
“When will the next model come out?” is the most searched topic with the least reliable information. That is why I split the information into three confidence levels: the provider's own official announcement, a credible press report, and an unverified rumour. Even official announcements often give no exact date.
| What | What was announced | Announced |
|---|---|---|
| GPT-6.1 Sol Ultrafast (OpenAI) | API option “in the coming days”; no exact date | 29 Sep 2026 |
| Claude Haiku 5.5 (Anthropic) | “In the coming weeks”; no date or model ID | 22 Sep 2026 |
| DeepSeek V4.1 Pro | V4-Pro traffic routed to Flash until V4.1-Pro launches; no date | 10 Sep 2026 |
| Gemini 4 Argon wider access | Limited programme for now; no general release date | 30 Sep 2026 |
Source: the providers' announcement pages.
On the press side there are two notable reports. The Associated Press reported that OpenAI delayed planned broad access to Astra over security concerns; this is a press report, not an OpenAI commitment. Axios, citing Bloomberg, reported that some Google employees felt Gemini 4 Argon's performance fell short of expectations; Google disputed the characterisation. On the rumour side, there was an unverified community post claiming a new Qwen model would arrive in September; that month has passed and there is no official confirmation.
If there is no official date, read any report saying a model will arrive “next month” as a prediction.
Open weights or a closed model?
Another question to ask when choosing a model is whether you will use it through an API or download its weights and run it on your own server. The weights of the frontier models from OpenAI, Anthropic and Google cannot be downloaded: GPT-6 Astra is offered through ChatGPT and the API, Claude models through the Claude app and the API, and Gemini 4 Argon is for now limited to the announced trusted cyber-defence programme. Meta's newest hosted model, Muse Spark 1.3, is available through Muse Code and the Meta Model API, without open weights. By contrast, Meta's Llama 4 family, Mistral's Medium 3.5, Large 3 and Small 4, and DeepSeek's V4 series come with downloadable weights.
The biggest advantage of open weights is data control: sensitive data never leaves your organisation, the model version doesn't change unless you want it to, and you pay no per-token fee. The cost is the hardware, setup and maintenance burden; running a large model with low latency needs serious GPU resources. Licences differ too: these Mistral releases come under Apache 2.0 or a modified MIT licence, while Llama comes under Meta's own licence. Always read the licence terms before enterprise use.
Reading benchmarks correctly
A benchmark table only means something when you know how it was produced. Checking the points below while reading model comparisons helps you separate marketing figures from real performance:
- Who measured it? A provider's own announcement and an independent leaderboard are not equally reliable.
- Are the settings the same? Reasoning level (high, max, xhigh), tool use and number of attempts change results considerably.
- Could test data have leaked into training data? Contamination is a known risk; a high score does not always prove general capability.
- Is the test saturated? If most models score above 90%, the test no longer separates them.
- Which version of the index? In composite indexes, the same model's score can change when the version changes.
- Under which conditions is the price quoted? Cache, long-context, peak-hour and batch prices all differ.
OpenAI itself published an analysis explaining why it considers SWE-bench Verified noisy for coding comparisons. In other words, saying “this model is the best at coding” based on a single number is risky even with the most reliable sources.
Which model for which job?
The table below is based on how providers position their models and on official specifications; it is not a definitive recommendation but a shortlist to start your own tests.
| Use case | Candidate models | What to watch |
|---|---|---|
| Complex reasoning and hard coding | GPT-6 Astra, Claude Opus 5.5 | Most expensive tier; Gemini 4 Argon only for eligible programme participants |
| Agentic coding workflows | GPT-6.1 Sol, Claude Sonnet 5.5, Grok 4.7 | Validate with your own codebase and toolchain |
| Long documents | GPT-6, Claude 5.5, Gemini 3.8 Flash / 3.1 Pro, DeepSeek V4.1 Flash, Qwen 3.8 | ≈1M context; test real recall quality |
| Low cost, high volume | GPT-6 Luna, Gemini 3.1 Flash-Lite, DeepSeek V4.1 Flash, Mistral Small 4, Qwen 3.8 Flash | Compare input and output prices separately; currencies and cache terms differ |
| Open weights / self-hosting | Llama 4, Mistral Medium 3.5 / Large 3 / Small 4, DeepSeek V4 | Licence terms and hardware cost |
| Image, video, audio | Gemini 3.8 Flash, Muse Spark 1.3, Qwen 3.8 | Don't assume equal quality across modalities |
Source: providers' model pages and announcements. Positioning is the providers' own claim, not a neutral endorsement.
In my experience, the soundest method is to run two or three candidates on your own real tasks with the same prompts and compare the results together with cost. A customer-support bot and a code-review tool have very different needs; a benchmark ranking does not show that difference. And because models are updated often, revisit your choice every few months.
Conclusion
As of October 2026, there is no single winner at the frontier. Claude Opus 5.5 leads the independent composite index and ranks near the top of Arena; Gemini 4 Argon is first on Arena but not publicly available; GPT-6 Astra is strong but expensive; and GPT-6 Luna, DeepSeek V4.1 Flash and Gemini 3.8 Flash change the game on cost. For those looking for open weights, Mistral and Llama 4 remain serious options.
Use this comparison as a starting point: the tables tell you which models to try, and a test with your own data tells you which one to choose. I will try to keep this article up to date as prices and rankings change; you will find every source I used below.
Sources
All model, date, price and score information in this article was checked against these primary sources on 6 October 2026:
- OpenAI: model catalog
- OpenAI: API pricing
- OpenAI: GPT-6 Astra
- OpenAI: GPT-6.1 Sol
- Anthropic: model overview
- Anthropic: pricing
- Anthropic: Claude Opus 5.5
- Anthropic: Claude Sonnet 5.5
- Google: Gemini API models
- Google: Gemini API pricing
- Google: Gemini 4 Argon
- xAI: Grok 4.7
- Meta: Muse Spark 1.3
- Meta: Llama 4
- Mistral: models and pricing
- DeepSeek: V4.1 Flash
- DeepSeek: API pricing
- Alibaba Model Studio: Qwen models
- Arena: Text leaderboard
- Artificial Analysis: Intelligence Index methodology
- Artificial Analysis: Index v4.3 report
- OpenAI: on noise in coding evaluations
- NAACL 2024: benchmark contamination
Model versions, prices and rankings change quickly; check the current state of each source before deciding. This article involves no partnership with any provider; brand and model names are used for identification only.