The short answer: what does the field look like in October 2026?

AI models change so quickly that a comparison from six months ago is almost entirely out of date today. In September 2026 alone, OpenAI announced GPT-6 Astra, GPT-6 Sol, Luna and GPT-6.1 Sol; Anthropic released Claude Fable 5.1, Opus 5.5 and Sonnet 5.5; Google shipped Gemini 3.8 Flash and the limited-access Gemini 4 Argon; xAI launched Grok 4.7; Meta released Muse Spark 1.3; and DeepSeek announced V4.1 Flash. This article takes a snapshot of that busy period as of 6 October 2026: which model came out when, what it costs, where it stands in independent measurements, and what is expected in the coming weeks.

The short version: on the independent human-preference ranking (Arena), the limited-access Gemini 4 Argon is on top, followed closely by Claude Opus 5.5 and Claude Fable 5.1. On Artificial Analysis's composite Intelligence Index, Claude Opus 5.5 scores highest (58), while GPT-6 Astra, Claude Fable 5.1 and Gemini 4 Argon meet at 53. On price, the gap is huge: GPT-6 Astra and Claude Fable 5.1 charge $50 per million output tokens, while GPT-6 Luna charges $0.50 and DeepSeek V4.1 Flash $1.20. So the question “which model is best?” has no single answer; you have to answer which task, which budget and which access conditions together.

Logo-free provider tiles for OpenAI, Anthropic, Google, xAI, Meta, Mistral, DeepSeek and Qwen with an October 2026 badge.
The eight providers covered in this comparison · October 2026.

Every number in this article is based on the primary source named under each table; I did not fill in values I could not verify with estimates, but left them as “unverified”. Scores from a provider's own announcement are marked separately as “vendor-reported”, because they don't carry the same weight as independent measurement.

Current models, provider by provider

OpenAI's autumn 2026 lineup has three tiers: GPT-6 Astra as the most capable model, GPT-6.1 Sol for general use and coding, and GPT-6 Luna for low-cost, high-volume work. All three offer a 1.05-million-token context window and a 128K-token output limit, accept text and image input and produce text. The real difference is price: Luna is one hundred times cheaper than Astra on both input and output.

OpenAI models
ModelReleaseContext / max outputAPI price (1M tokens, input / output)
GPT-6 Astra (gpt-6-astra)3 Sep 20261,05M / 128K$10 / $50
GPT-6.1 Sol (gpt-6.1-sol)29 Sep 20261,05M / 128K$2 / $10
GPT-6 Luna (gpt-6-luna)22 Sep 20261,05M / 128K$0.10 / $0.50

Source: providers' official model and pricing pages, checked on 6 October 2026. Prices are standard rates; cache, long-context and batch discounts excluded.

In Anthropic's Claude family, the newest broadly available model is Claude Opus 5.5, released on 22 September, followed by Sonnet 5.5 on 28 September. Claude Fable 5.1, whose access is restricted by policy, is the highest-capability and most expensive tier. Haiku 4.5 is still listed, but Anthropic has announced that a new Haiku will follow “in the coming weeks”. The whole 5.5 family offers a 1-million-token context window.

Anthropic models
ModelReleaseContext / max outputAPI price (1M tokens, input / output)
Claude Fable 5.1 (claude-fable-5-1)1 Sep 20261M / 128K$10 / $50
Claude Opus 5.5 (claude-opus-5-5)22 Sep 20261M / 128K$4 / $20
Claude Sonnet 5.5 (claude-sonnet-5-5)28 Sep 20261M / 128K$2 / $10
Claude Haiku 4.5 (claude-haiku-4-5-20251001)2025200K / 64K$1 / $5

Source: providers' official model and pricing pages, checked on 6 October 2026. Prices are standard rates; cache, long-context and batch discounts excluded.

On Google's side, the picture is a little mixed. Gemini 4 Argon, announced on 30 September, is for now available only through a programme for trusted cyber-defence teams; no public API pricing or technical specifications have been published. The newest stable model developers can use today is Gemini 3.8 Flash, released on 2 September: it handles text, image, audio and video input with a 1-million-token context and is offered at an introductory price until the end of the year.

Google Gemini models
ModelReleaseContext / max outputAPI price (1M tokens, input / output)
Gemini 4 Argon30 Sep 2026 (limited access)not publishednot published
Gemini 3.8 Flash (gemini-3.8-flash)2 Sep 20261M / 64K$0.75 / $3.75 (introductory, until 31 Dec 2026)
Gemini 3.1 Pro Previewday unverified1M / 64K$2 / $12 (below threshold)
Gemini 3.1 Flash-Liteday unverified1M / 64K$0.25 / $1.50

Source: providers' official model and pricing pages, checked on 6 October 2026. Prices are standard rates; cache, long-context and batch discounts excluded.

Beyond the big three there are strong options too. xAI's Grok 4.7 sits in the mid price band with a 500K-token context. Meta's hosted model Muse Spark 1.3 is not open-weight; Meta's newest downloadable weights are still the Llama 4 family from April 2025. Mistral's Medium 3.5, Large 3 and Small 4 are offered both through the API and as open weights. DeepSeek V4.1 Flash, with a 1M context and 384K output, is one of the cheapest strong options. Alibaba's Qwen 3.8 Max is priced in yuan, so it cannot be dropped straight into a dollar comparison.

Other providers
ModelReleaseContextAPI price (input / output)
xAI Grok 4.721 Sep 2026500K$2 / $6 (prompts under 200K)
Meta Muse Spark 1.32 Sep 2026≈1.05M$1.25 / $4.25
Meta Llama 4 Scout / Maverick5 Apr 2025 (open weights)Scout: 10M (Meta's claim)self-hosted
Mistral Medium 3.52026 (day unverified)256K$1.50 / $7.50
Mistral Small 416 Mar 2026unverified$0.15 / $0.60
DeepSeek V4.1 Flash10 Sep 20261M / 384K$0.30 / $1.20 (peak)
Qwen 3.8 Max 0902 (qwen3.8-max-0902)2 Sep 2026 (first 3.8 Max: 2 Aug)1M / 131KCNY 12 / CNY 36

Source: providers' official pages, 6 October 2026. The Qwen price is given in yuan (CNY) as on the official page and was not converted to dollars. Llama 4 is open-weight, so hosting cost depends on your hardware.

When did it come out? A timeline from 2024 to today

Two things stand out in the timeline. First, the gap between providers' main model families is shrinking: in 2024 major releases came months apart, while in September 2026 a frontier model was announced almost every week. Second, instead of a single “newest model”, different price tiers of the same family now arrive a few days apart; OpenAI's Astra, Sol and Luna and Anthropic's Fable, Opus and Sonnet are examples.

Swim-lane timeline of selected releases from February 2024 through September 2026, with the September cluster enlarged.
Source: providers' official announcements · 6 October 2026.

Looking further back also shows where today's competition came from. GPT-4o arrived in May 2024, Claude 3.5 Sonnet in June 2024 and DeepSeek-R1 in January 2025; GPT-5 came in August 2025 and Gemini 3 in November 2025. DeepSeek's open models and Mistral's open-weight series have kept up constant pressure to close the gap with closed models.

Benchmarks: what do independent measurements say?

The two independent sources most often used to compare models measure different things. Arena (formerly LMArena) is a human-preference ranking in which users blindly compare answers from two anonymous models; the score is calculated with the Elo system and shows which answer people prefer, not which is correct. The Artificial Analysis Intelligence Index runs its own test suite under the same conditions and produces a composite score; that score is not a percentage, and values can change when the index version changes.

Arena Text Elo bars sorted from Gemini 4 Argon High at 1525 ± 9 to Mistral Medium 3.5 at 1427 ± 6; axis begins at 1400.
Source: LMArena / Arena Text, checked October 6, 2026; human preference, not accuracy.
Arena Text leaderboard (6 October 2026)
Model / variantEloVotes
Gemini 4 Argon High1525 ± 94,932
Claude Opus 5.5 High1504 ± 94,552
Claude Fable 5.1 Max1501 ± 611,800
Gemini 3.8 Flash High1495 ± 526,298
Muse Spark 1.3 Max1494 ± 612,343
GPT-6.1 Sol Max1483 ± 113,071
Qwen 3.8 Max1482 ± 523,353
GPT-6 Astra Max1477 ± 79,156
DeepSeek V4.1 Flash Max1474 ± 78,728
Grok 4.7 xhigh1442 ± 86,405
Mistral Medium 3.51427 ± 611,757

Source: Arena Text Arena Overall, 6 October 2026. The ± value is the confidence interval; Arena marks the Gemini 4 Argon result as preliminary. This measures human preference, not accuracy or coding success.

Notice that the top nine models on Arena are within 51 Elo points of each other; once the confidence intervals are taken into account, many neighbouring models are not statistically separable. Also remember that the chart's axis starts at 1400, not at zero, which can make the differences look larger than they are.

Artificial Analysis composite index bars, not percentages; current Gemini 3.8 Flash v4.3.2 profile 41 and Mistral Medium 3.5 v4.3.2 profile 14. Footnote reports Gemini v4.2 at 59 and Mistral v4.3 at 15; estimated Llama 4 Maverick is excluded.
Source: Artificial Analysis, benchmark profiles and versions checked October 6, 2026; estimated Llama 4 Maverick is excluded.
Artificial Analysis Intelligence Index
Model / settingIndex score
Claude Opus 5.5 Max58
GPT-6 Astra xhigh53
Claude Fable 5.1 Max (fallback)53
Gemini 4 Argon High53
Grok 4.7 xhigh46
Muse Spark 1.3 xhigh45
Qwen 3.8 Max45
Gemini 3.8 Flash41 (v4.3.2; 59 on v4.2)
DeepSeek V4.1 Flash Max39
Mistral Medium 3.514 (v4.3.2; 15 on v4.3)

Source: Artificial Analysis model profiles and the v4.3 report, checked on 6 October 2026. A composite score, not a percentage. Different values were published for Gemini 3.8 Flash and Mistral Medium 3.5 under two index versions; I show both.

Scores from providers' own announcements are a separate category. OpenAI reports 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3 and 100% on ExploitBench for GPT-6 Astra, but has not published enough settings to reproduce those results independently. Anthropic reports 66.4% for Opus 5.5 and 70.6% for Sonnet 5.5 on Terminal-Bench 4.0, measured with different settings. In Artificial Analysis's own Terminal-Bench 4.0 run, GPT-6 Astra scored 59.1% and Claude Fable 5.1 52.0%. Putting vendor claims and independent runs in the same table is misleading.

For well-known tests such as SWE-bench Verified, GPQA Diamond, AIME, Humanity's Last Exam and MMLU-Pro, I could not find a reliable table comparing every provider's current flagship under the same settings. Rather than lining up numbers measured under different settings, I left these tests out of this comparison; if a test matters to you, check that test's own current leaderboard and settings.

Price and context window

Price is often more decisive than quality when choosing a model, especially for applications that send thousands of requests a day. The chart below shows the Artificial Analysis index score together with the price per million output tokens. It is not a “value for money verdict” but a screening tool. The prices used are published output rates, and some are conditional: the introductory price valid until 31 December 2026 for Gemini 3.8 Flash, the under-200K-token prompt price for Grok 4.7, and the peak (cache-miss) price for DeepSeek V4.1 Flash. Cache discounts and batch pricing are not included.

Log-scale USD output price versus AA Index for five models: DeepSeek V4.1 Flash $1.2/39; Gemini 3.8 Flash $3.75/41; Grok 4.7 $6/46; Claude Opus 5.5 $20/58; GPT-6 Astra $50/53. Qwen is omitted because verified USD pricing is unavailable. Price conditions: Gemini 3.8 Flash introductory until 31 Dec 2026, Grok prompt under 200K, DeepSeek peak cache-miss.
Sources: Artificial Analysis and official API pricing, 6 October 2026. Published output rates; Gemini 3.8 Flash introductory until 31 Dec 2026, Grok prompt under 200K, DeepSeek peak cache-miss. Not a value-for-money verdict.

Two clusters separate in this view. GPT-6 Astra is the most expensive model shown; Claude Fable 5.1 has the same $50 output rate and an AA score of 53, but is omitted because its point would overlap Astra. Claude Opus 5.5 gets a higher index score at a lower price. DeepSeek V4.1 Flash and Gemini 3.8 Flash deliver reasonable scores in a much cheaper band. I left Qwen 3.8 Max off the chart because it is priced in yuan; no converted dollar price exists in the official source.

To make the price gap concrete, let's work through a simple example. Imagine an application that sends 10,000 requests a month, each with an average of 2,000 input tokens and 500 output tokens. That is 20 million input and 5 million output tokens a month. The cost is 20 times the input price plus 5 times the output price. I used the official list prices; no cache discounts, long-context fees or batch pricing.

Example monthly cost: 10,000 requests (20M input + 5M output tokens)
ModelPrice (1M input / output)Monthly cost
GPT-6 Astra$10 / $50$200 + $250 = $450
Claude Fable 5.1$10 / $50$200 + $250 = $450
Claude Opus 5.5$4 / $20$80 + $100 = $180
GPT-6.1 Sol$2 / $10$40 + $50 = $90
Claude Sonnet 5.5$2 / $10$40 + $50 = $90
Grok 4.7$2 / $6$40 + $30 = $70
Gemini 3.8 Flash$0.75 / $3.75$15 + $18.75 = $33.75
DeepSeek V4.1 Flash$0.30 / $1.20$6 + $6 = $12
GPT-6 Luna$0.10 / $0.50$2 + $2.50 = $4.50

Calculation: 20 × input price + 5 × output price. Standard list prices (6 October 2026); the introductory price for Gemini 3.8 Flash, the peak price for DeepSeek and the under-200K prompt price for Grok were used.

For the same workload there is roughly a hundredfold difference between the most and least expensive options: a job that costs $450 a month on GPT-6 Astra drops to $4.50 on GPT-6 Luna. That does not mean the cheap model is good enough for every job; but for many teams the right architecture is a tiered setup that routes hard requests to an expensive model and routine requests to a cheap one. Such routing can cut the bill several times over without a serious drop in quality, which again you need to measure with your own data.

Context-window tokens: GPT-6 Astra 1,050,000; Muse Spark 1.3 1,048,576; four models 1,000,000; Grok 4.7 500,000; Mistral Medium 3.5 256,000.
Source: providers' official model specifications · 6 October 2026.

Most frontier models now meet at about 1 million tokens of context; Grok 4.7 stays at 500K and Mistral Medium 3.5 at 256K. But a large window does not mean the model recalls the whole text equally well, and because providers use different tokenizers, the same text can map to a different number of tokens. If you will process long documents, test with your own documents.

What's coming next? Announcements, reports and rumours

“When will the next model come out?” is the most searched topic with the least reliable information. That is why I split the information into three confidence levels: the provider's own official announcement, a credible press report, and an unverified rumour. Even official announcements often give no exact date.

Officially announced
WhatWhat was announcedAnnounced
GPT-6.1 Sol Ultrafast (OpenAI)API option “in the coming days”; no exact date29 Sep 2026
Claude Haiku 5.5 (Anthropic)“In the coming weeks”; no date or model ID22 Sep 2026
DeepSeek V4.1 ProV4-Pro traffic routed to Flash until V4.1-Pro launches; no date10 Sep 2026
Gemini 4 Argon wider accessLimited programme for now; no general release date30 Sep 2026

Source: the providers' announcement pages.

On the press side there are two notable reports. The Associated Press reported that OpenAI delayed planned broad access to Astra over security concerns; this is a press report, not an OpenAI commitment. Axios, citing Bloomberg, reported that some Google employees felt Gemini 4 Argon's performance fell short of expectations; Google disputed the characterisation. On the rumour side, there was an unverified community post claiming a new Qwen model would arrive in September; that month has passed and there is no official confirmation.

If there is no official date, read any report saying a model will arrive “next month” as a prediction.

Open weights or a closed model?

Another question to ask when choosing a model is whether you will use it through an API or download its weights and run it on your own server. The weights of the frontier models from OpenAI, Anthropic and Google cannot be downloaded: GPT-6 Astra is offered through ChatGPT and the API, Claude models through the Claude app and the API, and Gemini 4 Argon is for now limited to the announced trusted cyber-defence programme. Meta's newest hosted model, Muse Spark 1.3, is available through Muse Code and the Meta Model API, without open weights. By contrast, Meta's Llama 4 family, Mistral's Medium 3.5, Large 3 and Small 4, and DeepSeek's V4 series come with downloadable weights.

The biggest advantage of open weights is data control: sensitive data never leaves your organisation, the model version doesn't change unless you want it to, and you pay no per-token fee. The cost is the hardware, setup and maintenance burden; running a large model with low latency needs serious GPU resources. Licences differ too: these Mistral releases come under Apache 2.0 or a modified MIT licence, while Llama comes under Meta's own licence. Always read the licence terms before enterprise use.

Reading benchmarks correctly

A benchmark table only means something when you know how it was produced. Checking the points below while reading model comparisons helps you separate marketing figures from real performance:

  • Who measured it? A provider's own announcement and an independent leaderboard are not equally reliable.
  • Are the settings the same? Reasoning level (high, max, xhigh), tool use and number of attempts change results considerably.
  • Could test data have leaked into training data? Contamination is a known risk; a high score does not always prove general capability.
  • Is the test saturated? If most models score above 90%, the test no longer separates them.
  • Which version of the index? In composite indexes, the same model's score can change when the version changes.
  • Under which conditions is the price quoted? Cache, long-context, peak-hour and batch prices all differ.

OpenAI itself published an analysis explaining why it considers SWE-bench Verified noisy for coding comparisons. In other words, saying “this model is the best at coding” based on a single number is risky even with the most reliable sources.

Which model for which job?

The table below is based on how providers position their models and on official specifications; it is not a definitive recommendation but a shortlist to start your own tests.

Candidates by use case
Use caseCandidate modelsWhat to watch
Complex reasoning and hard codingGPT-6 Astra, Claude Opus 5.5Most expensive tier; Gemini 4 Argon only for eligible programme participants
Agentic coding workflowsGPT-6.1 Sol, Claude Sonnet 5.5, Grok 4.7Validate with your own codebase and toolchain
Long documentsGPT-6, Claude 5.5, Gemini 3.8 Flash / 3.1 Pro, DeepSeek V4.1 Flash, Qwen 3.8≈1M context; test real recall quality
Low cost, high volumeGPT-6 Luna, Gemini 3.1 Flash-Lite, DeepSeek V4.1 Flash, Mistral Small 4, Qwen 3.8 FlashCompare input and output prices separately; currencies and cache terms differ
Open weights / self-hostingLlama 4, Mistral Medium 3.5 / Large 3 / Small 4, DeepSeek V4Licence terms and hardware cost
Image, video, audioGemini 3.8 Flash, Muse Spark 1.3, Qwen 3.8Don't assume equal quality across modalities

Source: providers' model pages and announcements. Positioning is the providers' own claim, not a neutral endorsement.

In my experience, the soundest method is to run two or three candidates on your own real tasks with the same prompts and compare the results together with cost. A customer-support bot and a code-review tool have very different needs; a benchmark ranking does not show that difference. And because models are updated often, revisit your choice every few months.

Conclusion

As of October 2026, there is no single winner at the frontier. Claude Opus 5.5 leads the independent composite index and ranks near the top of Arena; Gemini 4 Argon is first on Arena but not publicly available; GPT-6 Astra is strong but expensive; and GPT-6 Luna, DeepSeek V4.1 Flash and Gemini 3.8 Flash change the game on cost. For those looking for open weights, Mistral and Llama 4 remain serious options.

Use this comparison as a starting point: the tables tell you which models to try, and a test with your own data tells you which one to choose. I will try to keep this article up to date as prices and rankings change; you will find every source I used below.

Sources

All model, date, price and score information in this article was checked against these primary sources on 6 October 2026:

  1. OpenAI: model catalog
  2. OpenAI: API pricing
  3. OpenAI: GPT-6 Astra
  4. OpenAI: GPT-6.1 Sol
  5. Anthropic: model overview
  6. Anthropic: pricing
  7. Anthropic: Claude Opus 5.5
  8. Anthropic: Claude Sonnet 5.5
  9. Google: Gemini API models
  10. Google: Gemini API pricing
  11. Google: Gemini 4 Argon
  12. xAI: Grok 4.7
  13. Meta: Muse Spark 1.3
  14. Meta: Llama 4
  15. Mistral: models and pricing
  16. DeepSeek: V4.1 Flash
  17. DeepSeek: API pricing
  18. Alibaba Model Studio: Qwen models
  19. Arena: Text leaderboard
  20. Artificial Analysis: Intelligence Index methodology
  21. Artificial Analysis: Index v4.3 report
  22. OpenAI: on noise in coding evaluations
  23. NAACL 2024: benchmark contamination

Model versions, prices and rankings change quickly; check the current state of each source before deciding. This article involves no partnership with any provider; brand and model names are used for identification only.

✳

The best model is the one you have tested and measured on your own work.

Explore more articles ↗