TL;DR
Get the little things that make your day delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
September 2026 brought five frontier models from Anthropic and OpenAI in three weeks. On the current Artificial Analysis Intelligence Index (v4.3), Claude Opus 5.5 posts the top score at 58, while GPT-6 Luna handles tasks for $0.07. Cost per task, not headline benchmarks, is the deciding factor for buyers.
Five frontier AI models released within three weeks in September 2026 — Claude Opus 5.5, Claude Fable 5.1, GPT-6 Astra, GPT-6 Sol and GPT-6 Luna — have been compared on a single independent yardstick. Anthropic’s Opus 5.5, released September 22, posts the highest score measured to date by the Artificial Analysis Intelligence Index at 58, while OpenAI’s GPT-6 Luna completes index tasks for roughly $0.07 each, placing the choice between them on workload and budget rather than raw intelligence.
Anthropic and OpenAI each shipped a full model lineup in September. Anthropic released Claude Fable 5.1 on September 1 and Claude Opus 5.5 on September 22. OpenAI released GPT-6 Astra on September 3, then GPT-6 Sol and GPT-6 Luna on the same day as Opus 5.5, September 22. All five were measured on the current version of the Artificial Analysis Intelligence Index (v4.3).
On pricing and performance, the models span an enormous range. Opus 5.5 charges $4 input / $20 output per million tokens and scores 58 at max effort, at a cost of $5.98 per index task. Fable 5.1, priced at $10/$50, scores 53 at $7.63 per task — outscored by its cheaper sibling at higher cost. GPT-6 Astra, also $10/$50, matches Fable’s 53 score but at $3.26 per task, less than half Fable’s cost. GPT-6 Sol at $2/$10 scores 48 for $1.06 per task, and GPT-6 Luna at $0.10/$0.50 scores 37 for about seven cents per task.
Expressed as throughput, $100 at max effort buys roughly 17 tasks on Opus 5.5, 13 on Fable 5.1, 31 on Astra, 94 on Sol and about 1,429 on Luna. The comparison indicates a practical split: Opus 5.5 for client-ready knowledge work and long agentic coding, Astra for computer use and token-efficient agents, Sol for scaled coding and business automation, and Luna for high-volume classification, extraction and routing.
Frontier AI · September 2026
Five Frontier Models, One Bill
Three weeks in September 2026 delivered five frontier releases from Anthropic and OpenAI — Opus 5.5, Fable 5.1, GPT-6 Astra, Sol and Luna. Measured on a single independent yardstick, the Artificial Analysis Intelligence Index v4.3, one model leads on score, another costs $0.07 a task. The deciding factor for buyers is cost per task, not headline benchmarks.
All scores: Index v4.3 · Max effortThe Scoreboard
Price, Score and Cost per Task
The five models span an enormous range: from $10/$50 per million tokens down to $0.10/$0.50 — and per-task costs from $7.63 to seven cents. Fable 5.1 is outscored by its cheaper sibling at higher cost; Astra matches Fable’s score at less than half the price per task.
| Model | Lab | Input / Output ($/M tok) | Index Score | Cost per Task | Tasks per $100 |
|---|---|---|---|---|---|
| Claude Opus 5.5Released Sep 22 | Anthropic | $4 / $20 | 58 | $5.98 | ~17 |
| Claude Fable 5.1Released Sep 1 | Anthropic | $10 / $50 | 53 | $7.63 | ~13 |
| GPT-6 AstraReleased Sep 3 | OpenAI | $10 / $50 | 53 | $3.26 | ~31 |
| GPT-6 SolReleased Sep 22 | OpenAI | $2 / $10 | 48 | $1.06 | ~94 |
| GPT-6 LunaReleased Sep 22 | OpenAI | $0.10 / $0.50 | 37 | $0.07 | ~1,429 |
Throughput
What $100 Buys at Max Effort
Expressed as tasks completed, the same $100 stretches from 17 tasks on Opus 5.5 to roughly 1,429 on GPT-6 Luna — an 84× spread across a single release cycle.
The Economics
Why Cost per Task Decides the Winner
Raw intelligence scores matter less than quality per dollar on a specific workload. Opus 5.5’s per-token price is less than half of Astra’s — yet at max effort it costs nearly twice as much per task, because it writes far more: roughly 119,000 output tokens per task versus about 27,000 for Astra.
~119,000 output tokens / task
Long, thorough answers drive the per-task bill to $5.98 despite cheap tokens. Best for client-ready knowledge work and long agentic coding where quality justifies volume.
~27,000 output tokens / task
Concise output keeps per-task cost at $3.26 on identical per-token pricing to Fable 5.1. Best suited for computer use and token-efficient agents.
“Opus 5.5 performs at the level of Fable 5.1 on most work at 40% less cost than Opus 5.”
— Anthropic, as cited in the comparisonThe results also differ from launch claims: the independent index places Opus 5.5 ahead of Fable 5.1 at 58 versus 53 — while Fable, which set a record at launch, is now outscored by a cheaper sibling in the same lineup.
Measurement Caveat
The Index Re-Basing Issue
Artificial Analysis rebuilt its Intelligence Index in the same month these models shipped. When Fable 5.1 launched, it scored 66 on the older v4.1.1 index — the highest recorded at the time. On the current v4.3 index, the same model scores 53, tied with GPT-6 Astra. The model did not get worse; the test got harder. Comparisons mixing launch-week numbers are effectively using different rulers.
Anthropic fallback behavior
Safety-flagged requests route to older Claude models, which can depress scores on security and biology work.
Max effort isn’t the default
Most real deployments run at medium or high effort, where both rankings and costs differ from the max-effort table.
Practical Split
Which Model for Which Job
Knowledge work
Client-ready deliverables and long agentic coding.
Agents
Computer use and token-efficient agent pipelines.
Scaled coding
High-volume coding and business automation.
High volume
Classification, extraction and routing at seven cents a task.
Action Plan
What Buyers Should Do Now
The analysis recommends workload testing rather than benchmark shopping — run a shadow test on your own tasks before switching models.
Medium-effort trials
Start with Opus 5.5 at medium effort — that’s where most real deployments run.
Test Astra
Evaluate GPT-6 Astra where token efficiency and output length matter.
Stress Sol & Luna
Evaluate the budget pair on high-volume pipelines and routable work.
Shadow test
Index scores alone can’t confirm regressions on your workload — test on your own tasks.
“Cost per task, not headline benchmarks, is the deciding factor for buyers.”
— The comparison’s core findingKey Questions
Frequently Asked
Which of the five scored highest?
Claude Opus 5.5 at 58 on Index v4.3 at max effort — the highest score the index has measured, ahead of Fable 5.1 and GPT-6 Astra, both at 53.
Why did Fable 5.1’s score drop from 66 to 53?
The index changed. Fable scored 66 on v4.1.1 and 53 on v4.3. The model did not get worse — the test got harder, so scores from different versions cannot be compared.
What does $100 buy on each model?
Roughly 17 tasks on Opus 5.5, 13 on Fable 5.1, 31 on Astra, 94 on Sol and about 1,429 on Luna — all at max effort.
Which model is best for budget deployments?
Sol ($1.06/task, score 48) suits scaled coding and automation; Luna ($0.07/task, score 37) suits classification, extraction and routing. Note Sol regressed on polished deliverables and Luna’s coding score slipped versus its predecessor.
Why does Opus 5.5 cost more per task than Astra despite cheaper tokens?
At max effort, Opus 5.5 writes about 119,000 output tokens per task versus roughly 27,000 for Astra. That volume difference means $5.98 per task despite less than half of Astra’s per-token price.
Why Cost Per Task Decides the Winner
The comparison reframes how buyers evaluate frontier models: raw intelligence scores matter less than quality per dollar on a specific workload. Opus 5.5’s per-token price is less than half of Astra’s, yet at max effort it costs nearly twice as much per task because it writes far more — roughly 119,000 output tokens per task versus about 27,000 for Astra, according to the measurement.
The results also differ from launch claims. Anthropic positioned Opus 5.5 as matching Fable 5.1 at 40% less cost than Opus 5; the independent index places it ahead at 58 versus 53. Fable 5.1, which set a record when it launched, is now outscored by a cheaper sibling in the same lineup. For budget-conscious deployments, Sol and Luna sit on the cost frontier below a score of 50, covering use cases that until recently required more expensive models.
As an affiliate, we earn on qualifying purchases.
The Index Re-Basing Issue
Artificial Analysis rebuilt its Intelligence Index in the same month these models shipped, and the change matters for anyone comparing headlines. When Fable 5.1 launched, it scored 66 on the older v4.1.1 index — the highest recorded at the time. On the current v4.3 index, the same model scores 53, tied with GPT-6 Astra. According to the analysis, the model did not get worse; the test got harder.
Comparisons that mix launch-week numbers from different weeks are effectively using different rulers. All figures cited in this comparison come from the current v4.3 index. Two further caveats: Anthropic models run with fallback behavior, in which safety-flagged requests route to older Claude models and can depress scores on security and biology work; and max effort is not the default — most real deployments run at medium or high effort, where both rankings and costs differ from the max-effort table.
“Opus 5.5 performs at the level of Fable 5.1 on most work at 40% less cost than Opus 5.”
— Anthropic (as cited in the comparison)
As an affiliate, we earn on qualifying purchases.
Caveats in the Comparison
Several limitations remain. The context window for Fable 5.1 is not listed in the sources used. Opus 5.5’s speed was measured only at medium and extra-high effort, while Fable 5.1 is shown at extra-high and max, and Sol and Luna at low and max — so speed figures are not directly comparable across all five models. GPT-6 Sol reportedly regressed on polished, complete deliverables (the GDPval-AA evaluation), and GPT-6 Luna’s Coding Agent Index slipped 2 points versus its predecessor, according to the analysis. Whether these regressions affect a given buyer’s workload is not established by index scores alone; the analysis recommends running a shadow test on your own tasks before switching models.
As an affiliate, we earn on qualifying purchases.
What Buyers Should Do Now
For teams choosing among the five, the analysis recommends workload testing rather than benchmark shopping: run medium-effort trials of Opus 5.5 first, test Astra where token efficiency matters, and evaluate Sol and Luna on high-volume pipelines. Further movement on the index is possible, since Artificial Analysis has re-based its methodology before, which could reshuffle scores again. Additional data points still to be clarified include Fable 5.1’s context window and consistent speed measurements across all effort levels. Future releases from both labs — and any successor to Opus 5.5 — will be measured against this new, higher-scoring baseline.
As an affiliate, we earn on qualifying purchases.
Key Questions
Which of the five September 2026 models scored highest?
Claude Opus 5.5 scored 58 on the Artificial Analysis Intelligence Index v4.3 at max effort, the highest score the index has measured, several points ahead of Claude Fable 5.1 and GPT-6 Astra, both at 53.
Why did Claude Fable 5.1’s score drop from 66 to 53?
The index itself changed. Fable 5.1 scored 66 on the older v4.1.1 version and 53 on the current v4.3 version. According to the analysis, the model did not get worse — the test got harder, so scores from different index versions cannot be compared.
What does $100 buy on each model at max effort?
Roughly 17 index tasks on Claude Opus 5.5, 13 on Claude Fable 5.1, 31 on GPT-6 Astra, 94 on GPT-6 Sol and about 1,429 on GPT-6 Luna.
Which model is best for budget deployments?
GPT-6 Sol ($1.06 per task, index score 48) suits scaled coding and automation, while GPT-6 Luna ($0.07 per task, score 37) suits high-volume classification, extraction and routing. Buyers should note Sol regressed on polished deliverables and Luna’s coding score slipped versus its predecessor.
Why does Opus 5.5 cost more per task than GPT-6 Astra despite cheaper tokens?
At max effort, Opus 5.5 writes far more output — about 119,000 tokens per task versus roughly 27,000 for Astra. That volume difference means Opus 5.5 costs about $5.98 per task despite having less than half of Astra’s per-token price.
Source: Thorsten Meyer AI
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
