AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

September 2026 brought five frontier models from Anthropic and OpenAI in three weeks. On the current Artificial Analysis Intelligence Index (v4.3), Claude Opus 5.5 posts the top score at 58, while GPT-6 Luna handles tasks for $0.07. Cost per task, not headline benchmarks, is the deciding factor for buyers.

Five frontier AI models released within three weeks in September 2026Claude Opus 5.5, Claude Fable 5.1, GPT-6 Astra, GPT-6 Sol and GPT-6 Luna — have been compared on a single independent yardstick. Anthropic’s Opus 5.5, released September 22, posts the highest score measured to date by the Artificial Analysis Intelligence Index at 58, while OpenAI’s GPT-6 Luna completes index tasks for roughly $0.07 each, placing the choice between them on workload and budget rather than raw intelligence.

Anthropic and OpenAI each shipped a full model lineup in September. Anthropic released Claude Fable 5.1 on September 1 and Claude Opus 5.5 on September 22. OpenAI released GPT-6 Astra on September 3, then GPT-6 Sol and GPT-6 Luna on the same day as Opus 5.5, September 22. All five were measured on the current version of the Artificial Analysis Intelligence Index (v4.3).

On pricing and performance, the models span an enormous range. Opus 5.5 charges $4 input / $20 output per million tokens and scores 58 at max effort, at a cost of $5.98 per index task. Fable 5.1, priced at $10/$50, scores 53 at $7.63 per task — outscored by its cheaper sibling at higher cost. GPT-6 Astra, also $10/$50, matches Fable’s 53 score but at $3.26 per task, less than half Fable’s cost. GPT-6 Sol at $2/$10 scores 48 for $1.06 per task, and GPT-6 Luna at $0.10/$0.50 scores 37 for about seven cents per task.

Expressed as throughput, $100 at max effort buys roughly 17 tasks on Opus 5.5, 13 on Fable 5.1, 31 on Astra, 94 on Sol and about 1,429 on Luna. The comparison indicates a practical split: Opus 5.5 for client-ready knowledge work and long agentic coding, Astra for computer use and token-efficient agents, Sol for scaled coding and business automation, and Luna for high-volume classification, extraction and routing.

At a glance
reportWhen: all five models released September 2026…
The developmentAn independent side-by-side comparison of the five frontier AI models released in September 2026 — on a single re-based index and a cost-per-task basis — shows Opus 5.5 leading on score and OpenAI’s Sol and Luna dominating the budget end.
Five Frontier Models, One Bill

Frontier AI · September 2026

Five Frontier Models, One Bill

Three weeks in September 2026 delivered five frontier releases from Anthropic and OpenAI — Opus 5.5, Fable 5.1, GPT-6 Astra, Sol and Luna. Measured on a single independent yardstick, the Artificial Analysis Intelligence Index v4.3, one model leads on score, another costs $0.07 a task. The deciding factor for buyers is cost per task, not headline benchmarks.

All scores: Index v4.3 · Max effort
58
Opus 5.5 — highest index score to date
$0.07/task
GPT-6 Luna — cheapest per index task
3 weeks
Five releases from two labs
Sep 1
Claude Fable 5.1 ships
Sep 3
GPT-6 Astra ships
Sep 22
Opus 5.5 · Sol · Luna
5
Models on one index
v4.3
Re-based methodology

The Scoreboard

Price, Score and Cost per Task

The five models span an enormous range: from $10/$50 per million tokens down to $0.10/$0.50 — and per-task costs from $7.63 to seven cents. Fable 5.1 is outscored by its cheaper sibling at higher cost; Astra matches Fable’s score at less than half the price per task.

ModelLabInput / Output ($/M tok)Index ScoreCost per TaskTasks per $100
Claude Opus 5.5Released Sep 22 Anthropic$4 / $20 58 $5.98~17
Claude Fable 5.1Released Sep 1 Anthropic$10 / $50 53 $7.63~13
GPT-6 AstraReleased Sep 3 OpenAI$10 / $50 53 $3.26~31
GPT-6 SolReleased Sep 22 OpenAI$2 / $10 48 $1.06~94
GPT-6 LunaReleased Sep 22 OpenAI$0.10 / $0.50 37 $0.07~1,429

Throughput

What $100 Buys at Max Effort

Expressed as tasks completed, the same $100 stretches from 17 tasks on Opus 5.5 to roughly 1,429 on GPT-6 Luna — an 84× spread across a single release cycle.

GPT-6 Luna
1,429
GPT-6 Sol
94
GPT-6 Astra
31
Opus 5.5
17
Fable 5.1
13

The Economics

Why Cost per Task Decides the Winner

Raw intelligence scores matter less than quality per dollar on a specific workload. Opus 5.5’s per-token price is less than half of Astra’s — yet at max effort it costs nearly twice as much per task, because it writes far more: roughly 119,000 output tokens per task versus about 27,000 for Astra.

Opus 5.5 · Max Effort

~119,000 output tokens / task

Long, thorough answers drive the per-task bill to $5.98 despite cheap tokens. Best for client-ready knowledge work and long agentic coding where quality justifies volume.

GPT-6 Astra · Max Effort

~27,000 output tokens / task

Concise output keeps per-task cost at $3.26 on identical per-token pricing to Fable 5.1. Best suited for computer use and token-efficient agents.

“Opus 5.5 performs at the level of Fable 5.1 on most work at 40% less cost than Opus 5.”

— Anthropic, as cited in the comparison

The results also differ from launch claims: the independent index places Opus 5.5 ahead of Fable 5.1 at 58 versus 53 — while Fable, which set a record at launch, is now outscored by a cheaper sibling in the same lineup.

Measurement Caveat

The Index Re-Basing Issue

Artificial Analysis rebuilt its Intelligence Index in the same month these models shipped. When Fable 5.1 launched, it scored 66 on the older v4.1.1 index — the highest recorded at the time. On the current v4.3 index, the same model scores 53, tied with GPT-6 Astra. The model did not get worse; the test got harder. Comparisons mixing launch-week numbers are effectively using different rulers.

Fable 5.1 · v4.1.1 → 66
Fable 5.1 · v4.3 → 53
Opus 5.5 · v4.3 → 58
Older index (v4.1.1)Current index (v4.3) — all figures in this report
Caveat · Fallback

Anthropic fallback behavior

Safety-flagged requests route to older Claude models, which can depress scores on security and biology work.

Caveat · Effort Levels

Max effort isn’t the default

Most real deployments run at medium or high effort, where both rankings and costs differ from the max-effort table.

Practical Split

Which Model for Which Job

Opus 5.5

Knowledge work

Client-ready deliverables and long agentic coding.

GPT-6 Astra

Agents

Computer use and token-efficient agent pipelines.

GPT-6 Sol

Scaled coding

High-volume coding and business automation.

GPT-6 Luna

High volume

Classification, extraction and routing at seven cents a task.

Action Plan

What Buyers Should Do Now

The analysis recommends workload testing rather than benchmark shopping — run a shadow test on your own tasks before switching models.

1

Medium-effort trials

Start with Opus 5.5 at medium effort — that’s where most real deployments run.

2

Test Astra

Evaluate GPT-6 Astra where token efficiency and output length matter.

3

Stress Sol & Luna

Evaluate the budget pair on high-volume pipelines and routable work.

4

Shadow test

Index scores alone can’t confirm regressions on your workload — test on your own tasks.

“Cost per task, not headline benchmarks, is the deciding factor for buyers.”

— The comparison’s core finding

Key Questions

Frequently Asked

Which of the five scored highest?

Claude Opus 5.5 at 58 on Index v4.3 at max effort — the highest score the index has measured, ahead of Fable 5.1 and GPT-6 Astra, both at 53.

Why did Fable 5.1’s score drop from 66 to 53?

The index changed. Fable scored 66 on v4.1.1 and 53 on v4.3. The model did not get worse — the test got harder, so scores from different versions cannot be compared.

What does $100 buy on each model?

Roughly 17 tasks on Opus 5.5, 13 on Fable 5.1, 31 on Astra, 94 on Sol and about 1,429 on Luna — all at max effort.

Which model is best for budget deployments?

Sol ($1.06/task, score 48) suits scaled coding and automation; Luna ($0.07/task, score 37) suits classification, extraction and routing. Note Sol regressed on polished deliverables and Luna’s coding score slipped versus its predecessor.

Why does Opus 5.5 cost more per task than Astra despite cheaper tokens?

At max effort, Opus 5.5 writes about 119,000 output tokens per task versus roughly 27,000 for Astra. That volume difference means $5.98 per task despite less than half of Astra’s per-token price.

Why Cost Per Task Decides the Winner

The comparison reframes how buyers evaluate frontier models: raw intelligence scores matter less than quality per dollar on a specific workload. Opus 5.5’s per-token price is less than half of Astra’s, yet at max effort it costs nearly twice as much per task because it writes far more — roughly 119,000 output tokens per task versus about 27,000 for Astra, according to the measurement.

The results also differ from launch claims. Anthropic positioned Opus 5.5 as matching Fable 5.1 at 40% less cost than Opus 5; the independent index places it ahead at 58 versus 53. Fable 5.1, which set a record when it launched, is now outscored by a cheaper sibling in the same lineup. For budget-conscious deployments, Sol and Luna sit on the cost frontier below a score of 50, covering use cases that until recently required more expensive models.

Amazon

AI language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Index Re-Basing Issue

Artificial Analysis rebuilt its Intelligence Index in the same month these models shipped, and the change matters for anyone comparing headlines. When Fable 5.1 launched, it scored 66 on the older v4.1.1 index — the highest recorded at the time. On the current v4.3 index, the same model scores 53, tied with GPT-6 Astra. According to the analysis, the model did not get worse; the test got harder.

Comparisons that mix launch-week numbers from different weeks are effectively using different rulers. All figures cited in this comparison come from the current v4.3 index. Two further caveats: Anthropic models run with fallback behavior, in which safety-flagged requests route to older Claude models and can depress scores on security and biology work; and max effort is not the default — most real deployments run at medium or high effort, where both rankings and costs differ from the max-effort table.

“Opus 5.5 performs at the level of Fable 5.1 on most work at 40% less cost than Opus 5.”

— Anthropic (as cited in the comparison)

Amazon

AI model performance comparison

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Caveats in the Comparison

Several limitations remain. The context window for Fable 5.1 is not listed in the sources used. Opus 5.5’s speed was measured only at medium and extra-high effort, while Fable 5.1 is shown at extra-high and max, and Sol and Luna at low and max — so speed figures are not directly comparable across all five models. GPT-6 Sol reportedly regressed on polished, complete deliverables (the GDPval-AA evaluation), and GPT-6 Luna’s Coding Agent Index slipped 2 points versus its predecessor, according to the analysis. Whether these regressions affect a given buyer’s workload is not established by index scores alone; the analysis recommends running a shadow test on your own tasks before switching models.

Amazon

cost-effective AI chatbot

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Buyers Should Do Now

For teams choosing among the five, the analysis recommends workload testing rather than benchmark shopping: run medium-effort trials of Opus 5.5 first, test Astra where token efficiency matters, and evaluate Sol and Luna on high-volume pipelines. Further movement on the index is possible, since Artificial Analysis has re-based its methodology before, which could reshuffle scores again. Additional data points still to be clarified include Fable 5.1’s context window and consistent speed measurements across all effort levels. Future releases from both labs — and any successor to Opus 5.5 — will be measured against this new, higher-scoring baseline.

Amazon

enterprise AI assistant

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Which of the five September 2026 models scored highest?

Claude Opus 5.5 scored 58 on the Artificial Analysis Intelligence Index v4.3 at max effort, the highest score the index has measured, several points ahead of Claude Fable 5.1 and GPT-6 Astra, both at 53.

Why did Claude Fable 5.1’s score drop from 66 to 53?

The index itself changed. Fable 5.1 scored 66 on the older v4.1.1 version and 53 on the current v4.3 version. According to the analysis, the model did not get worse — the test got harder, so scores from different index versions cannot be compared.

What does $100 buy on each model at max effort?

Roughly 17 index tasks on Claude Opus 5.5, 13 on Claude Fable 5.1, 31 on GPT-6 Astra, 94 on GPT-6 Sol and about 1,429 on GPT-6 Luna.

Which model is best for budget deployments?

GPT-6 Sol ($1.06 per task, index score 48) suits scaled coding and automation, while GPT-6 Luna ($0.07 per task, score 37) suits high-volume classification, extraction and routing. Buyers should note Sol regressed on polished deliverables and Luna’s coding score slipped versus its predecessor.

Why does Opus 5.5 cost more per task than GPT-6 Astra despite cheaper tokens?

At max effort, Opus 5.5 writes far more output — about 119,000 tokens per task versus roughly 27,000 for Astra. That volume difference means Opus 5.5 costs about $5.98 per task despite having less than half of Astra’s per-token price.

Source: Thorsten Meyer AI

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Jakarta Dinaungi Langit Cerah Hari Ini – Beritajakarta.id

Jakarta experiences clear skies today, providing favorable weather conditions for residents and outdoor activities, according to Beritajakarta.id.

Will The **High Temp In Philadelphia** Be 81-82° On Jul 28, 2026?

Market activity suggests a prediction that Philadelphia’s high temperature on July 28, 2026, will be between 81-82°F, but official forecasts are not yet available.

Lotto 6Aus49 Jackpot

The Lotto 6aus49 jackpot has grown to a record high, attracting widespread attention. Here’s what is confirmed and what remains uncertain.

AI Breakthrough: What The Weights First Approach Reveals About Machine Thinking

Thinking Machines Lab released Inkling’s weights under Apache 2.0, prioritizing model ownership over benchmark leadership.