Real cost per task across frontier models

Updated September 22, 2026 · first published September 22, 2026

List prices tell you what a token costs. They do not tell you what a finished task costs, because models spend very different numbers of tokens to finish the same work. This page puts measured cost per task and a quality score side by side for Claude Opus 5.5, Claude Opus 5, Claude Fable 5.1, GPT-6 Astra and GPT-6 Sol at every effort level. Every figure has a source, and nothing is estimated.

The short answer: GPT-6 Sol is the cheapest way to reach any score up to 47.5. Above 47.5, Opus 5.5 is the cheapest at every level. No setting of GPT-6 Astra above low, Opus 5 or Fable 5.1 is cost-efficient: each one is beaten on both score and price by some setting of Sol or Opus 5.5.

Method

All scores and costs come from the Artificial Analysis Intelligence Index v4.3.2, read on 22 September 2026. The index runs ten evaluations: AA-Briefcase, GDPval-AA, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR. Artificial Analysis runs each model at each effort setting through the provider API and records what it pays at list prices. Cost per task is that spend per task, split into input and output. The same suite runs for every row, so the rows compare directly.

Two limits apply before you read the numbers. First, the index is one task mix. Your mix will shift the costs, which is why the last section explains how to measure your own. Second, scores from different index versions do not compare, so do not mix these figures with older published numbers.

The full ladder: 26 settings

Output tokens per task are the tokens the model writes, including reasoning. Artificial Analysis does not publish the input and output split for the four middle GPT-6 Sol settings, so those cells say n/a.

ModelEffortIntelligence IndexCost per taskInput costOutput costOutput tokens per taskOn the frontier
GPT-6 Solnone28.1$0.33$0.28$0.054,900No
GPT-6 Sollow33.9$0.13n/an/a3,400Yes
GPT-6 Solmedium (default)39.8$0.25n/an/a6,500Yes
GPT-6 Solhigh42.8$0.37n/an/a10,200Yes
GPT-6 Solxhigh44.1$0.53n/an/a16,000Yes
GPT-6 Solmax47.5$1.06$0.74$0.3131,200Yes
GPT-6 Astralow45.8$0.82$0.60$0.224,400Yes
GPT-6 Astramedium49.6$1.54$1.06$0.489,600No
GPT-6 Astrahigh50.9$1.73$1.13$0.5911,800No
GPT-6 Astraxhigh52.4$2.31$1.46$0.8516,900No
GPT-6 Astramax52.7$3.26$1.90$1.3627,200No
Claude Opus 5low39.4$1.10$0.73$0.3714,700No
Claude Opus 5medium44.8$2.19$1.47$0.7229,000No
Claude Opus 5high (default)48.1$3.61$2.46$1.1646,200No
Claude Opus 5xhigh49.7$4.88$3.36$1.5260,700No
Claude Opus 5max50.8$5.86$4.05$1.8172,500No
Claude Opus 5.5low42.3$0.55$0.35$0.2010,200No
Claude Opus 5.5medium (default)51.2$1.34$0.82$0.5125,700Yes
Claude Opus 5.5high53.6$1.82$1.11$0.7135,600Yes
Claude Opus 5.5xhigh56.0$3.46$2.15$1.3165,700Yes
Claude Opus 5.5max57.6$5.98$3.60$2.38119,200Yes
Claude Fable 5.1low46.8$2.37$1.30$1.0821,600No
Claude Fable 5.1medium48.9$2.98$1.59$1.3927,900No
Claude Fable 5.1high51.2$3.91$2.01$1.9038,100No
Claude Fable 5.1xhigh53.2$5.98$2.96$3.0260,500No
Claude Fable 5.1max53.4$7.63$3.73$3.9078,100No

On the frontier means that no other setting in this table scores higher for less money.

The cost frontier

Ten of the 26 settings are on the frontier. Sorted by cost, they are the only settings worth considering on price alone. Everything else pays more for the same score or less.

SettingIntelligence IndexCost per taskPer 10,000 tasks
GPT-6 Sol · low33.9$0.13$1,300
GPT-6 Sol · medium39.8$0.25$2,500
GPT-6 Sol · high42.8$0.37$3,700
GPT-6 Sol · xhigh44.1$0.53$5,300
GPT-6 Astra · low45.8$0.82$8,200
GPT-6 Sol · max47.5$1.06$10,600
Claude Opus 5.5 · medium51.2$1.34$13,400
Claude Opus 5.5 · high53.6$1.82$18,200
Claude Opus 5.5 · xhigh56.0$3.46$34,600
Claude Opus 5.5 · max57.6$5.98$59,800

The frontier has two regions. From $0.13 to $1.06 it is almost all GPT-6 Sol, with GPT-6 Astra low as the one exception at $0.82. From $1.34 up it is Opus 5.5 only. The step between the two regions is small: Opus 5.5 medium scores 3.7 points more than Sol max for $0.28 more per task.

Head to head at equal or better score

These pairs compare a cheaper setting with a dearer setting that scores the same or lower.

Cheaper settingDearer settingScore gapCost saved
Opus 5.5 medium: 51.2, $1.34Opus 5 high: 48.1, $3.61+3.1 for Opus 5.563%
Opus 5.5 medium: 51.2, $1.34Opus 5 max: 50.8, $5.86+0.4 for Opus 5.577%
Opus 5.5 medium: 51.2, $1.34Fable 5.1 high: 51.2, $3.91tie66%
Opus 5.5 high: 53.6, $1.82GPT-6 Astra max: 52.7, $3.26+0.9 for Opus 5.544%
Opus 5.5 high: 53.6, $1.82Fable 5.1 max: 53.4, $7.63+0.2 for Opus 5.576%
GPT-6 Sol max: 47.5, $1.06Opus 5 high: 48.1, $3.61+0.6 for Opus 571%

The Opus 5 row matters for teams that have not migrated yet. Opus 5 ran at high effort by default. Opus 5.5 runs at medium by default, and at that default it scores 3.1 points more for 63% less per task. Anthropic says Opus 5.5 is 40% cheaper than Opus 5; the measured gap on this index is wider than the claim.

Where the money goes: input, not output

Across every model, input is between about half and three quarters of the cost per task. For Opus 5.5 at medium, $0.82 of the $1.34 is input, which is 61%. For GPT-6 Astra at max it is 58%, for GPT-6 Sol at max 70%, and for Fable 5.1 at max 49%. Agent tasks re-send their growing context on every step, so input spend grows with the number of steps, not just with the length of the answer.

This is why GPT-6 Sol at no reasoning costs more than at low: $0.33 against $0.13, even though it writes only 4,900 output tokens. $0.28 of the $0.33 is input. The likely cause is more steps, each one re-sending context. Turning reasoning off does not make an agent cheaper when the agent then needs more turns.

For your bill, this means caching matters as much as the model choice. Opus 5.5 and GPT-6 Sol both charge $0.20 per million cached input tokens. GPT-6 Astra charges $1.00 and Opus 5 charges $0.50.

Token volume against list price

GPT-6 Astra is the most token-efficient model here. At max effort it writes 27,200 output tokens per task, against 119,200 for Opus 5.5 at max. But Astra charges $10 and $50 per million, 2.5× the Opus 5.5 rate. The list price wins: Opus 5.5 high costs $1.82 per task and scores 53.6, while Astra max costs $3.26 and scores 52.7.

Opus 5.5 at max is the one place where verbosity bites. Its 119,200 output tokens cost $2.38 before any input, and the full task costs $5.98. That buys the highest score in the table, 57.6, but it is 4.5× the cost of medium.

Diminishing returns by effort

Each row shows what one step up in effort buys, as index points against extra cost per task.

Modelmedium → highxhigh → max
GPT-6 Sol+3.0 points for +48%+3.4 points for +100%
GPT-6 Astra+1.3 points for +12%+0.3 points for +41%
Claude Opus 5+3.3 points for +65%+1.1 points for +20%
Claude Opus 5.5+2.4 points for +36%+1.6 points for +73%
Claude Fable 5.1+2.3 points for +31%+0.2 points for +28%

For Fable 5.1 and GPT-6 Astra, max effort buys almost nothing: 0.2 and 0.3 points for 28% and 41% more. For Opus 5.5 the step to max still buys 1.6 points, but at 73% more cost. Use max only on tasks where you have measured that the extra points change the outcome.

When GPT-6 Sol max is enough

The index is an average, and the gap between Sol max and Opus 5.5 medium is not even across it. Per-evaluation scores from Artificial Analysis:

EvaluationOpus 5.5 medium ($1.34)GPT-6 Sol max ($1.06)
Terminal-Bench 4.052.543.9
AutomationBench-AA61.261.6
GDPval-AA (Elo)15761487
AA-Briefcase (Elo)16421483
Humanity's Last Exam54.747.9
SciCode59.357.6
CritPt27.730.9
AA-LCR84.383.7

On AutomationBench and CritPt, GPT-6 Sol matches or beats Opus 5.5 at medium for $1.06 per task against $1.34. On Terminal-Bench, the knowledge-work evaluations and Humanity's Last Exam, Opus 5.5 leads by a wide margin. Workflow automation can run on Sol. Terminal and coding agents and document-heavy knowledge work are where Opus 5.5 earns its price.

Vendor claims against independent numbers

Vendor figures and independent figures often differ, because they use different harnesses and settings. Anthropic reports 66.4% on Terminal-Bench 4.0 for Opus 5.5 at xhigh. Artificial Analysis measures 59.6% at the same effort.

EffortOpus 5.5Fable 5.1
low31.3%40.4%
medium52.5%44.9%
high56.6%52.0%
xhigh59.6%55.1%
max59.6%52.0%

The ranking can also change with the index. On the Vals AI Index, Fable 5.1 is first at 68.83%, GPT-6 Astra third at 66.61% and Opus 5.5 fourth at 66.16%. Cost per task is not published there on the same basis, so it does not move the frontier above. It is a reminder that one index is one view.

List prices and a cache-heavy task

Digital Applied priced one cache-heavy agent task on each model: 8M cache-read tokens, 400K uncached input, 600K cache-write tokens and 300K output tokens. The result follows the frontier, with one extra penalty for GPT-6 Astra.

ModelInput $/MTokOutput $/MTokCache read $/MTokCache-heavy task (Digital Applied)
GPT-6 Sol$2.00$10.00$0.20$6.90
Claude Opus 5.5$4.00$20.00$0.20$12.20
Claude Opus 5$5.00$25.00$0.50$17.25
Claude Fable 5.1$10.00$50.00$0.25$28.50
GPT-6 Astra (≤272K)$10.00$50.00$1.00$34.50
GPT-6 Astra (>272K)$20.00$75.00n/a$61.50

GPT-6 Astra bills the whole request at $20 and $75 per million once input passes 272K tokens. Opus 5.5 keeps one price across its 1M-token window. Opus 5.5 cache writes cost $5 per million for the 5-minute cache and $8 for the 1-hour cache, batch jobs are billed at 50%, and fast mode doubles the price. For the details, see the Claude Opus 5.5 cost model and Opus 5.5 fast mode economics.

How to measure your own cost per task

  1. Log tokens per task, not per request: uncached input, cached input, cache writes and output, summed over every step of the task.
  2. Take a sample of a few hundred real tasks and replay it on two or three candidate settings, for example Sol max, Opus 5.5 medium and Opus 5.5 high.
  3. Grade each result with the same check you use in production, and count the tasks that pass.
  4. Divide total spend by passed tasks. Cost per successful task is the number to compare, because a cheap setting that fails more often is not cheap.
  5. Pick the cheapest setting that clears your quality bar, then route only the tasks that fail to a higher setting.
  6. Re-run the replay when a price or a default effort changes.

For how effort settings change the bill on one model, see Opus 5.5 effort levels are the real price.

Caveats

Sources

Related


Want this applied to your own LLM spend? FinOps LLM runs a free audit of your AI costs and shows where the savings are. Book free audit →

Back to research