Real cost per task across frontier models
Updated September 22, 2026 · first published September 22, 2026
List prices tell you what a token costs. They do not tell you what a finished task costs, because models spend very different numbers of tokens to finish the same work. This page puts measured cost per task and a quality score side by side for Claude Opus 5.5, Claude Opus 5, Claude Fable 5.1, GPT-6 Astra and GPT-6 Sol at every effort level. Every figure has a source, and nothing is estimated.
The short answer: GPT-6 Sol is the cheapest way to reach any score up to 47.5. Above 47.5, Opus 5.5 is the cheapest at every level. No setting of GPT-6 Astra above low, Opus 5 or Fable 5.1 is cost-efficient: each one is beaten on both score and price by some setting of Sol or Opus 5.5.
Method
All scores and costs come from the Artificial Analysis Intelligence Index v4.3.2, read on 22 September 2026. The index runs ten evaluations: AA-Briefcase, GDPval-AA, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR. Artificial Analysis runs each model at each effort setting through the provider API and records what it pays at list prices. Cost per task is that spend per task, split into input and output. The same suite runs for every row, so the rows compare directly.
Two limits apply before you read the numbers. First, the index is one task mix. Your mix will shift the costs, which is why the last section explains how to measure your own. Second, scores from different index versions do not compare, so do not mix these figures with older published numbers.
The full ladder: 26 settings
Output tokens per task are the tokens the model writes, including reasoning. Artificial Analysis does not publish the input and output split for the four middle GPT-6 Sol settings, so those cells say n/a.
| Model | Effort | Intelligence Index | Cost per task | Input cost | Output cost | Output tokens per task | On the frontier |
|---|---|---|---|---|---|---|---|
| GPT-6 Sol | none | 28.1 | $0.33 | $0.28 | $0.05 | 4,900 | No |
| GPT-6 Sol | low | 33.9 | $0.13 | n/a | n/a | 3,400 | Yes |
| GPT-6 Sol | medium (default) | 39.8 | $0.25 | n/a | n/a | 6,500 | Yes |
| GPT-6 Sol | high | 42.8 | $0.37 | n/a | n/a | 10,200 | Yes |
| GPT-6 Sol | xhigh | 44.1 | $0.53 | n/a | n/a | 16,000 | Yes |
| GPT-6 Sol | max | 47.5 | $1.06 | $0.74 | $0.31 | 31,200 | Yes |
| GPT-6 Astra | low | 45.8 | $0.82 | $0.60 | $0.22 | 4,400 | Yes |
| GPT-6 Astra | medium | 49.6 | $1.54 | $1.06 | $0.48 | 9,600 | No |
| GPT-6 Astra | high | 50.9 | $1.73 | $1.13 | $0.59 | 11,800 | No |
| GPT-6 Astra | xhigh | 52.4 | $2.31 | $1.46 | $0.85 | 16,900 | No |
| GPT-6 Astra | max | 52.7 | $3.26 | $1.90 | $1.36 | 27,200 | No |
| Claude Opus 5 | low | 39.4 | $1.10 | $0.73 | $0.37 | 14,700 | No |
| Claude Opus 5 | medium | 44.8 | $2.19 | $1.47 | $0.72 | 29,000 | No |
| Claude Opus 5 | high (default) | 48.1 | $3.61 | $2.46 | $1.16 | 46,200 | No |
| Claude Opus 5 | xhigh | 49.7 | $4.88 | $3.36 | $1.52 | 60,700 | No |
| Claude Opus 5 | max | 50.8 | $5.86 | $4.05 | $1.81 | 72,500 | No |
| Claude Opus 5.5 | low | 42.3 | $0.55 | $0.35 | $0.20 | 10,200 | No |
| Claude Opus 5.5 | medium (default) | 51.2 | $1.34 | $0.82 | $0.51 | 25,700 | Yes |
| Claude Opus 5.5 | high | 53.6 | $1.82 | $1.11 | $0.71 | 35,600 | Yes |
| Claude Opus 5.5 | xhigh | 56.0 | $3.46 | $2.15 | $1.31 | 65,700 | Yes |
| Claude Opus 5.5 | max | 57.6 | $5.98 | $3.60 | $2.38 | 119,200 | Yes |
| Claude Fable 5.1 | low | 46.8 | $2.37 | $1.30 | $1.08 | 21,600 | No |
| Claude Fable 5.1 | medium | 48.9 | $2.98 | $1.59 | $1.39 | 27,900 | No |
| Claude Fable 5.1 | high | 51.2 | $3.91 | $2.01 | $1.90 | 38,100 | No |
| Claude Fable 5.1 | xhigh | 53.2 | $5.98 | $2.96 | $3.02 | 60,500 | No |
| Claude Fable 5.1 | max | 53.4 | $7.63 | $3.73 | $3.90 | 78,100 | No |
On the frontier means that no other setting in this table scores higher for less money.
The cost frontier
Ten of the 26 settings are on the frontier. Sorted by cost, they are the only settings worth considering on price alone. Everything else pays more for the same score or less.
| Setting | Intelligence Index | Cost per task | Per 10,000 tasks |
|---|---|---|---|
| GPT-6 Sol · low | 33.9 | $0.13 | $1,300 |
| GPT-6 Sol · medium | 39.8 | $0.25 | $2,500 |
| GPT-6 Sol · high | 42.8 | $0.37 | $3,700 |
| GPT-6 Sol · xhigh | 44.1 | $0.53 | $5,300 |
| GPT-6 Astra · low | 45.8 | $0.82 | $8,200 |
| GPT-6 Sol · max | 47.5 | $1.06 | $10,600 |
| Claude Opus 5.5 · medium | 51.2 | $1.34 | $13,400 |
| Claude Opus 5.5 · high | 53.6 | $1.82 | $18,200 |
| Claude Opus 5.5 · xhigh | 56.0 | $3.46 | $34,600 |
| Claude Opus 5.5 · max | 57.6 | $5.98 | $59,800 |
The frontier has two regions. From $0.13 to $1.06 it is almost all GPT-6 Sol, with GPT-6 Astra low as the one exception at $0.82. From $1.34 up it is Opus 5.5 only. The step between the two regions is small: Opus 5.5 medium scores 3.7 points more than Sol max for $0.28 more per task.
Head to head at equal or better score
These pairs compare a cheaper setting with a dearer setting that scores the same or lower.
| Cheaper setting | Dearer setting | Score gap | Cost saved |
|---|---|---|---|
| Opus 5.5 medium: 51.2, $1.34 | Opus 5 high: 48.1, $3.61 | +3.1 for Opus 5.5 | 63% |
| Opus 5.5 medium: 51.2, $1.34 | Opus 5 max: 50.8, $5.86 | +0.4 for Opus 5.5 | 77% |
| Opus 5.5 medium: 51.2, $1.34 | Fable 5.1 high: 51.2, $3.91 | tie | 66% |
| Opus 5.5 high: 53.6, $1.82 | GPT-6 Astra max: 52.7, $3.26 | +0.9 for Opus 5.5 | 44% |
| Opus 5.5 high: 53.6, $1.82 | Fable 5.1 max: 53.4, $7.63 | +0.2 for Opus 5.5 | 76% |
| GPT-6 Sol max: 47.5, $1.06 | Opus 5 high: 48.1, $3.61 | +0.6 for Opus 5 | 71% |
The Opus 5 row matters for teams that have not migrated yet. Opus 5 ran at high effort by default. Opus 5.5 runs at medium by default, and at that default it scores 3.1 points more for 63% less per task. Anthropic says Opus 5.5 is 40% cheaper than Opus 5; the measured gap on this index is wider than the claim.
Where the money goes: input, not output
Across every model, input is between about half and three quarters of the cost per task. For Opus 5.5 at medium, $0.82 of the $1.34 is input, which is 61%. For GPT-6 Astra at max it is 58%, for GPT-6 Sol at max 70%, and for Fable 5.1 at max 49%. Agent tasks re-send their growing context on every step, so input spend grows with the number of steps, not just with the length of the answer.
This is why GPT-6 Sol at no reasoning costs more than at low: $0.33 against $0.13, even though it writes only 4,900 output tokens. $0.28 of the $0.33 is input. The likely cause is more steps, each one re-sending context. Turning reasoning off does not make an agent cheaper when the agent then needs more turns.
For your bill, this means caching matters as much as the model choice. Opus 5.5 and GPT-6 Sol both charge $0.20 per million cached input tokens. GPT-6 Astra charges $1.00 and Opus 5 charges $0.50.
Token volume against list price
GPT-6 Astra is the most token-efficient model here. At max effort it writes 27,200 output tokens per task, against 119,200 for Opus 5.5 at max. But Astra charges $10 and $50 per million, 2.5× the Opus 5.5 rate. The list price wins: Opus 5.5 high costs $1.82 per task and scores 53.6, while Astra max costs $3.26 and scores 52.7.
Opus 5.5 at max is the one place where verbosity bites. Its 119,200 output tokens cost $2.38 before any input, and the full task costs $5.98. That buys the highest score in the table, 57.6, but it is 4.5× the cost of medium.
Diminishing returns by effort
Each row shows what one step up in effort buys, as index points against extra cost per task.
| Model | medium → high | xhigh → max |
|---|---|---|
| GPT-6 Sol | +3.0 points for +48% | +3.4 points for +100% |
| GPT-6 Astra | +1.3 points for +12% | +0.3 points for +41% |
| Claude Opus 5 | +3.3 points for +65% | +1.1 points for +20% |
| Claude Opus 5.5 | +2.4 points for +36% | +1.6 points for +73% |
| Claude Fable 5.1 | +2.3 points for +31% | +0.2 points for +28% |
For Fable 5.1 and GPT-6 Astra, max effort buys almost nothing: 0.2 and 0.3 points for 28% and 41% more. For Opus 5.5 the step to max still buys 1.6 points, but at 73% more cost. Use max only on tasks where you have measured that the extra points change the outcome.
When GPT-6 Sol max is enough
The index is an average, and the gap between Sol max and Opus 5.5 medium is not even across it. Per-evaluation scores from Artificial Analysis:
| Evaluation | Opus 5.5 medium ($1.34) | GPT-6 Sol max ($1.06) |
|---|---|---|
| Terminal-Bench 4.0 | 52.5 | 43.9 |
| AutomationBench-AA | 61.2 | 61.6 |
| GDPval-AA (Elo) | 1576 | 1487 |
| AA-Briefcase (Elo) | 1642 | 1483 |
| Humanity's Last Exam | 54.7 | 47.9 |
| SciCode | 59.3 | 57.6 |
| CritPt | 27.7 | 30.9 |
| AA-LCR | 84.3 | 83.7 |
On AutomationBench and CritPt, GPT-6 Sol matches or beats Opus 5.5 at medium for $1.06 per task against $1.34. On Terminal-Bench, the knowledge-work evaluations and Humanity's Last Exam, Opus 5.5 leads by a wide margin. Workflow automation can run on Sol. Terminal and coding agents and document-heavy knowledge work are where Opus 5.5 earns its price.
Vendor claims against independent numbers
Vendor figures and independent figures often differ, because they use different harnesses and settings. Anthropic reports 66.4% on Terminal-Bench 4.0 for Opus 5.5 at xhigh. Artificial Analysis measures 59.6% at the same effort.
| Effort | Opus 5.5 | Fable 5.1 |
|---|---|---|
| low | 31.3% | 40.4% |
| medium | 52.5% | 44.9% |
| high | 56.6% | 52.0% |
| xhigh | 59.6% | 55.1% |
| max | 59.6% | 52.0% |
The ranking can also change with the index. On the Vals AI Index, Fable 5.1 is first at 68.83%, GPT-6 Astra third at 66.61% and Opus 5.5 fourth at 66.16%. Cost per task is not published there on the same basis, so it does not move the frontier above. It is a reminder that one index is one view.
List prices and a cache-heavy task
Digital Applied priced one cache-heavy agent task on each model: 8M cache-read tokens, 400K uncached input, 600K cache-write tokens and 300K output tokens. The result follows the frontier, with one extra penalty for GPT-6 Astra.
| Model | Input $/MTok | Output $/MTok | Cache read $/MTok | Cache-heavy task (Digital Applied) |
|---|---|---|---|---|
| GPT-6 Sol | $2.00 | $10.00 | $0.20 | $6.90 |
| Claude Opus 5.5 | $4.00 | $20.00 | $0.20 | $12.20 |
| Claude Opus 5 | $5.00 | $25.00 | $0.50 | $17.25 |
| Claude Fable 5.1 | $10.00 | $50.00 | $0.25 | $28.50 |
| GPT-6 Astra (≤272K) | $10.00 | $50.00 | $1.00 | $34.50 |
| GPT-6 Astra (>272K) | $20.00 | $75.00 | n/a | $61.50 |
GPT-6 Astra bills the whole request at $20 and $75 per million once input passes 272K tokens. Opus 5.5 keeps one price across its 1M-token window. Opus 5.5 cache writes cost $5 per million for the 5-minute cache and $8 for the 1-hour cache, batch jobs are billed at 50%, and fast mode doubles the price. For the details, see the Claude Opus 5.5 cost model and Opus 5.5 fast mode economics.
How to measure your own cost per task
- Log tokens per task, not per request: uncached input, cached input, cache writes and output, summed over every step of the task.
- Take a sample of a few hundred real tasks and replay it on two or three candidate settings, for example Sol max, Opus 5.5 medium and Opus 5.5 high.
- Grade each result with the same check you use in production, and count the tasks that pass.
- Divide total spend by passed tasks. Cost per successful task is the number to compare, because a cheap setting that fails more often is not cheap.
- Pick the cheapest setting that clears your quality bar, then route only the tasks that fail to a higher setting.
- Re-run the replay when a price or a default effort changes.
For how effort settings change the bill on one model, see Opus 5.5 effort levels are the real price.
Caveats
- All costs are at list API prices on the Artificial Analysis task mix. Discounts, batch pricing and your own caching change them.
- Artificial Analysis labels its Opus 5.5 and Fable 5.1 entries Default Fallback. Check its methodology notes before treating every row as the same API configuration.
- Scores are from index v4.3.2 only. Earlier versions used different evaluations and do not compare.
- The Digital Applied cache-heavy task is one illustrative workload, not a measurement.
- Numbers change when providers change prices or defaults. The read date is 22 September 2026.
Sources
- Artificial Analysis Intelligence Index
- Artificial Analysis: Claude Opus 5.5
- Artificial Analysis: Claude Opus 5
- Artificial Analysis: Claude Fable 5.1
- Artificial Analysis: GPT-6 Astra
- Artificial Analysis: GPT-6 Sol
- Digital Applied: GPT-6 Sol vs Claude Opus 5.5 cost benchmarks
- Digital Applied: Claude Opus 5.5 vs GPT-6 Astra
- Digital Applied: Is Claude Fable 5.1 still worth it?
- Digital Applied: Claude Opus 5.5 launch pricing and benchmarks
- Digital Applied: Opus 5.5 vs Grok 4.7 vs Muse Spark 1.3 cost per task
Related
Want this applied to your own LLM spend? FinOps LLM runs a free audit of your AI costs and shows where the savings are. Book free audit →