Claude Opus 5.5 cost model

Updated September 22, 2026 · first published September 22, 2026

Anthropic released Claude Opus 5.5 (claude-opus-5-5) on September 22, 2026. The headline price is $4 per million input tokens and $20 per million output tokens, down from $5 and $25 on Opus 5. That is a 20% cut. For most production workloads the actual cut is larger, because the rate that moved most is the one the headline leaves out.

The rates that matter

Token typeOpus 5.5Opus 5Fable 5.1
Input (uncached)$4.00 / MTok$5.00 / MTok$10.00 / MTok
Output$20.00 / MTok$25.00 / MTok$50.00 / MTok
Cache write (5-minute)$5.00 / MTok$6.25 / MTok$12.50 / MTok
Cache read$0.20 / MTok$0.50 / MTok$0.25 / MTok
Fast mode (input / output)$8 / $40

Every row is 20% lower except the cache read, which is 60% lower. At $0.20 it is now cheaper than Fable 5.1's $0.25. That reverses the one row where Fable used to win. In our Fable 5.1 cost model, a heavily cached workload could come out cheaper on Fable than on Opus 5. Against Opus 5.5 that no longer happens. Fable costs more on every token type.

Why a typical task costs about 40% less

Anthropic's claim is a 40% drop in cost for a typical job, not 20%. The extra comes from token efficiency. Anthropic says Opus 5.5 reaches the same result with about 20% fewer input and output tokens and writes about 40% less verbose output. A lower rate times fewer tokens compounds: 0.8 × 0.8 = 0.64, which is roughly 36–40% off depending on the mix.

Here is one agent step: 200,000 input tokens, 90% served from cache, and 20,000 output tokens.

LineOpus 5Opus 5.5 (same tokens)Fable 5.1
20K uncached input$0.100$0.080$0.200
180K cached input$0.090$0.036$0.045
20K output$0.500$0.400$1.000
Total$0.690$0.516$1.245

On identical tokens, Opus 5.5 is 25% cheaper than Opus 5, not 20%, because of the cache discount. If the 20% token reduction holds for your workload, the step drops to about $0.41, which is 40% below Opus 5 and a third of the Fable cost. Output is still about 80% of the bill in this example. That is where the verbosity reduction pays off.

Cost per task against GPT-6 Sol, GPT-6 Astra and Fable 5.1

List prices per million tokens, and the same agent step as above (200K input, 90% cached, 20K output) priced on each model.

ModelInput $/MTokOutput $/MTokCache read $/MTokSame agent step
GPT-6 Sol$2.00$10.00$0.20$0.276
Claude Opus 5.5$4.00$20.00$0.20$0.516
Claude Opus 5$5.00$25.00$0.50$0.690
Claude Fable 5.1$10.00$50.00$0.25$1.245
GPT-6 Astra$10.00$50.00$1.00$1.380

Measured cost per task on the Artificial Analysis Intelligence Index v4.3.2, read 22 September 2026, sorted by score. GPT-6 Sol is the cheapest way to any score up to 47.5. Above that, Opus 5.5 is the cheapest at every level: Opus 5.5 high scores 53.6 for $1.82, while GPT-6 Astra max scores 52.7 for $3.26 and Fable 5.1 max scores 53.4 for $7.63.

ModelIntelligence IndexCost per taskOutput tokens per task
GPT-6 Sol · medium39.8$0.256,500
GPT-6 Sol · xhigh44.1$0.5316,000
GPT-6 Astra · low45.8$0.824,400
GPT-6 Sol · max47.5$1.0631,200
Claude Opus 5 · high48.1$3.6146,200
Claude Opus 5 · max50.8$5.8672,500
Claude Opus 5.5 · medium51.2$1.3425,700
Claude Fable 5.1 · high51.2$3.9138,100
GPT-6 Astra · max52.7$3.2627,200
Claude Fable 5.1 · max53.4$7.6378,100
Claude Opus 5.5 · high53.6$1.8235,600
Claude Opus 5.5 · max57.6$5.98119,200

Every GPT-6 Astra setting above low, every Opus 5 setting and every Fable 5.1 setting costs more than a Sol or Opus 5.5 setting that scores the same or higher. Cost per task depends on the benchmark's task mix. Replay your own traffic before switching models.

Is it good enough to replace Fable?

Anthropic positions Opus 5.5 as matching Fable 5.1 on most tasks. The published numbers support that for coding and agentic work:

Artificial Analysis ranked it first on its Intelligence Index on launch day. Anthropic also says that at default (medium) effort it beats GPT-6 Astra at max effort on knowledge work for about a fifth of the cost per task. That is a vendor comparison, so treat it as a hypothesis to test. The cost math above is the part you can verify yourself.

Fast mode costs 2×

Fast mode costs $8 in and $40 out, exactly double the standard rate, for up to 2.5× the speed. The standard model is also about 30% faster than Opus 5. Use fast mode for latency-bound interactive paths where a user is waiting. Keep it off batch jobs, background agents and evals. When fast mode is enabled by default, the unit cost doubles and the dashboard shows nothing unusual until the invoice arrives.

What to do this week

  1. Pin the model ID. Point production at claude-opus-5-5 explicitly and record the ID on every request, so cost per request can be split by model after the migration.
  2. Re-baseline on your own traffic. Replay a sample of real requests on both models and compare tokens per completed task, not per call. The 20% token-efficiency claim is the part most likely to vary by workload.
  3. Re-check effort settings. If a workload moved to Fable or to max effort for quality, test Opus 5.5 at medium first. That is the setting Anthropic benchmarks against.
  4. Re-run the Fable routing decision. Any route that sent cache-heavy traffic to Fable to get the cheaper cache read should be re-evaluated. That cost advantage is gone.
  5. Watch cache hit ratio. At $0.20 against $4.00, a cache miss now costs 20× a hit. A prompt-template change that breaks the prefix costs more than it did on Opus 5.

Anthropic says Sonnet 5.5 and Haiku 5.5 will follow in the coming weeks. Keep the replay harness from step 2 so each release can be re-baselined in an afternoon.

Related


Want this applied to your own LLM spend? FinOps LLM runs a free audit of your AI costs and shows where the savings are. Book free audit →

Back to research