Claude Sonnet 5.5 pricing: what changes, what does not

Updated September 28, 2026 · first published September 27, 2026

Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens on the Anthropic API, the same list price as Sonnet 5. The economic change is efficiency: Anthropic says Sonnet 5.5 can cost up to 30% less per task because it uses fewer tokens, while generating output more than 30% faster. That is a vendor-reported maximum, not a guaranteed discount on every workload.

For a budget forecast, keep the token rates fixed and model the possible reduction in tokens and agent steps separately. Measure your own task mix before applying any savings assumption.

Sonnet 5.5 API prices

Token typePrice per million
Input$2.00
Output$10.00
5-minute cache write$2.50
1-hour cache write$4.00
Cache read$0.20

Anthropic lists these rates in its Sonnet 5.5 launch announcement. Do not interpret “up to 30% less per task” as a 30% reduction in per-token rates: the input and output prices are unchanged from Sonnet 5. Sonnet 5's $2/$10 price is also now standard, after Anthropic cancelled its planned September increase.

Example: token efficiency versus rate changes

Consider a task that uses 100,000 input tokens and 10,000 output tokens, without caching. At Sonnet 5.5's published rates, that is $0.20 for input plus $0.10 for output, or $0.30 per task. If the same task used 30% fewer tokens in both categories, its modeled cost would be $0.21. This is an illustration of the arithmetic, not a promise that every task will use 30% fewer tokens.

For a workload of 10,000 such tasks per month, the baseline is $3,000. A 30% reduction in token volume would save $900, taking the modeled monthly bill to $2,100. Actual results vary with prompt length, cache reuse, output length, reasoning effort, retries, and tool calls.

What the launch benchmarks do and do not show

Anthropic reports gains on its coding and knowledge-work evaluations, including Terminal-Bench 4.0 and CursorBench. The company also reports that Sonnet 5.5 generates output more than 30% faster than Sonnet 5, and that in several benchmark comparisons it reaches previous Sonnet 5 scores at a fraction of the per-task cost. Those results are useful hypotheses for an evaluation plan, but they are vendor-published measurements rather than independent production cost data.

One practical caveat: a faster model does not automatically lower a token-metered API bill. The bill falls when usage per completed task falls, or when fewer tasks need retries or escalation. Faster generation may instead improve throughput or reduce latency while leaving token spend unchanged.

A safe way to forecast adoption

  1. Keep the rate card constant. Use $2/$10 per million input/output tokens as the starting point for Sonnet 5 and 5.5.
  2. Measure completed work. Compare cost per accepted task, not only tokens per request. Include retries, tool calls, and human rework where you can.
  3. Split by workload. Evaluate coding, support, extraction, and long-context tasks separately; a single blended savings rate can hide regressions.
  4. Use a range. Budget with 0% savings until internal tests provide evidence, then add measured scenarios. Treat Anthropic's “up to 30%” as an upper bound claim, not your forecast.
  5. Recheck quality and latency. A lower token bill is not a saving if acceptance drops or more expensive fallback models are needed.

FAQ

Did Sonnet 5.5 get cheaper per token?

No. Anthropic lists the same $2 input and $10 output per million token rates as Sonnet 5. Its cost-per-task claim comes from using fewer tokens to do comparable work.

Is 30% savings guaranteed?

No. Anthropic describes cost reductions of up to 30% for most work in its launch materials. Your result depends on workload, prompt caching, effort settings, and task completion quality.

What does a cached request cost?

Anthropic lists cache reads at $0.20 per million tokens, 5-minute writes at $2.50, and 1-hour writes at $4.00. Check the current Claude API pricing documentation before committing a forecast.

Should we switch production traffic immediately?

Run a representative evaluation first. Shadow or canary the model, compare cost per accepted result and latency, and keep a rollback route for workloads that regress.

Sources and related reading

Related


Want this applied to your own LLM spend? FinOps LLM runs a free audit of your AI costs and shows where the savings are. Book free audit →

Back to research