Claude Sonnet 5.5: the cost-aware model routing guide
Updated September 28, 2026 · first published September 28, 2026
Claude Sonnet 5.5 launched on September 28, 2026 at Sonnet 5's API rates: $2 per million input tokens and $10 per million output tokens. Anthropic positions it for well-scoped everyday work, bug fixes, coding, and polished documents, while Opus 5.5 is aimed at more complex work requiring careful judgment. For teams, the useful question is not which model wins a benchmark; it is which one meets each task's quality and latency target at the lowest cost per accepted result.
Anthropic says Sonnet 5.5 is more than 30% faster than Sonnet 5 and can cost up to 30% less per task. Those are launch claims. Use them to choose what to evaluate, not as automatic budget reductions.
What Anthropic announced
- Sonnet 5.5 retains the Sonnet 5 API list prices: $2/MTok input, $10/MTok output, $2.50/MTok for 5-minute cache writes, $4/MTok for 1-hour writes, and $0.20/MTok for cache reads.
- Anthropic reports output generation more than 30% faster than Sonnet 5 and up to 30% lower cost per task for most work due to token efficiency.
- The company describes Sonnet as strongest at well-scoped everyday tasks; Opus is designed for complex tasks that require careful judgment.
- Anthropic says Haiku 5.5 is expected in the coming weeks. It was not part of today's launch.
See Anthropic's launch announcement for the complete benchmark descriptions and caveats.
A routing policy to test
| Workload | Starting candidate | What to measure |
|---|---|---|
| Routine coding changes, bug fixes, document edits | Sonnet 5.5 | Accepted change rate, retries, output tokens, elapsed time |
| Ambiguous, long-horizon work with high judgment requirements | Compare Sonnet 5.5 with Opus 5.5 | Quality at each effort setting, escalation frequency, end-to-end cost |
| High-volume simple classification or extraction | Compare Sonnet with lower-cost models already in your stack | Quality threshold, cost per valid record, exception handling |
| Latency-sensitive interactive work | Test Sonnet 5.5 at the effort setting that meets the quality bar | Tail latency, completion rate, and cost per completed task |
This is a test matrix, not a claim that one model is always the right choice for a workload category. Prompt shape, context length, tool environment, and evaluation criteria can change the result.
Why list-price comparison is incomplete
Sonnet 5.5's per-token price did not fall at launch. The potential saving is in how many tokens and agent steps it needs to complete a task. A model that generates a shorter answer but triggers more tool calls may cost more overall; a model that finishes faster but uses the same tokens may improve responsiveness without changing the token bill.
For an agent, calculate a task's total as input tokens × input rate + output tokens × output rate + cache writes/reads + any separately billed tools or infrastructure. Compare the total only after confirming the task passes the same acceptance checks.
Measure the migration before changing your budget
- Choose a few hundred representative tasks, including failures and edge cases.
- Run Sonnet 5 and 5.5 with the same tools, prompts, context limits, and effort settings where possible.
- Record token charges, retries, tool calls, time to completion, and human acceptance.
- Compare cost per accepted task and report results by task family, not only as a global average.
- Canary the workloads that meet your gates. Keep a route back to the previous model while monitoring quality and spend.
Start forecasts with zero assumed savings. Replace that assumption with measured deltas as the evaluation produces evidence. For API rates and cache pricing, use the Anthropic pricing documentation.
FAQ
Is Sonnet 5.5 cheaper than Sonnet 5?
Its listed per-token rates are the same. Anthropic says it can cost up to 30% less per task because it uses fewer tokens, but results depend on workload.
Does Sonnet 5.5 replace Opus 5.5?
No. Anthropic presents the models for different task profiles. Evaluate both where task difficulty, failure cost, or quality requirements justify the comparison.
Is it available on AWS?
AWS announced availability on September 28 through Amazon Bedrock and Claude Platform on AWS. Check each service's current region, model access, and pricing details before planning deployment.
Sources and related reading
- Anthropic: Introducing Claude Sonnet 5.5
- AWS: Claude Sonnet 5.5 now available on AWS
- Sonnet 5.5 pricing and task-cost model
- Claude Opus 5.5 cost model
Related
Want this applied to your own LLM spend? FinOps LLM runs a free audit of your AI costs and shows where the savings are. Book free audit →