Claude Sonnet 5.5: the cost-aware model routing guide

Updated September 28, 2026 · first published September 28, 2026

Claude Sonnet 5.5 launched on September 28, 2026 at Sonnet 5's API rates: $2 per million input tokens and $10 per million output tokens. Anthropic positions it for well-scoped everyday work, bug fixes, coding, and polished documents, while Opus 5.5 is aimed at more complex work requiring careful judgment. For teams, the useful question is not which model wins a benchmark; it is which one meets each task's quality and latency target at the lowest cost per accepted result.

Anthropic says Sonnet 5.5 is more than 30% faster than Sonnet 5 and can cost up to 30% less per task. Those are launch claims. Use them to choose what to evaluate, not as automatic budget reductions.

What Anthropic announced

See Anthropic's launch announcement for the complete benchmark descriptions and caveats.

A routing policy to test

WorkloadStarting candidateWhat to measure
Routine coding changes, bug fixes, document editsSonnet 5.5Accepted change rate, retries, output tokens, elapsed time
Ambiguous, long-horizon work with high judgment requirementsCompare Sonnet 5.5 with Opus 5.5Quality at each effort setting, escalation frequency, end-to-end cost
High-volume simple classification or extractionCompare Sonnet with lower-cost models already in your stackQuality threshold, cost per valid record, exception handling
Latency-sensitive interactive workTest Sonnet 5.5 at the effort setting that meets the quality barTail latency, completion rate, and cost per completed task

This is a test matrix, not a claim that one model is always the right choice for a workload category. Prompt shape, context length, tool environment, and evaluation criteria can change the result.

Why list-price comparison is incomplete

Sonnet 5.5's per-token price did not fall at launch. The potential saving is in how many tokens and agent steps it needs to complete a task. A model that generates a shorter answer but triggers more tool calls may cost more overall; a model that finishes faster but uses the same tokens may improve responsiveness without changing the token bill.

For an agent, calculate a task's total as input tokens × input rate + output tokens × output rate + cache writes/reads + any separately billed tools or infrastructure. Compare the total only after confirming the task passes the same acceptance checks.

Measure the migration before changing your budget

  1. Choose a few hundred representative tasks, including failures and edge cases.
  2. Run Sonnet 5 and 5.5 with the same tools, prompts, context limits, and effort settings where possible.
  3. Record token charges, retries, tool calls, time to completion, and human acceptance.
  4. Compare cost per accepted task and report results by task family, not only as a global average.
  5. Canary the workloads that meet your gates. Keep a route back to the previous model while monitoring quality and spend.

Start forecasts with zero assumed savings. Replace that assumption with measured deltas as the evaluation produces evidence. For API rates and cache pricing, use the Anthropic pricing documentation.

FAQ

Is Sonnet 5.5 cheaper than Sonnet 5?

Its listed per-token rates are the same. Anthropic says it can cost up to 30% less per task because it uses fewer tokens, but results depend on workload.

Does Sonnet 5.5 replace Opus 5.5?

No. Anthropic presents the models for different task profiles. Evaluate both where task difficulty, failure cost, or quality requirements justify the comparison.

Is it available on AWS?

AWS announced availability on September 28 through Amazon Bedrock and Claude Platform on AWS. Check each service's current region, model access, and pricing details before planning deployment.

Sources and related reading

Related


Want this applied to your own LLM spend? FinOps LLM runs a free audit of your AI costs and shows where the savings are. Book free audit →

Back to research