Claude Sonnet 5 intro pricing ends August 31: your bill rises 50%
Published 19 July 2026
Anthropic launched Claude Sonnet 5 on June 30, 2026 at introductory pricing of $2/1M input tokens and $10/1M output tokens. That price is a hard deadline: August 31, 2026. Starting September 1, Sonnet 5 moves to standard pricing of $3/$15 per 1M—a flat 50% cost increase. If you're running Sonnet 5 workloads, the next 6 weeks are your window to lock in costs and plan for the hike.
- Intro pricing valid through August 31, 2026 only
- Standard pricing ($3/$15) begins September 1, 2026
- Cost increase is exactly 50% for Sonnet 5 workloads
- Applies to Anthropic API, Bedrock, and Vertex AI
- 43 days to re-benchmark, lock commitments, and re-forecast
The Math: What 50% Looks Like at Scale
The jump from intro to standard pricing is clean and large. Let's model a realistic workload:
| Workload Profile | Monthly Input Tokens | Monthly Output Tokens | August Cost (Intro) | September Cost (Standard) | Monthly Delta | Annual Impact |
|---|---|---|---|---|---|---|
| Light (batch/internal) | 50M | 20M | $200 | $300 | +$100 | +$1,200 |
| Mid-tier (customer-facing) | 500M | 150M | $2,500 | $3,750 | +$1,250 | +$15,000 |
| Enterprise (high-volume reasoning) | 2B | 800M | $12,000 | $18,000 | +$6,000 | +$72,000 |
| Heavy (multi-request flows) | 5B | 2.5B | $35,000 | $52,500 | +$17,500 | +$210,000 |
Key insight: A company with a mid-tier Sonnet 5 workload (500M input + 150M output tokens/month) will see a $1,250 monthly increase that compounds to $15,000 annually if nothing changes. For enterprise workloads, the annual impact can exceed $70K.
Who Is Affected
The deadline applies across all channels:
Anthropic API (direct)
Introductory pricing ends August 31, 2026 at 11:59 PM UTC. Verified on Anthropic's pricing docs. No extension or delay has been announced.
AWS Bedrock
Bedrock passes through Anthropic's pricing. Intro pricing applies to Sonnet 5 on Bedrock through August 31. The hike will appear in your Bedrock bill on September 1. Check your AWS contract for any commitment discounts that might offset the increase.
Google Vertex AI
Vertex AI integrates Claude models with Anthropic pricing pass-through. The transition happens on September 1. Vertex bills are quoted in your local currency; confirm the exact date with your Google account team if using a different timezone.
GitHub Copilot Enterprise
If Copilot uses Sonnet 5 as a backend, Microsoft may or may not pass through the price increase to your plan. Copilot Enterprise is a fixed seat license, so the impact depends on whether Microsoft absorbs or reflects the backend cost. Check your Copilot invoice or contact your Microsoft account team.
Three Strategies: What to Do Before August 31
Strategy 1: Lock Usage Commitments (If Available)
Check if Anthropic offers usage commitments or volume discounts for Sonnet 5. Volume-based discounts or annual prepayment agreements could reduce the effective price increase. Contact Anthropic sales now—they may hold intro pricing on commitments signed before August 31.
Action: Email Anthropic sales by August 15 if you're spending >$5K/month on Sonnet 5. Ask whether an annual commitment locks intro pricing or offers tiered discounts. Lead time for commitment contracts is typically 2–3 weeks.
Strategy 2: Re-Benchmark Routing and Model Mix
Not all workloads need Sonnet 5. The 50% price jump is a forcing function to audit which tasks actually require Sonnet 5 and which can shift to cheaper alternatives.
| Task Category | Sonnet 5 Input/$1M | Alternative Model | Alt Price Input/$1M | Savings per 100M Tokens | When to Shift |
|---|---|---|---|---|---|
| Classification (strict rules) | $3 (post-Sep 1) | Claude Haiku 4.5 | $1.00 | $200 | Accuracy ≥95% |
| Summarization | $3 | Gemini 2.5 Flash | $0.15 | $285 | No reasoning needed |
| Multi-turn conversation | $3 | GPT-5.6 Luna | $1.00 | $200 | Quality equivalent |
| Complex reasoning (keep Sonnet 5) | $3 | — | — | — | No alternative |
| Agentic loops (multi-step) | $3 | Haiku + Sonnet 5 hybrid | $1.00 (Haiku) + $3 (Sonnet 5) | Save by using Haiku for simple steps | Benchmarks support |
Action: Over the next 2 weeks, pull your Sonnet 5 call logs and categorize by workload. Benchmark 10% of classification tasks on Haiku, summarization on Gemini 2.5 Flash, and simple multi-turn on GPT-5.6 Luna. If quality is equivalent (or close), shift that workload before September 1 and lock the savings.
Strategy 3: Implement (or Expand) Prompt Caching
Prompt caching is the most direct way to reduce token consumption and preserve the value of intro pricing before it expires. Cache writes to Sonnet 5 cost $2.50 per 1M tokens (1.25× input price on intro; will be $3.75 on standard). Cache reads cost $0.20 per 1M (0.1× input price on intro; will be $0.30 on standard).
The opportunity: A workload with 40% cache hit rate saves 36% on input cost. Locking cache infrastructure now at intro pricing means your ongoing hit-rate savings are grandfathered in. After September 1, the same cache hits are still 90% off, but the base price they discount from is 50% higher.
Example: System prompt + context cached 100M tokens/month on Sonnet 5:
- August (intro): Write cost: 100M × $2.50/M = $250. Monthly read cost (assume 1B read tokens from cache hits): 1B × $0.20/M = $200. Total: $450/month for caching.
- September (standard): Write cost: 100M × $3.75/M = $375. Monthly read cost: 1B × $0.30/M = $300. Total: $675/month (50% higher).
- Savings from implementing caching now: If you add 500M cached input tokens between now and August 31, and they accumulate 30 cache hits each month starting September, you lock in cumulative read discounts. The cache-write cost is a sunk investment that pays off for months.
Action: Identify your top 3 repeated prompts or document contexts (RAG documents, system instructions, boilerplate). Test caching on Anthropic API. If hit rate is ≥20%, implement before August 31. Cache writes are the one cost that doesn't care about when the rate hike happens—investing in write cost now yields fixed read-cost returns forever.
Timeline: The 6-Week Countdown
Week 1 (by July 25)
Audit current Sonnet 5 spend. Extract call logs (model, tokens, workload type) from your billing dashboard or LLM observability platform. Calculate monthly run rate. Model Q4 spend if nothing changes.
Week 2 (by August 1)
Contact Anthropic sales to inquire about usage commitments. Start benchmarking alternative models (Haiku, Luna, Gemini 2.5 Flash) on 10% of production traffic. Identify top 3 candidates for routing shift.
Week 3 (by August 8)
Finalize benchmarks and commit to a routing plan. If shifts are needed, start deploying A/B tests. Implement prompt caching for high-volume, repeated contexts. Target ≥100M tokens cached before cutover.
Week 4 (by August 15)
Lock any usage commitments with Anthropic (if applicable). Deploy routing changes to production (shift low-reasoning tasks to cheaper models). Verify cache infrastructure is live and hitting 15%+ hit rate.
Week 5 (by August 22)
Monitor production costs and quality metrics. Adjust routing if needed. Forecast final September costs based on August run rate.
Week 6 (by August 31)
Confirm all routing changes are live. Set calendar alert for September 5 to verify billing matches forecast. Publish post-mortem on cost impact and savings achieved.
Comparison: Sonnet 5 vs. Alternatives After September 1
Once intro pricing expires, how does Sonnet 5 stack up against other models at their standard pricing?
| Model | Provider | Input $/1M | Output $/1M | Typical Use Case | vs. Sonnet 5 Post-Sep 1 |
|---|---|---|---|---|---|
| Claude Sonnet 5 (standard) | Anthropic | $3.00 | $15.00 | Balance of speed & reasoning | Baseline |
| GPT-5.6 Terra | OpenAI | $2.50 | $15.00 | Mid-tier reasoning | 17% cheaper input |
| GPT-5.6 Luna | OpenAI | $1.00 | $6.00 | Fast classification, simple tasks | 67% cheaper overall |
| Gemini 3.1 Pro | $2.00 | $12.00 | Agentic loops, tool use | 33% cheaper input | |
| Gemini 2.5 Flash | $0.15 | $0.60 | Ultra-fast, high-volume | 95% cheaper | |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | Low-latency classification | 67% cheaper input |
| DeepSeek V4 Flash | DeepSeek | $0.14 | $0.28 | Cost-optimized, open model | 95% cheaper (but less polished) |
The reality: Sonnet 5 at $3/$15 is no longer the cheapest or most cost-effective model for most tasks. GPT-5.6 Terra is 17% cheaper on input. Haiku is 67% cheaper for classification. The value of Sonnet 5 post-hike is in quality for complex reasoning and speed for balanced tasks—not in cost leadership. Workloads that don't require that quality should move to alternatives now.
FAQs
Q: Is there any chance the deadline extends?
A: Anthropic has not announced an extension. The June 30 launch announcement clearly states "through August 31, 2026." Prepare for the hard date. If an extension is announced, it's a bonus; treat the deadline as final.
Q: Does the newTokenizer affect token counts?
A: Yes. Sonnet 5 uses a new tokenizer that produces approximately 30% more tokens than older Claude models for equivalent text. This means your actual token spend is higher than you might expect. Factor this into cost forecasts.
Q: Can I lock intro pricing with an annual prepayment?
A: Possibly, if Anthropic offers prepayment options. Contact sales by August 15 to ask. Prepayment is often the only way to grandfather pricing through a rate increase, but you need to commit well in advance.
Q: What if I'm on a custom Anthropic contract?
A: Custom contracts may have different terms. Verify the exact cutover date and rate with your account team. If your contract specifies pricing terms, they take precedence over the public pricing page—confirm now before September 1 surprises.
The August 31 deadline is real, the 50% hike is certain, and you have 6 weeks to act. The teams that re-benchmark, lock commitments, and shift workloads before the cutover will cut their Q4 LLM spend significantly. Teams that wait until September 1 will absorb the full cost increase.