Silent model downgrades and cost models
Updated September 27, 2026 · first published September 27, 2026
On 23 September 2026, an r/GroundTruthAINews thread pointed out a footnote buried in Anthropic's Opus 5.5 launch notes: an Opus 5.5 request may be served by Opus 4.8. Anthropic and OpenAI both shipped price cuts the same afternoon. Both labs sold the same idea: keep the capability, spend far less. The footnote is the part that matters to FinOps.
What actually happened
Claude Opus 5.5 launched at $4/$20 per million tokens, roughly 20% under Opus 5, with cache reads cut 60% to $0.20. Anthropic says fewer tokens per task and 30% faster output make typical workloads about 40% cheaper. Roughly 90 minutes later OpenAI shipped GPT-6 Sol at $2/$10 and Luna at $0.10/$0.50, both half the GPT-5.6 prices.
Every launch post from this period carries the same kind of footnote: the endpoint you call is an alias, and the alias is load-balanced across model versions during a transition window. That is a normal capacity practice. It is also a cost-modeling problem.
Why it breaks the numbers
- Cost per task drifts. Your forecast assumed Opus 5.5 output pricing. If a share of calls land on 4.8, both the unit price and the token count move, in opposite directions.
- Evals measure a mixture. An eval suite run during a transition window scores a blend of versions. The score is not reproducible later.
- Price cuts get attributed wrongly. If a 40% saving shows up as 15%, the gap is usually version mixing, not a broken optimization.
- Latency baselines go stale. 4.8 is not as fast as 5.5. Any p50 you captured mid-transition is not your steady-state p50.
How to detect the swap
Every major provider returns the resolved model in the response body. Log it. One field, four lines of code:
resolved = response.model # e.g. "claude-opus-4-8-20260115"
assert resolved.startswith(EXPECTED_PREFIX) or resolved in ALLOWED_FALLBACKSThen split every cost metric by resolved model, not by the model you asked for. The moment one bucket starts drifting, you have your answer: the alias is mixing versions and your per-task cost is a weighted average of two different products.
The budget implication
During a transition, plan against the blended cost, not the headline. If 20% of calls drop to an older model, your effective input rate is not $4.00 per million — it is whatever the mix produces. The same logic applies to the headline saving: the 40% cheaper claim is a typical-workload average, which already assumes a token-count reduction that may not hold if the cheaper-but-older model is the one generating the longer answers.
Rules to hold
- Never model on an alias. Pin a dated model ID for anything you budget or evaluate.
- Log the resolved model on every request. Non-negotiable if you ever re-run an eval.
- Re-baseline after the transition window. Usual duration is weeks, not quarters, but confirm.
- Price a transition as a mixture. Ask your provider for the expected split, or measure it.
Bottom line
A price cut and a silent downgrade can land in the same afternoon. The cut is the headline; the downgrade is the footnote that quietly invalidates last month's forecast. Log the resolved model, pin dated IDs, and price the mixture.
Related
- Claude Opus 5.5 cost model
- GPT-6 pricing: Sol, Luna, Astra
- Model price cuts and routing budgets
- Provider prices are not comparable
- Evals cost control
Want this applied to your own LLM spend? FinOps LLM runs a free audit of your AI costs and shows where the savings are. Book free audit →