Skip to content

Claude Haiku 5.5 and the 100k-token price threshold

Updated October 8, 2026 · first published October 8, 2026

Quick answer: Claude Haiku 5.5's headline price depends on prompt length. Anthropic released the model on October 7, 2026 for high-volume work such as classification, summaries, and subagent tasks. Its...

Claude Haiku 5.5's headline price depends on prompt length. Anthropic released the model on October 7, 2026 for high-volume work such as classification, summaries, and subagent tasks. Its 1-million-token context window is not a promise that all of those tokens cost the short-prompt rate: a prompt over 100,000 tokens uses a higher price tier. That makes request sizing a budget decision, especially for agents that accumulate documents and tool results.

The rate card has two sides

The following are Anthropic's published Claude API standard prices in US dollars per million tokens, checked October 8. The boundary is up to 100,000 prompt tokens versus over 100,000; it is not a price charged only to the tokens beyond the threshold. Check the live pricing documentation before using the figures in a forecast.

Token categoryPrompt up to 100kPrompt over 100k
Uncached input$0.10$0.50
Output$0.50$2.50
5-minute cache write$0.125$0.625
1-hour cache write$0.20$1.00
Cache read$0.01$0.05

Each listed token rate is five times higher above the boundary. That does not mean every application bill rises fivefold: most calls may remain short, and output volume, caching, tools, batch processing, retries, and task success all affect total spend. Cached context still occupies the prompt; a cache hit does not turn a 120k-token request into a short prompt.

What the boundary looks like on one request

Consider two illustrative calls with no caching or separately billed tools, each producing 1,000 output tokens. A 99,000-token prompt costs about $0.0104 in input plus output at the short-prompt rates. A 101,000-token prompt costs about $0.0530 at the long-prompt rates. That is roughly five times as much for a request only 2,000 input tokens longer. The example holds output fixed to isolate the pricing boundary; it is not a measured production saving. Real prompts also include system instructions, conversation history, tool schemas, and tool results.

Recount before migrating

Anthropic's migration notes say the same text produces approximately 30% more tokens on Haiku 5.5 than on Haiku 4.5, with variation by content. A prompt that looked comfortably under 100k using old token counts may cross the new boundary. Do not multiply every old count by 1.3 and treat that as an invoice forecast; recount representative requests with claude-haiku-5-5.

Anthropic's token-counting endpoint accepts the structured message, system prompt, and supported client tools before sending the generation request. Count against the target model, then log actual usage after the response. The preflight count is an estimate and can differ slightly from billed input; some server tools and URL/file blocks are not supported by the counter. For those paths, rely on observed request usage and keep a safety margin near the threshold.

A prompt-budget rule for subagents

  1. Measure the distribution. Record Haiku 5.5 prompt-token counts by workload, model ID, and task result. Watch how many calls cluster near or exceed 100k rather than looking only at the average.
  2. Warn before the boundary. Flag near-threshold requests in the agent orchestrator. Treat the warning level as your own operational margin, not an Anthropic price rule.
  3. Trim the cause. Remove duplicated retrieval chunks, cap tool-result size, or compact stale history where quality tests permit. Do not discard evidence an agent needs merely to save tokens.
  4. Compare the long tail. For unavoidable long prompts, test Haiku at its higher tier against another route on the same tasks. Compare cost per accepted result, latency, and fallback frequency—not just list prices.
  5. Reconcile to the invoice. Price actual input, output, cache reads and writes at the applicable tier, then include tools and retries. Recheck provider or cloud-platform rates and any contract discounts.

Haiku 5.5 may still be the best choice above 100k, but that should be a measured routing decision. The important control is to know which requests cross the line and why.

Questions teams ask

Does the higher price apply only to tokens above 100k?

No. Anthropic describes Haiku 5.5 as priced by prompt length: a prompt over 100,000 tokens pays the higher published rates for that request.

Can prompt caching keep a long request in the cheaper tier?

No. Reused tokens can receive a cache-read rate, but cached context still occupies prompt length. Count the full prompt when checking the threshold.

Is a 30% tokenizer increase guaranteed?

No. Anthropic says approximately 30% more tokens for the same text than on Haiku 4.5, depending on content. Count your own prompts on the target model.

Related


Want this applied to your stack? Bring the provider bills, gateway logs, and top workflows; we will map the cost drivers and savings path. Book a free audit →

Back to finopsllm.com