Skip to content

98% of our Codex input tokens were cached. Without caching the bill is 12x bigger

Updated October 9, 2026 · first published October 9, 2026

Quick answer: In 12 days of heavy Codex use on ChatGPT Pro $100 we sent 13.0 billion input tokens and got back 47 million output tokens. 97.6% of that input was a cache hit. At OpenAI API prices the work came to...

In 12 days of heavy Codex use on ChatGPT Pro $100 we sent 13.0 billion input tokens and got back 47 million output tokens. 97.6% of that input was a cache hit. At OpenAI API prices the work came to $1,211. Price the same tokens with no cache hits and it is $14,198, 12 times more.

API value of 12 days of Codex by model: cached vs uncached$0$2k$4k$6k$8k$10k$12k$14k$1,073$13,299Sol$95$655Luna$27$144GPT-5.5$16$99AstraOrange: what we paid in API terms, with cachingBlue: the same tokens with no cache hits
API value of 12 days of Codex by model: cached vs uncached

Why coding agents are almost all cached input

An agent loop re-sends the whole conversation on every step: system prompt, tool definitions, the files it has read, every earlier tool result. Each step adds a little and re-reads everything. So input grows with the square of the session length while output stays small. Over 12 days that gave us a 279:1 input to output ratio.

Prompt caching is what keeps this affordable. Cached input on OpenAI's API costs a tenth or less of normal input. GPT-6 Sol, which did most of our work, charges $2 per million input tokens and $0.10 per million cached.

Where the money went

ModelInput tokensCachedOutput tokensAPI valueIf uncached
Sol6.56B98.1%17.5M$1,073$13,299
Luna6.41B97.2%29.0M$95$655
GPT-5.50.03B93.5%0.2M$27$144
Astra0.01B94.5%0.0M$16$99
Total13.01B97.6%47M$1,211$14,198
Cost componentAPI valueShare
Cached input$72860%
Uncached input$28824%
Output$19616%

Even at a 95% discount, cached input is the single largest line. That is the signature of an agent workload: the cheapest token type, in enormous volume. Per-model totals are in codex-pro-100-model-mix.csv.

What breaks the cache

A cache hit needs an identical prefix. Anything that changes the start of the context forces a full-price re-read of everything after it.

Why this matters for your Codex limits

If the Codex gauge weighs cached input anywhere near its API price, a 1-point drop in your cache hit rate is expensive. At our volume, Sol going from 98.1% to 90% cached would add about $1,004 of Sol input alone, roughly 6 extra Pro $100 windows. Cache discipline is quota discipline. For what a window holds, see our Pro $100 multiplier measurement; for how we got eight windows in 12 days, see the reset log.

How we measured it

Codex writes a JSONL log of every session under ~/.codex/sessions/. Each request records cumulative input, cached input and output tokens, plus the weekly gauge (used_percent) and its resets_at time. We priced every request's token delta at the API list price of the model it ran on and grouped requests by weekly window. Dollar figures are API-equivalent value, not money we paid.

Sources

Related


Want this applied to your stack? Bring the provider bills, gateway logs, and top workflows; we will map the cost drivers and savings path. Book a free audit →

Back to finopsllm.com