98% of our Codex input tokens were cached. Without caching the bill is 12x bigger
Updated October 9, 2026 · first published October 9, 2026
In 12 days of heavy Codex use on ChatGPT Pro $100 we sent 13.0 billion input tokens and got back 47 million output tokens. 97.6% of that input was a cache hit. At OpenAI API prices the work came to $1,211. Price the same tokens with no cache hits and it is $14,198, 12 times more.
- Input to output ratio: about 279 input tokens for every output token.
- Cached input share of the bill: 60%, even at a 95% discount.
- Output share of the bill: 16%.
Why coding agents are almost all cached input
An agent loop re-sends the whole conversation on every step: system prompt, tool definitions, the files it has read, every earlier tool result. Each step adds a little and re-reads everything. So input grows with the square of the session length while output stays small. Over 12 days that gave us a 279:1 input to output ratio.
Prompt caching is what keeps this affordable. Cached input on OpenAI's API costs a tenth or less of normal input. GPT-6 Sol, which did most of our work, charges $2 per million input tokens and $0.10 per million cached.
Where the money went
| Model | Input tokens | Cached | Output tokens | API value | If uncached |
|---|---|---|---|---|---|
| Sol | 6.56B | 98.1% | 17.5M | $1,073 | $13,299 |
| Luna | 6.41B | 97.2% | 29.0M | $95 | $655 |
| GPT-5.5 | 0.03B | 93.5% | 0.2M | $27 | $144 |
| Astra | 0.01B | 94.5% | 0.0M | $16 | $99 |
| Total | 13.01B | 97.6% | 47M | $1,211 | $14,198 |
| Cost component | API value | Share |
|---|---|---|
| Cached input | $728 | 60% |
| Uncached input | $288 | 24% |
| Output | $196 | 16% |
Even at a 95% discount, cached input is the single largest line. That is the signature of an agent workload: the cheapest token type, in enormous volume. Per-model totals are in codex-pro-100-model-mix.csv.
What breaks the cache
A cache hit needs an identical prefix. Anything that changes the start of the context forces a full-price re-read of everything after it.
- Switching models partway through a session. The new model has no cache for your context.
- Editing instructions or tool lists mid-session. They sit at the top of the prompt, so everything below is re-read.
- Long pauses. Caches expire. Coming back after a break re-reads the session at full price.
- Compaction. Summarizing the history rewrites the prefix. It saves tokens later but costs one uncached pass now.
Why this matters for your Codex limits
If the Codex gauge weighs cached input anywhere near its API price, a 1-point drop in your cache hit rate is expensive. At our volume, Sol going from 98.1% to 90% cached would add about $1,004 of Sol input alone, roughly 6 extra Pro $100 windows. Cache discipline is quota discipline. For what a window holds, see our Pro $100 multiplier measurement; for how we got eight windows in 12 days, see the reset log.
How we measured it
Codex writes a JSONL log of every session under ~/.codex/sessions/. Each request records cumulative input, cached input and output tokens, plus the weekly gauge (used_percent) and its resets_at time. We priced every request's token delta at the API list price of the model it ran on and grouped requests by weekly window. Dollar figures are API-equivalent value, not money we paid.
Sources
Related
- ChatGPT Pro $100 is really 2.6x Plus
- 8 Codex weekly windows in 12 days
- Codex resets turned $40 into $1,211
- How much Codex ChatGPT Pro includes
- Claude Max 20x vs ChatGPT Pro
Want this applied to your stack? Bring the provider bills, gateway logs, and top workflows; we will map the cost drivers and savings path. Book a free audit →