Order of LLM cost optimizations
Updated September 27, 2026 · first published September 27, 2026
A September 2026 r/LLMDevs thread asked how people actually control LLM costs in production, and what they'd do if engineering time were not the constraint. The top-voted reply is worth more than any vendor playbook: the order matters way more than people admit. They started with prompt caching and trimming system prompts way down. What failed spectacularly was fine-tuning a small model to handle the repetitive stuff — maintenance overhead was a nightmare and the quality kept drifting.
The order that works
- Prompt caching. Cheapest engineering hour you will ever spend. Cache reads on Opus 5.5 are $0.20 per million against $4.00 for fresh input — a 95% discount on the largest line item. If your system prompt, tool definitions, or retrieved context repeat across calls, you are leaving this on the table before you do anything else.
- Trim the system prompt. The same thread, same reply: system prompts down. Verbose instructions are billed on every call forever. This is a writing task, not an engineering task.
- Delete unused tools. Every tool definition ships in the context on every request whether or not the model calls it. The same mechanism as MCP tool definition overhead. Removing a tool is a one-line config change with a permanent effect.
- Route by task. Only now does model routing pay. A router built on step 1 is already cheaper before it switches a single model.
- Batch or defer. Batch API is a straight discount for anything that is not user-facing. Non-production traffic should be on batch from day one.
- Fine-tune last, if ever. The step teams reach for first, and the one that failed for this team. Maintenance overhead plus drifting quality means a permanent ops tax for a saving that caching and routing already captured.
Why the order is the whole point
Each step changes the token counts the next step operates on. Trim the system prompt before routing, and the cheap model handles more of your traffic for the same routing logic. Enable caching before fine-tuning, and the fine-tuning budget only has to cover the cache misses. Do it in the wrong order and you pay twice: a fine-tune that a prompt rewrite would have made unnecessary, or a router that routes 60K tokens of instructions through a cheap model it was never meant to carry.
The mistake pattern
The failure mode in the thread was fine-tuning a small model for repetitive work. The intended saving was real. What arrived instead was a model that needed re-training every time the repetitive work shifted, quality that drifted away from the underlying model as that model was upgraded upstream, and an ops burden that made the fine-tune unmaintainable. The repetitive work was probably 15% of volume and 2% of spend — a small number wrapped in a large ongoing cost.
Rehearse the sequence on one request
Take a single representative request and apply each step in order, recording total billed tokens:
| Step | What changes | Typical billed-token cut |
|---|---|---|
| Caching on | Repeat prefix becomes a cache read | 40–70% |
| System prompt trimmed | Fewer uncached tokens on every call | 10–25% |
| Unused tools removed | Smaller tool block | 5–15% |
| Routing | Cheaper model on most traffic | 30–60% |
| Batch for non-prod | Discount tier | 25–50% of that traffic |
These compound. Applied in order, the same request often lands 80–90% cheaper. Applied in the order teams usually reach for, you get a fine-tune on 15% of volume and a routing rule that never pays for itself.
Bottom line
Caching, trimming, and deleting unused tools are cheap, compounding, and low-risk. Routing and batching come after. Fine-tuning comes last or not at all. The list is not the hard part; the sequence is.
Related
- The verbose prompt tax
- MCP tool definition overhead
- Subagent fan-out concurrency limits
- Free tier is an acquisition cost
- Three agent budgets
Want this applied to your own LLM spend? FinOps LLM runs a free audit of your AI costs and shows where the savings are. Book free audit →