Committed spend and reserved capacity discounts
Updated September 1, 2026 · first published September 1, 2026
Once your model spend is large enough to be noticed, a provider will offer you a discount for committing to a volume. Provisioned throughput, an annual commitment, a reserved capacity tier — the shapes differ, the trade does not. You are being paid to be predictable.
The discount is real. So is the risk, and it is not the one most teams price.
The three risks, in order of how much they cost
Underuse is the obvious one and the smallest. You commit to a volume, you use less, you pay for the gap. It is easy to model and easy to bound: commit under your conservative case, not your plan.
Being locked to a model is bigger. A commitment usually names a model or a family. The frontier moves in months. A twelve-month commitment to today's flagship is a bet that nothing meaningfully cheaper arrives — and something meaningfully cheaper has arrived every quarter for three years. The discount has to beat the price decline you would have got by doing nothing, not the price you pay today.
Losing the ability to switch is biggest and never appears in the model. A commit removes your leverage in the next negotiation, because both sides know you cannot walk. It also quietly kills routing work: nobody optimises traffic away from a model the company has prepaid for, even when a cheaper route is right there.
How to price it
Compare against the honest counterfactual. Not "list price times volume", but what you would plausibly have paid on-demand over the same period, including the price cuts and cheaper models that historically arrive within it. On a twelve-month horizon that counterfactual is often 20–40% below today's list, which eats most of a typical commit discount before you start.
Then check the shape of the deal. Commit to a floor, not a plan — the volume you are confident about even if a product underperforms. Prefer spend commitments to model commitments, since a dollar commitment lets you move to whatever the provider ships next. Insist on rollover or reallocation for unused volume. And keep the term short; in a market repricing every quarter, a twelve-month term is already long and a multi-year one is a different kind of decision.
When it is clearly right
Stable, high-volume, latency-sensitive production traffic on a model you have already run for two quarters without wanting to change it. That is a real forecast, and being paid for it is a good deal. Everything experimental, agentic, or growing fast should stay on-demand — those are exactly the workloads whose volume you cannot promise and whose model you will want to change.
Related
Related
Want this applied to your own LLM spend? FinOps LLM runs a free audit of your AI costs and shows where the savings are. Book free audit →