FREE TOOL · UPDATED DAILY

LLM API Pricing Tracker.

Live prices for every major model. Input, output, cache-read, and cache-write costs across all providers. Updated daily at 00:00 UTC.

All models

Model Provider Input / 1M Output / 1M Cache read / 1M Cache write / 1M Context Updated

Why list prices don't predict your bill

The per-million-token numbers in this table are the starting point of a cost estimate, not the estimate. Four things move the real figure, usually by more than the gap between two providers on this page.

How often do these prices change?

Less often than model availability, but often enough that a quarterly review is the minimum. Providers cut prices when competitors do, introduce cheaper tiers of existing models, and deprecate older ones with a migration window. The practical consequence is that routing rules and budget assumptions both go stale — a routing table written six months ago is usually still correct and occasionally sending traffic to a model that is no longer the cheapest option that clears your quality bar.

Context windows are a ceiling, not a target

The context column shows the maximum a model accepts. It is not a budget to fill. Cost scales with what you actually send, and retrieval that grows context faster than it improves answers is one of the most common causes of a bill that rises without the product getting better. Measure the p95 context length you send in production; that number, not the ceiling, is what your invoice reflects.

How we track prices

Prices are scraped daily from provider pricing pages and verified against API responses. Cache prices apply only where providers support prompt caching. Context windows are maximum supported; effective context may be lower for reasoning models.

Last updated:

Related

← Back to finopsllm.com