LLM API Pricing Tracker
Updated 12 July 2026 · first published 4 July 2026
Track per-token pricing across all major LLM providers. Input, output, cache read/write, context windows, and batch discounts in one normalised table. Updated monthly with verifiable provider data.
- All prices in USD per 1 million tokens
- Estimated values shown in gray
- Cache prices reflect prompt caching where available
Methodology
Collection
Pricing data collected from official provider pricing pages and API documentation.
Normalisation
All prices normalised to USD per 1 million tokens for direct comparison.
Cache pricing
Cache Read/Write columns show prompt caching discounts where providers support them.
Display policy
Official values and estimated values are clearly separated. Estimates marked in gray.
Update Feed
LLM API Pricing Tracker launched
Published per-token pricing comparison across 12+ major LLM providers as a FinOps LLM research article.
Per-Token Pricing Table
| Provider ↕ | Model ↕ | Input $/1M ↕ | Output $/1M ↕ | Cache Read $/1M ↕ | Cache Write $/1M ↕ | Context Window ↕ | Batch Discount |
|---|---|---|---|---|---|---|---|
| OpenAI | GPT-5.6-Sol | $5.00 | $30.00 | $0.50 | $6.25 | 128K | 50% |
| OpenAI | GPT-5.6-Terra | $2.50 | $15.00 | $0.25 | $3.13 | 128K | 50% |
| OpenAI | GPT-5.6-Luna | $1.00 | $6.00 | $0.10 | $1.25 | 128K | 50% |
| OpenAI | GPT-5.5 | $5.00 | $30.00 | $0.50 | $6.25 | 128K | 50% |
| OpenAI | GPT-5.4 | $2.50 | $15.00 | $0.25 | $3.13 | 128K | 50% |
| OpenAI | o3 | $2.00 | $8.00 | $0.50 | $2.00 | 200K | 50% |
| OpenAI | o4-mini | $1.10 | $4.40 | $0.28 | $1.10 | 200K | 50% |
| Anthropic | Claude Fable 5 | $10.00 | $50.00 | $1.00 | $12.50 | 200K | 50% |
| Anthropic | Claude Opus 4.8 | $5.00 | $25.00 | $0.50 | $6.25 | 200K | 50% |
| Anthropic | Claude Sonnet 5* | $2.00 | $10.00 | $0.20 | $2.50 | 200K | 50% |
| Anthropic | Claude Haiku 4.5 | $1.00 | $5.00 | $0.10 | $1.25 | 200K | 50% |
| Gemini 3.5 Flash | $1.50 | $9.00 | $0.15 | $1.88 | 1M | 50% | |
| Gemini 3.1 Pro Preview | $2.00 | $12.00 | $0.20 | $2.50 | 1M | 50% | |
| Gemini 2.5 Pro | $1.25 | $5.00 | $0.31 | $4.50 | 1M | 50% | |
| Gemini 2.5 Flash | $0.15 | $0.60 | $0.04 | $0.15 | 1M | 50% | |
| xAI | Grok 4.5 | $2.00 | $6.00 | $0.50 | $1.50 | 1M | 50% |
| xAI | Grok 4.3 | $1.25 | $2.50 | - | - | 1M | 50% |
| Meta | Llama 4 Maverick (DeepInfra) | $0.20 | $0.80 | - | - | 1M | - |
| Meta | Llama 4 Scout (Groq) | $0.11 | $0.34 | - | - | 5M | - |
| Meta | Llama 3.3 70B (Groq) | $0.59 | $0.79 | - | - | 8M | - |
| Mistral | Mistral Large 2/3 | $2.00 | $6.00 | - | - | 128K | - |
| Mistral | Mistral Medium 3.5 | $1.50 | $7.50 | - | - | 128K | - |
| Mistral | Mistral Small 3 | $0.10 | $0.30 | - | - | 128K | - |
| Mistral | Codestral | $0.30 | $0.90 | - | - | 128K | - |
| DeepSeek | DeepSeek V4 Flash | $0.14 | $0.28 | $0.0028 | - | 128K | - |
| DeepSeek | DeepSeek V4 Pro | $0.435 | $0.87 | $0.0036 | - | 128K | - |
| Alibaba Qwen | Qwen 3.7-Max | $2.50 | $7.50 | - | - | 131K | - |
| Moonshot Kimi | Kimi K2.6 | $0.95 | $4.00 | $0.16 | - | 200K | - |
| Xiaomi MiMO Sign up - both get $2 ↗ |
MiMo-V2-Pro | $0.50 | $2.00 | - | - | 128K | - |
| Xiaomi MiMO Sign up - both get $2 ↗ |
MiMo-V2-Omni | $0.40 | $1.60 | - | - | 128K | - |
| Zhipu AI | GLM-5 | $1.00 | $4.00 | - | - | 128K | - |
| Zhipu AI | GLM-5-Turbo | $0.30 | $1.20 | - | - | 128K | - |
| MiniMax | M2.7 | $0.40 | $1.60 | - | - | 128K | - |
| StepFun | step-3.5-flash | $0.30 | $1.20 | - | - | 128K | - |
* Claude Sonnet 5: introductory pricing of $2/$10 per 1M tokens in effect through August 31, 2026; standard pricing $3/$15 begins September 1, 2026.
Values not explicitly provided by providers are shown in gray.
Meta Llama prices reflect hosted API pricing (Fireworks, Together, etc.), not self-hosted cost. Self-hosting economics are covered in our Open-Source vs Closed-Source Cost Comparison.
Context windows are maximum supported; actual availability may vary by tier.
Cache Read/Write pricing reflects provider-specific prompt caching where available. "-" means the provider does not offer caching or pricing is not published.
Batch Discount shows the approximate savings when using batch/async APIs where available.
Some links may include referral codes.
Prices can change frequently. Always verify the latest information on official provider pages.
Practical Guide
How to use this pricing data for cost planning
1) Calculate your token volume first
Before comparing prices, estimate your monthly input and output token volumes. Output tokens are typically 3-5x more expensive than input, so reducing output length has the biggest impact on cost.
2) Factor in caching savings
If your workload has repeated system prompts or context (RAG, multi-turn conversations), prompt caching can reduce input costs by 50-90%. Check the Cache Read column for applicable discounts.
3) Use batch APIs for non-urgent work
Most major providers offer 50% discounts on batch/async API calls. If your workload tolerates latency (evaluations, data processing, content generation), batch pricing cuts costs in half.
4) Compare total cost, not just per-token price
A cheap model that needs 3x more tokens to match quality may cost more overall. Factor in model quality and task-specific performance alongside raw pricing.
Next step: Compare your spend across providers—see Azure OpenAI vs Direct OpenAI Cost for a breakdown of managed vs direct pricing models.
Want this applied to your own LLM spend? FinOps LLM runs a free audit of your AI costs and shows where the savings are. Book free audit →