LLM Cost Calculator
Updated 12 July 2026
Estimate your monthly LLM spend across popular models. Enter your usage pattern below and see costs compared side-by-side. All prices are per-million-tokens as of July 2026.
Pre-calculated costs by usage tier
Here's what each usage tier costs per month across popular models, based on typical request patterns:
| Model | Light (1K req/mo) | Medium (10K req/mo) | Heavy (100K req/mo) |
|---|---|---|---|
| GPT-5.6-Sol | $8.75 | $87.50 | $875.00 |
| GPT-5.6-Terra | $4.38 | $43.75 | $437.50 |
| GPT-5.6-Luna | $1.75 | $17.50 | $175.00 |
| Claude Fable 5 | $17.50 | $175.00 | $1,750.00 |
| Claude Opus 4.8 | $8.75 | $87.50 | $875.00 |
| Claude Sonnet 5 | $3.50 | $35.00 | $350.00 |
| Claude Haiku 4.5 | $1.75 | $17.50 | $175.00 |
| Gemini 3.5 Flash | $2.63 | $26.25 | $262.50 |
| Grok 4.5 | $4.00 | $40.00 | $400.00 |
| DeepSeek V4 Pro | $0.95 | $9.50 | $95.00 |
| Kimi K2.6 | $1.70 | $17.00 | $170.00 |
Light: 1K requests/mo, 500 input + 200 output tokens. Medium: 10K requests/mo, 1,000 input + 500 output tokens. Heavy: 100K requests/mo, 2,000 input + 1,000 output tokens.
Interactive calculator
How to use this calculator
- Pick your model from the dropdown. Prices reflect standard API rates as of July 2026.
- Enter your request volume - how many API calls you make per month.
- Estimate token counts - the average number of input and output tokens per request. If you're unsure, start with 1,000 input / 500 output for chat, or 3,000 input / 1,000 output for code generation.
- Hit Calculate to see your estimated monthly cost split by input and output.
How to reduce your estimate
- Use prompt caching. If your prompts have a stable prefix, cached reads can cut input costs by 50–90%. See prompt caching explained.
- Switch models by task. Route simple queries to cheaper models and reserve expensive ones for complex tasks. This is model routing.
- Use batch APIs. Non-urgent workloads can use batch pricing at 50% off.
- Try cheaper providers. MiMO and DeepSeek offer significantly lower rates for comparable quality on many tasks.
Related
- Hidden LLM Costs - infrastructure, caching, retry, and observability costs beyond token pricing.
- How much does GPT-5 cost? - detailed GPT-5 pricing breakdown.
- How do LLM providers charge? - understanding token-based billing.
- Cheapest way to run AI code generation - cost comparison for coding tools.
- Provider arbitrage - finding the cheapest model that meets your quality bar.
Want this applied to your own LLM spend? FinOps LLM runs a free audit of your AI costs and shows where the savings are. Book free audit →