LLM API Pricing Tracker

Updated 12 July 2026 · first published 4 July 2026

Track per-token pricing across all major LLM providers. Input, output, cache read/write, context windows, and batch discounts in one normalised table. Updated monthly with verifiable provider data.

12+Providers tracked
25+Models listed
6Price columns
Jul 2026Last updated

Methodology

Collection

Pricing data collected from official provider pricing pages and API documentation.

Normalisation

All prices normalised to USD per 1 million tokens for direct comparison.

Cache pricing

Cache Read/Write columns show prompt caching discounts where providers support them.

Display policy

Official values and estimated values are clearly separated. Estimates marked in gray.

Update Feed

Jul 4, 2026

LLM API Pricing Tracker launched

Published per-token pricing comparison across 12+ major LLM providers as a FinOps LLM research article.

Per-Token Pricing Table

Provider Model Input $/1M Output $/1M Cache Read $/1M Cache Write $/1M Context Window Batch Discount
OpenAI GPT-5.6-Sol $5.00 $30.00 $0.50 $6.25 128K 50%
OpenAI GPT-5.6-Terra $2.50 $15.00 $0.25 $3.13 128K 50%
OpenAI GPT-5.6-Luna $1.00 $6.00 $0.10 $1.25 128K 50%
OpenAI GPT-5.5 $5.00 $30.00 $0.50 $6.25 128K 50%
OpenAI GPT-5.4 $2.50 $15.00 $0.25 $3.13 128K 50%
OpenAI o3 $2.00 $8.00 $0.50 $2.00 200K 50%
OpenAI o4-mini $1.10 $4.40 $0.28 $1.10 200K 50%
Anthropic Claude Fable 5 $10.00 $50.00 $1.00 $12.50 200K 50%
Anthropic Claude Opus 4.8 $5.00 $25.00 $0.50 $6.25 200K 50%
Anthropic Claude Sonnet 5* $2.00 $10.00 $0.20 $2.50 200K 50%
Anthropic Claude Haiku 4.5 $1.00 $5.00 $0.10 $1.25 200K 50%
Google Gemini 3.5 Flash $1.50 $9.00 $0.15 $1.88 1M 50%
Google Gemini 3.1 Pro Preview $2.00 $12.00 $0.20 $2.50 1M 50%
Google Gemini 2.5 Pro $1.25 $5.00 $0.31 $4.50 1M 50%
Google Gemini 2.5 Flash $0.15 $0.60 $0.04 $0.15 1M 50%
xAI Grok 4.5 $2.00 $6.00 $0.50 $1.50 1M 50%
xAI Grok 4.3 $1.25 $2.50 - - 1M 50%
Meta Llama 4 Maverick (DeepInfra) $0.20 $0.80 - - 1M -
Meta Llama 4 Scout (Groq) $0.11 $0.34 - - 5M -
Meta Llama 3.3 70B (Groq) $0.59 $0.79 - - 8M -
Mistral Mistral Large 2/3 $2.00 $6.00 - - 128K -
Mistral Mistral Medium 3.5 $1.50 $7.50 - - 128K -
Mistral Mistral Small 3 $0.10 $0.30 - - 128K -
Mistral Codestral $0.30 $0.90 - - 128K -
DeepSeek DeepSeek V4 Flash $0.14 $0.28 $0.0028 - 128K -
DeepSeek DeepSeek V4 Pro $0.435 $0.87 $0.0036 - 128K -
Alibaba Qwen Qwen 3.7-Max $2.50 $7.50 - - 131K -
Moonshot Kimi Kimi K2.6 $0.95 $4.00 $0.16 - 200K -
Xiaomi MiMO
Sign up - both get $2 ↗
MiMo-V2-Pro $0.50 $2.00 - - 128K -
Xiaomi MiMO
Sign up - both get $2 ↗
MiMo-V2-Omni $0.40 $1.60 - - 128K -
Zhipu AI GLM-5 $1.00 $4.00 - - 128K -
Zhipu AI GLM-5-Turbo $0.30 $1.20 - - 128K -
MiniMax M2.7 $0.40 $1.60 - - 128K -
StepFun step-3.5-flash $0.30 $1.20 - - 128K -

* Claude Sonnet 5: introductory pricing of $2/$10 per 1M tokens in effect through August 31, 2026; standard pricing $3/$15 begins September 1, 2026.

Values not explicitly provided by providers are shown in gray.

Meta Llama prices reflect hosted API pricing (Fireworks, Together, etc.), not self-hosted cost. Self-hosting economics are covered in our Open-Source vs Closed-Source Cost Comparison.

Context windows are maximum supported; actual availability may vary by tier.

Cache Read/Write pricing reflects provider-specific prompt caching where available. "-" means the provider does not offer caching or pricing is not published.

Batch Discount shows the approximate savings when using batch/async APIs where available.

Some links may include referral codes.

Prices can change frequently. Always verify the latest information on official provider pages.

Practical Guide

How to use this pricing data for cost planning

1) Calculate your token volume first

Before comparing prices, estimate your monthly input and output token volumes. Output tokens are typically 3-5x more expensive than input, so reducing output length has the biggest impact on cost.

2) Factor in caching savings

If your workload has repeated system prompts or context (RAG, multi-turn conversations), prompt caching can reduce input costs by 50-90%. Check the Cache Read column for applicable discounts.

3) Use batch APIs for non-urgent work

Most major providers offer 50% discounts on batch/async API calls. If your workload tolerates latency (evaluations, data processing, content generation), batch pricing cuts costs in half.

4) Compare total cost, not just per-token price

A cheap model that needs 3x more tokens to match quality may cost more overall. Factor in model quality and task-specific performance alongside raw pricing.

Next step: Compare your spend across providers—see Azure OpenAI vs Direct OpenAI Cost for a breakdown of managed vs direct pricing models.


Want this applied to your own LLM spend? FinOps LLM runs a free audit of your AI costs and shows where the savings are. Book free audit →

Back to research