AI Model Release Tracker.
Every new model the major labs have put into production in the last 45 days, with the price you would actually pay, and every model they have announced they are switching off. Two lists, because the second one is the one that reaches your invoice unannounced.
New in the last 45 days 40 models
A model appears here the day it first became purchasable through a public router. It usually trails the lab's own announcement by a few days. Prices are USD per 1M tokens, straight from the router's rate card.
| First available | Model | Lab | Input / 1M | Output / 1M | Context |
|---|---|---|---|---|---|
| z-ai/glm-5.3-prime | Z.ai | $2.8 | $8.8 | 1M | |
| upstage/solar-mini4 | Upstage | $0.05 | $0.2 | 524K | |
| qwen/qwen3.8-max-prime | Alibaba | $4 | $12 | 1M | |
| aion-labs/aion-3.5-mini | Aion Labs | $0.7 | $1.4 | 262K | |
| aion-labs/aion-3.5 | Aion Labs | $3 | $6 | 262K | |
| openai/gpt-6-sol-pro | OpenAI | $2 | $10 | 1.05M | |
| openai/gpt-6-sol | OpenAI | $2 | $10 | 1.05M | |
| openai/gpt-6-luna-pro | OpenAI | $0.1 | $0.5 | 1.05M | |
| openai/gpt-6-luna | OpenAI | $0.1 | $0.5 | 1.05M | |
| cohere/command-a-plus | Cohere | $0.3 | $1.5 | 192K | |
| anthropic/claude-opus-5.5 | Anthropic | $4 | $20 | 1M | |
| xiaomi/mimo-v2.6-pro-ultraspeed | Xiaomi | $4.35 | $8.7 | 1.05M | |
| xiaomi/mimo-v2.6-pro | Xiaomi | $0.435 | $0.87 | 1.05M | |
| xiaomi/mimo-v2.6-flash | Xiaomi | $0.14 | $0.28 | 1.05M | |
| x-ai/grok-4.7 | xAI | $1.6 | $4.8 | 500K | |
| qwen/qwen3.8-omni-flash | Alibaba | $0.15 | $0.47 | 1M | |
| z-ai/glm-5.3-flashx | Z.ai | $0.37 | $1.25 | 1.05M | |
| sakana/fugu-ultra-v2 | Sakana | $5 | $30 | 1M | |
| sakana/fugu-max | Sakana | $2 | $6 | 1M | |
| inclusionai/ling-3.0-flash-vl | InclusionAI | $0.021 | $0.0616 | 262K | |
| deepseek/deepseek-v4.1-flash | DeepSeek | $0.035 | $0.29 | 1.05M | |
| openai/gpt-6-astra-pro | OpenAI | $10 | $50 | 1.05M | |
| openai/gpt-6-astra | OpenAI | $10 | $50 | 1.05M | |
| qwen/qwen3.8-max-0902 | Alibaba | $2 | $6 | 1M | |
| meta/muse-spark-1.3-contributor | Meta | $0.1 | $0.2 | 1.05M | |
| meta/muse-spark-1.3 | Meta | $1.25 | $4.25 | 1.05M | |
| google/gemini-3.8-flash | $0.75 | $3.75 | 1.05M | ||
| anthropic/claude-fable-5.1 | Anthropic | $10 | $50 | 1M | |
| tencent/hy4-preview | Tencent | $0.834 | $2.501 | 1.05M | |
| inclusionai/ling-3.0-flash-fin | InclusionAI | $0.06 | $0.18 | 262K | |
| z-ai/glm-5.3-flash | Z.ai | $0.045 | $0.14 | 1.31M | |
| qwen/qwen3.8-flash | Alibaba | $0.15 | $0.47 | 1M | |
| meta/muse-spark-1.2-contributor | Meta | $0.1 | $0.2 | 1.05M | |
| deepseek/deepseek-v4-flash-vision-exp | DeepSeek | $0.2156 | $0.6468 | 1.05M | |
| tencent/hy-mt2-30b-a3b | Tencent | $0.074 | $0.295 | 8K | |
| tencent/hy-mt2-1.8b | Tencent | $0.044 | $0.177 | 8K | |
| tencent/hy-mt2-7b | Tencent | $0.074 | $0.295 | 8K | |
| z-ai/glm-5.3 | Z.ai | $1.4 | $4.4 | 1.31M | |
| qwen/qwen3.8-27b | Alibaba | $0.42 | $3 | 1M | |
| google/gemini-3.7-flash | $0.75 | $3.75 | 1.05M |
No models from that lab in this window.
Announced retirements 26 models
Dates a lab or router has published for switching a model off. A model on this list is not deprecated-yet: on the date shown, calls to it start failing. Price is included so the replacement can be costed in the same column.
| Switch-off date | In | Model | Lab | Input / 1M | Output / 1M |
|---|---|---|---|---|---|
| 1d | deepseek/deepseek-v3.2 | DeepSeek | $0.269 | $0.4 | |
| 1d | deepseek/deepseek-v3.2-exp | DeepSeek | $0.27 | $0.41 | |
| 1d | deepseek/deepseek-v3.1-terminus | DeepSeek | $0.27 | $1 | |
| 1d | deepseek/deepseek-r1-distill-llama-70b | DeepSeek | $0.8 | $0.8 | |
| 11d | minimax/minimax-m2.1 | MiniMax | $0.3 | $1.2 | |
| 11d | baidu/ernie-4.5-vl-424b-a47b | Baidu | $0.42 | $1.25 | |
| 12d | qwen/qwen3.6-max-preview | Alibaba | $1.027 | $6.162 | |
| 12d | qwen/qwen3-max-thinking | Alibaba | $0.78 | $3.9 | |
| 12d | qwen/qwen3-vl-32b-instruct | Alibaba | $0.104 | $0.416 | |
| 12d | qwen/qwen3-vl-8b-thinking | Alibaba | $0.18 | $2.1 | |
| 12d | qwen/qwen3-vl-8b-instruct | Alibaba | $0.117 | $0.455 | |
| 12d | qwen/qwen3-vl-30b-a3b-thinking | Alibaba | $0.2 | $2.4 | |
| 12d | qwen/qwen3-vl-235b-a22b-thinking | Alibaba | $0.4 | $4 | |
| 12d | qwen/qwen3-max | Alibaba | $0.78 | $3.9 | |
| 12d | qwen/qwen3-coder-plus | Alibaba | $0.65 | $3.25 | |
| 12d | qwen/qwen-plus-2025-07-28 | Alibaba | $0.26 | $0.78 | |
| 12d | qwen/qwen3-30b-a3b-thinking-2507 | Alibaba | $0.2 | $2.4 | |
| 12d | qwen/qwen3-235b-a22b-thinking-2507 | Alibaba | $0.23 | $2.3 | |
| 12d | qwen/qwen3-8b | Alibaba | $0.117 | $0.455 | |
| 12d | qwen/qwen3-235b-a22b | Alibaba | $0.455 | $1.82 | |
| 23d | google/gemini-2.5-flash-lite | $0.1 | $0.4 | ||
| 23d | google/gemini-2.5-flash | $0.3 | $2.5 | ||
| 23d | google/gemini-2.5-pro | $1.25 | $10 | ||
| 45d | bytedance-seed/seed-2.0-code | ByteDance | $0.5 | $3 | |
| 45d | bytedance-seed/seed-1.6-flash | ByteDance | $0.075 | $0.3 | |
| 45d | bytedance-seed/seed-1.6 | ByteDance | $0.25 | $2 |
Nothing announced for that lab.
Why the second table is the one that costs you
A new model is a decision. A retirement is an event, and it arrives whether or not anyone on your team was watching. Three things break quietly on a switch-off date:
- Routing rules. A load balancer or fallback chain that still references the retired id returns errors for that slice of traffic. If the fallback was never load-tested, the retry path becomes the primary path and your latency bill follows.
- Evaluation baselines. Any score, prompt or guardrail tuned against the old model's behaviour is now measuring a system that no longer exists. Regression alerts that fire afterwards are noise, and the ones that don't fire are worse.
- Prompts tuned to a context window. A replacement with a smaller ceiling truncates mid-document. The failure is a plausible answer built on half the input, which is far more expensive to catch than an error.
The cheapest time to price a replacement is before the switch-off date, while the old model still runs and you can measure both side by side on real traffic. That is the whole argument for keeping this list rather than reading about it after the fact.
How we build this
Both tables come from one public model feed that lists every served model with the timestamp it first appeared and, where a switch-off has been announced, the date it ends. The filter keeps first-party models only: third-party routers, quantised mirrors and anonymous drops all carry timestamps but are not a lab shipping anything, and :batch or :free entries are the same weights behind a different endpoint, so they are dropped as duplicates of a row already here. Named tiers such as -pro are kept, because they are separate price points and choosing between them is the reason to be on this page.
Two dates matter and they are not the same date. 2026-09-27 is when this page's table was generated. The live feed behind it is at most six hours old and is checked when you load this page; a new model found in the live check is added to the top of the table and marked. Prices come from the router's rate card for that model, which is what you would be billed, and a model whose lab has not published a price shows a dash rather than a number.
What this page cannot tell you is what a lab said, or on what date they said it. For the announcement itself — the reasoning behind a price, the scope of a deprecation, the caveat on a benchmark — read the lab's own newsroom and social accounts. This is the operational layer underneath those posts: what is purchasable, what it costs, and what is about to stop working.
Related
- LLM API pricing tracker — full rate cards for every model listed here
- Trending AI models — what the current releases actually score on
- Silent model downgrades — when quality moves without a price change
- Committed spend and procurement — locking in before a repricing
- LLM cost calculator — cost a workload on a specific model