The stealth-model free-token playbook
Updated September 1, 2026 · first published September 1, 2026
Ox Alpha appeared on the routing aggregators as an unbranded endpoint, free, no rate card, no lab attached. Then it was confirmed as GLM-5.3 and the free window closed. GLM-5.3-Flash now lists at $0.15 per million input tokens and $0.50 per million output.
Treat this as a repeatable play, because it is one. It has run before under other names, and it will run again. What matters for your bill is what you did during the free window and what happens the moment it ends.
What the lab buys with free tokens
Three things, all of them cheaper than the alternatives. Blind evaluation: an unbranded endpoint gets judged on output rather than on the logo, which is the only honest read a lab can get on a frontier model. Distribution: the aggregators route real traffic to it immediately, no partnerships required. An anchor: by the time the name is revealed, developers have already decided the model is good, and the published price lands against that impression rather than against a spec sheet.
None of that is dishonest. It is just not free for you, and the invoice arrives later than the tokens do.
The number that matters when the window closes
Sixteen trillion input tokens flowed through the free endpoint. At the now-published $0.15 per million, that is roughly $2.4 million of inference — which is both what the lab spent on the campaign and, more usefully, the size of the step if that same traffic simply stayed where it was.
Run that arithmetic on your own share. Take the volume you sent through the anonymous endpoint during the window, multiply by the published rate, and you have your monthly step change. Teams that ran an entire evaluation cycle on free tokens are the ones most likely to be surprised, because the workload that proved the model was also the workload that cost nothing.
Three rules for the next one
Never depend on an endpoint whose owner you do not know. Anonymous means no rate card, no deprecation policy, no support path, no data-handling commitment you can point at. It is fine for evaluation. It is not fine for a production path with no fallback.
Price every free evaluation at a hypothetical rate. Pick a plausible per-million figure before you start — the closest comparable model is fine — and carry the shadow cost in your own telemetry. Then a reveal is a repricing, not a discovery.
Assume the free window is the shortest part of the relationship. Six days here. What you should be measuring during it is not just quality but switching cost: how much routing, prompt tuning and evaluation work you would throw away if the published price came in high.
Where it leaves the price
$0.15 / $0.50 is aggressive against the closed frontier tier and roughly in line with the open-weight hosted market. That is the point of the whole exercise: the reveal is the pricing announcement, and it is designed to land while the model is already in your router. Whether it wins your traffic should be decided by your own cost-per-request numbers on your own workload, measured after the free window closed — not by how good it felt when the tokens were free.
Related
Related
Want this applied to your own LLM spend? FinOps LLM runs a free audit of your AI costs and shows where the savings are. Book free audit →