OpenAI Ultra Fast API pricing and speed tiers
Updated September 27, 2026 · first published September 27, 2026
OpenAI's response API playground now shows a speed selector with three options: standard, fast, and ultra fast. This isn't just a UI tweak — it's a new pricing dimension. Ultra fast mode, powered by Cerebras inference, was previewed last month for select customers where GPT-5 Soul hit 750 tokens per second — up to 14x faster than standard. The Dev Day leak suggests OpenAI is preparing to expand ultra fast access to a much larger developer group.
The three tiers
| Tier | Relative speed | Use case | Pricing signal |
|---|---|---|---|
| Standard | 1x (baseline) | Batch, async, non-urgent | Base rate |
| Fast | ~3-5x | Interactive, user-facing | Likely 2-3x base |
| Ultra fast | ~10-14x | Real-time agents, trading, live assist | Likely 5-10x base |
Exact multipliers aren't public yet. The leaked documentation shows ultra fast as an "access controlled service tier" — meaning you'll need approval or a specific plan (the rumored $500 ChatGPT Pro Max) to unlock it.
What this does to your cost model
Speed tiers turn latency into a first-class cost lever. Today you optimize by model selection (GPT-5 mini vs GPT-5 vs GPT-5 Soul). Tomorrow you'll optimize by speed tier within the same model.
For a 10K token request/response pair at standard GPT-5 Soul pricing (~$5/$15 per million):
- Standard: ~$0.20 per call
- Fast (est. 3x): ~$0.60 per call
- Ultra fast (est. 8x): ~$1.60 per call
The breakeven depends entirely on what latency buys you. A support agent that resolves in one turn instead of three because ultra fast kept the user engaged? That's a win. A batch enrichment job? Standard every time.
How to prepare
- Instrument latency per feature. You can't decide if fast is worth it without knowing your current p50/p95 and the business value of each millisecond.
- Tag requests by urgency. Build a router that sends batch work to standard, chat to fast, and real-time agents to ultra fast — automatically.
- Model the Pro Max unlock. If the $500/mo plan includes ultra fast quota, calculate your break-even volume vs pay-per-call at the ultra fast multiplier.
Related reading
See model routing for the pattern this extends, and LLM cost per user benchmarks to benchmark your own latency-to-value ratio.
Related
- Model routing
- LLM cost per user benchmarks
- ChatGPT Pro Max: $500 plan economics
- How to cap inference costs
Want this applied to your own LLM spend? FinOps LLM runs a free audit of your AI costs and shows where the savings are. Book free audit →