OpenAI Ultra Fast API pricing and speed tiers

Updated September 27, 2026 · first published September 27, 2026

OpenAI's response API playground now shows a speed selector with three options: standard, fast, and ultra fast. This isn't just a UI tweak — it's a new pricing dimension. Ultra fast mode, powered by Cerebras inference, was previewed last month for select customers where GPT-5 Soul hit 750 tokens per second — up to 14x faster than standard. The Dev Day leak suggests OpenAI is preparing to expand ultra fast access to a much larger developer group.

The three tiers

TierRelative speedUse casePricing signal
Standard1x (baseline)Batch, async, non-urgentBase rate
Fast~3-5xInteractive, user-facingLikely 2-3x base
Ultra fast~10-14xReal-time agents, trading, live assistLikely 5-10x base

Exact multipliers aren't public yet. The leaked documentation shows ultra fast as an "access controlled service tier" — meaning you'll need approval or a specific plan (the rumored $500 ChatGPT Pro Max) to unlock it.

What this does to your cost model

Speed tiers turn latency into a first-class cost lever. Today you optimize by model selection (GPT-5 mini vs GPT-5 vs GPT-5 Soul). Tomorrow you'll optimize by speed tier within the same model.

For a 10K token request/response pair at standard GPT-5 Soul pricing (~$5/$15 per million):

The breakeven depends entirely on what latency buys you. A support agent that resolves in one turn instead of three because ultra fast kept the user engaged? That's a win. A batch enrichment job? Standard every time.

How to prepare

  1. Instrument latency per feature. You can't decide if fast is worth it without knowing your current p50/p95 and the business value of each millisecond.
  2. Tag requests by urgency. Build a router that sends batch work to standard, chat to fast, and real-time agents to ultra fast — automatically.
  3. Model the Pro Max unlock. If the $500/mo plan includes ultra fast quota, calculate your break-even volume vs pay-per-call at the ultra fast multiplier.

Related reading

See model routing for the pattern this extends, and LLM cost per user benchmarks to benchmark your own latency-to-value ratio.

Related


Want this applied to your own LLM spend? FinOps LLM runs a free audit of your AI costs and shows where the savings are. Book free audit →

Back to research