OpenAI O / AON always-on agent cost model
Updated September 27, 2026 · first published September 27, 2026
Leaked references reveal OpenAI's "O" (codename AON) — an always-on assistant designed for long-horizon tasks. Unlike chat models that respond and stop, O gets its own cloud environment, can set up execution contexts, write code, research, and operate continuously for hours, days, or potentially weeks. It may involve multiple agents communicating on the same task. Expected reveal: OpenAI Dev Day, September 29, 2026.
What changes for cost modeling
Traditional LLM cost: per-request tokens × rate. O/AON cost: time × allocated compute + token throughput.
| Dimension | Chat/Completion API | O / AON (projected) |
|---|---|---|
| Billing unit | Per 1M tokens | Per hour (compute) + per 1M tokens |
| Idle cost | $0 | >$0 (environment reservation) |
| Max duration | Minutes (timeout) | Days/weeks |
| Parallelism | Sequential requests | Multi-agent swarms |
| State | Passed in context | Persistent environment |
Projected cost model (speculative, based on leak patterns)
OpenAI hasn't published pricing. Two likely structures:
Option A: Compute-hour + tokens (like Codespaces)
- Environment reservation: ~$0.50-2.00/hour (CPU + memory + storage)
- Token throughput: standard GPT-5 rates on top
- Ultra fast tier: 5-10x token multiplier when enabled
A 24-hour agent run with 5M input / 2M output tokens:
| Component | Cost (low) | Cost (high) |
|---|---|---|
| 24h environment ($1/hr) | $24 | $48 |
| 5M input @ $5/MTok | $25 | $25 |
| 2M output @ $15/MTok | $30 | $30 |
| Total | $79 | $103 |
Option B: Flat subscription (Pro Max bundle)
- $500/mo ChatGPT Pro Max includes O/AON access + ultra fast quota
- Effective hourly: $500 / 720h = $0.69/hr (if 24/7)
- Break-even vs Option A: ~15 hours/day sustained usage
Cost optimization strategies for always-on agents
- Hibernate, don't terminate. If the environment persists, pause the agent loop during idle periods instead of spinning down — cold start costs more than idle reservation.
- Batch subtask tokens. Accumulate subtask results and write context in larger chunks to improve cache hit rates on the persistent environment.
- Route subtasks to cheaper models. O/AON orchestrates; delegate coding to Sonnet 5.5, research to Flash models, only escalate to premium for critical decisions.
- Set hard time/money budgets per objective. "Solve this bug, max $10 / 2 hours" — enforce via orchestrator, not hope.
When to use O/AON vs Opus 5.5 long-running
- O/AON: Multi-day research, continuous monitoring, multi-agent coordination, tasks needing persistent environment state
- Opus 5.5 agent loop: Single-objective coding tasks (hours, not days), well-defined deliverables, one-shot execution
Related reading
See ChatGPT Pro Max economics, OpenAI Ultra Fast API pricing, Opus 5.5 25-hour agent cost, and agent economics.
Related
Want this applied to your own LLM spend? FinOps LLM runs a free audit of your AI costs and shows where the savings are. Book free audit →