OpenAI O / AON always-on agent cost model

Updated September 27, 2026 · first published September 27, 2026

Leaked references reveal OpenAI's "O" (codename AON) — an always-on assistant designed for long-horizon tasks. Unlike chat models that respond and stop, O gets its own cloud environment, can set up execution contexts, write code, research, and operate continuously for hours, days, or potentially weeks. It may involve multiple agents communicating on the same task. Expected reveal: OpenAI Dev Day, September 29, 2026.

What changes for cost modeling

Traditional LLM cost: per-request tokens × rate. O/AON cost: time × allocated compute + token throughput.

DimensionChat/Completion APIO / AON (projected)
Billing unitPer 1M tokensPer hour (compute) + per 1M tokens
Idle cost$0>$0 (environment reservation)
Max durationMinutes (timeout)Days/weeks
ParallelismSequential requestsMulti-agent swarms
StatePassed in contextPersistent environment

Projected cost model (speculative, based on leak patterns)

OpenAI hasn't published pricing. Two likely structures:

Option A: Compute-hour + tokens (like Codespaces)

A 24-hour agent run with 5M input / 2M output tokens:

ComponentCost (low)Cost (high)
24h environment ($1/hr)$24$48
5M input @ $5/MTok$25$25
2M output @ $15/MTok$30$30
Total$79$103

Option B: Flat subscription (Pro Max bundle)

Cost optimization strategies for always-on agents

  1. Hibernate, don't terminate. If the environment persists, pause the agent loop during idle periods instead of spinning down — cold start costs more than idle reservation.
  2. Batch subtask tokens. Accumulate subtask results and write context in larger chunks to improve cache hit rates on the persistent environment.
  3. Route subtasks to cheaper models. O/AON orchestrates; delegate coding to Sonnet 5.5, research to Flash models, only escalate to premium for critical decisions.
  4. Set hard time/money budgets per objective. "Solve this bug, max $10 / 2 hours" — enforce via orchestrator, not hope.

When to use O/AON vs Opus 5.5 long-running

Related reading

See ChatGPT Pro Max economics, OpenAI Ultra Fast API pricing, Opus 5.5 25-hour agent cost, and agent economics.

Related


Want this applied to your own LLM spend? FinOps LLM runs a free audit of your AI costs and shows where the savings are. Book free audit →

Back to research