Skip to content

OpenAI Decisions API cost model: input-only pricing and fast typed outputs

Updated October 7, 2026 · first published October 7, 2026

Quick answer: OpenAI’s Decisions API is a beta endpoint for turning shared text or image evidence into typed classifications, choices, or scores. With GPT-6 Luna, it charges $0.10 per million input tokens and no...

OpenAI’s Decisions API is a beta endpoint for turning shared text or image evidence into typed classifications, choices, or scores. With GPT-6 Luna, it charges $0.10 per million input tokens and no cache or output tokens. Treat its “10x faster” claim as a latency signal, then test quality and total task cost on your own workload.

What did OpenAI release on October 6, 2026?

OpenAI added the Decisions API in public beta on October 6, 2026. The dedicated Decisions guide says it returns typed answers about 10 times faster than the Responses API. At launch, GPT-6 Luna is the only supported model and requests use POST /v1/decisions. The speed comparison is OpenAI’s published claim, not a promise that every application will see the same end-to-end improvement.

A request carries one shared input and one or more questions. Each question can ask for a predicate (a probability that a condition is true), a choice from a supplied set, or a score against ordered levels. The endpoint returns those answers in typed form, with probabilities or confidence information. That makes it suited to bounded decisions such as routing a support ticket, flagging likely damage in a product photo, or scoring issue severity.

How does Decisions API billing work?

For the Decisions endpoint, OpenAI lists $0.10 per million input tokens with GPT-6 Luna. The guide says only input tokens are charged: cache-read, cache-write, and output-token charges do not apply. This differs from an ordinary GPT-6 Luna Responses request, where the standard short-context rates are $0.10 per million uncached input tokens, $0.01 cached input, $0.125 cache writes, and $0.50 output tokens. Use the dedicated Decisions pricing terms for this endpoint, not the full model rate card.

Here is a hypothetical estimate, before any applicable regional or long-context adjustment: if an average request uses 1,200 input tokens, its token charge is 1,200 ÷ 1,000,000 × $0.10 = $0.00012. One million such requests would use 1.2 billion input tokens and cost about $120. This is a planning example, not a quote; measure real token usage and confirm the live rate before setting a budget.

Because output tokens are not a billed category here, count the complete input carefully: evidence, system context included in the request, and every question’s instructions and choices. Ask independent questions about the same input in a single request where that makes sense. OpenAI documents that independent questions can share evidence; questions that depend on earlier answers should be sent in separate calls. Keep questions short and observable so the request does not spend tokens on ambiguous rubrics.

When should a team use Decisions instead of Responses?

Use Decisions when the result is inherently bounded: “Is this image visibly damaged?”, “Which of these three queues owns this ticket?”, or “How severe is this incident against these explicit levels?” It returns an answer the application can route or tally without asking a general-purpose model to write a paragraph and then parsing that prose.

Keep Responses for tasks that need explanation, arbitrary JSON schemas, generated text, tool calls, or a decision that requires a conversation with the model. OpenAI’s guide explicitly directs developers to Structured Outputs when they need a custom JSON object and to function calling when the model should request a tool call. A typed choice is not a substitute for a reasoned explanation or an executable action.

The right comparison is cost per correct business outcome, not tokens or latency in isolation. A quick classifier that creates costly false positives can increase review labor. A low-price endpoint can still lose if teams compensate with repeated prompts, manual review, or downstream correction. Track the Decision response’s usage and answer fields, sample outcomes against labeled examples, and include human review and downstream handling in the cost per successfully routed case.

How can FinOps validate the savings claim?

  1. Choose one bounded workload. Pick a high-volume classification or scoring step with a clear answer set and a measurable error cost.
  2. Build a labeled test set. Measure accuracy by category and inspect confidence or probability calibration. Set human-review thresholds from the cost of false positives and false negatives, as OpenAI recommends.
  3. Run a comparable pilot. Compare the current production path and Decisions on equivalent inputs. Record endpoint latency, tokens, retries, review rate, and the cost of corrected errors.
  4. Budget the full route. Add storage, application compute, reviewer time, and any later Responses or tool calls. Decisions token pricing does not make the rest of the workflow free.
  5. Keep an exit path. The endpoint is in public beta. Pin the endpoint and model in configuration, watch the changelog, and retest before broad rollout as behavior or availability changes.

OpenAI documents support for Zero Data Retention and HIPAA use for eligible customers, plus US and Europe (EEA and Switzerland) residency options with eligibility and agreement requirements. Confirm those controls against your organization’s actual configuration before putting regulated data into a pilot.

What should a buyer remember about Decisions API pricing?

The Decisions API offers a new price-and-latency shape for compact, typed decisions: at the published GPT-6 Luna rate, input is priced at $0.10 per million tokens and generated output tokens carry no charge. Its best candidate workloads have reusable evidence, clear answer types, and low-cost ways to catch uncertain results. Measure task quality and total operating cost before shifting production volume.

Which related research should teams read?

Related


Want this applied to your stack? Bring the provider bills, gateway logs, and top workflows; we will map the cost drivers and savings path. Book a free audit →

Back to finopsllm.com