Skip to content

OpenAI's Math Release and AI Research Costs

Updated October 8, 2026 · first published October 8, 2026

Quick answer: OpenAI's October 6 mathematics release is a useful cost-governance case, not a public price benchmark. The company says the average result used compute equivalent to roughly three hours of ChatGPT Pro...

OpenAI's October 6 mathematics release is a useful cost-governance case, not a public price benchmark. The company says the average result used compute equivalent to roughly three hours of ChatGPT Pro thinking, but it does not publish a per-result API price for the internal model. Finance teams should track experiment spend, review effort, and verified results separately instead of converting that compute equivalent into an invented dollar figure.

What OpenAI disclosed

OpenAI published 722 mathematical manuscripts grouped into 372 families, alongside some Lean formalizations and details about its research process. It says the average result used compute equivalent to about three hours of ChatGPT Pro thinking. The repository also says the collection includes results at different stages of verification and that some unformalized results could contain issues. See OpenAI's release note and the research repository for the source material.

These disclosures describe research output and an approximate compute equivalent. They do not provide the internal model's API rate, exact per-manuscript metering, a conversion from Pro compute to dollars, or an independently verified success rate. A finance model should keep those unknowns visible.

Do not treat a compute equivalent as a price

A consumer plan's time-equivalent measure is not the same as metered API usage. It may be a useful way to communicate scale, but the public disclosure does not specify a token count, hardware allocation, marginal serving cost, or billable unit. Multiplying three hours by a guessed hourly GPU rate would imply precision the source does not support.

For a deployable workload, use the provider's actual billing records or your own instrumented inference ledger. For an unreleased internal model, record the disclosed compute equivalent as a separate research-unit measure and label any internal transfer price as an assumption. Do not place that assumption beside observed invoice costs without a clear distinction.

Measure the work that survives review

For research programs, a raw output count rewards volume even when results need correction or cannot be reused. Pair spend with an acceptance funnel:

Report compute or API spend per attempted problem and per verified result. Include reviewer time, orchestration, storage, and follow-up runs if the decision is about total program economics. Keep failed attempts and revisions in the denominator; excluding them makes the unit cost look artificially low. For broader definitions, see cost per successful task and reasoning-model cost tracking.

Put budget gates around research runs

Before a pilot, set a project ceiling, a maximum spend or compute allowance per problem, and a review checkpoint before the next batch. Stop or pause the run when one of three conditions occurs: the work-unit budget is exhausted, the rate of independently accepted results falls below the agreed floor, or reviewers cannot validate outputs fast enough to keep up.

At each checkpoint, compare the marginal cost of the next batch with the value of the results already verified. Keep separate owners for the technical acceptance decision and the budget exception. A large batch of plausible-looking outputs is not, on its own, evidence to raise the limit.

What finance can conclude today

The announcement makes high-compute scientific work a concrete planning scenario. It does not establish a commercial price, enterprise ROI, or a production workload's cost per verified answer. For now, use the release to define the ledger and review gates you would require before approving a similar program. When actual usage and verified outcomes are available, replace the compute-equivalent proxy with measured costs. A companion guide explains how to allocate those costs to research projects and results: AI research cost allocation.

FAQ

How much compute did an average OpenAI math result use?

OpenAI says it used the compute equivalent of roughly three hours of ChatGPT Pro thinking. This is not a published dollar cost or per-result API price.

Can finance estimate the dollar cost from that number?

Not from the public disclosure alone. It does not state the internal model's API price, tokens per result, or a conversion rate from the Pro compute equivalent to dollars.

What should FinOps measure for AI research?

Track actual compute or API charges, attempts, independently verified results, reviewer effort, and reusable outputs. Report cost per verified result as well as total spend, and label estimates separately from invoice data.

Related

Related


Want this applied to your stack? Bring the provider bills, gateway logs, and top workflows; we will map the cost drivers and savings path. Book a free audit →

Back to finopsllm.com