Opus 5.5 long-running agent cost analysis

Updated September 27, 2026 · first published September 27, 2026

A demo revealed Opus 5.5 working autonomously for ~25 hours to build a complete Minecraft clone: ~100K lines of game code, 26K lines of tests, 1,000 automated tests, 775 files, 783 generated images, 184 generated sounds, and an in-game guide. This is the first public glimpse of a long-horizon coding agent at scale. Let's model the cost.

Estimating token volume

No official token counts published. We'll model from output artifacts:

ArtifactEst. tokensNotes
100K lines game code~3M output~30 tokens/line avg
26K lines tests~0.8M output~30 tokens/line
775 files (structure, config)~0.5M outputBoilerplate + imports
783 images (prompts + metadata)~1.5M input~2K tokens/image prompt
184 sounds (prompts + metadata)~0.4M input~2K tokens/sound prompt
In-game guide / docs~0.3M outputMarkdown + code snippets
Context / reasoning (25 hrs)~10M inputTool calls, observations, planning
Total estimate~12.4M input / ~4.6M output

Cost at Opus 5.5 pricing

Token typeVolumeRateCost
Input (uncached, est. 30%)3.7M$4.00/MTok$14.80
Input (cached, est. 70%)8.7M$0.20/MTok$1.74
Output4.6M$20.00/MTok$92.00
Total~$108.54

~$109 for a 25-hour autonomous agent run producing a shippable game. That's ~$4.36/hour of agent compute.

Sensitivity: cache hit rate matters most

Cache hit rateInput costTotal costCost/hour
50%$24.80$131.60$5.26
70% (baseline)$16.54$108.54$4.34
90%$8.88$92.88$3.72

At 90% cache hit (achievable with good context management), the run drops to $93 total / $3.72/hr.

Comparison: human developer

A senior dev building this solo: 3-6 weeks (~120-240 hours) at $100-200/hr = $12K-$48K. The agent run at $109 is 100-400x cheaper — but only if the output quality passes your release gate.

What this means for your agent budget

  1. Long-horizon agents are affordable. Sub-$150 for a 25-hour complex build changes the ROI calculus for agent-first workflows.
  2. Cache is the lever. Invest in prompt caching infrastructure — every 10% cache hit improvement saves ~$1.60 on a run like this.
  3. Output tokens dominate. At $20/MTok, output is 85% of the bill. Models that are more concise (Anthropic claims 40% less verbose) directly cut cost.
  4. Monitor per-run, not per-month. Agent workloads are bursty. Track cost per task completion, not aggregate monthly spend.

Related reading

See Claude Opus 5.5 cost model, agent economics, prompt caching economics, and cost per successful task.

Related


Want this applied to your own LLM spend? FinOps LLM runs a free audit of your AI costs and shows where the savings are. Book free audit →

Back to research