Real-SWE cost per resolved issue
Updated September 27, 2026 · first published September 27, 2026
Real-SWE landed on 12 September 2026 and picked up 275 points and 158 comments on Hacker News. The premise is the part most cost models get wrong: benchmark suites are public, memorised, and unrepresentative of the code your agents actually touch. Real-SWE runs models against private, real-world enterprise repositories.
Why the cost number matters more than the score
A public benchmark gives you a pass rate. It does not give you a cost. Cost per resolved issue needs three things at once:
- Pass rate on the hard, private tasks.
- Tokens consumed per attempt, including the failures.
- Attempt count — did the model need a retry, a plan revision, a diff-and-rethink loop?
Multiply those and you get the only figure that maps to a line item. A model that scores 4 points below the leader but resolves issues in one attempt at a third of the tokens is cheaper per resolved issue, and on a real backlog that difference compounds faster than the score gap suggests.
Why public benchmarks overstate quality
Enterprise codebases have the properties that break public benchmark assumptions: long internal conventions, undocumented invariants, tests that encode historical accidents, and build systems nobody fully understands. A model that pattern-matches a public repo cannot pattern-match a 400-file monorepo with three years of accreted decisions. That is why a widely discussed result this month was a 27B open-weights creative-writing model performing at Fable 5 level at a reported 40x cheaper price — open weights start from a local-repository prior that API models lack.
Running the calculation
Take a real backend repo where 40 issues are in scope. Say the cheaper model resolves 22 in one attempt at 60K input / 8K output, and the premium model resolves 26 with a 1.4x retry multiplier at 90K input / 14K output. At Opus 5.5 rates ($4/$20, cache reads $0.20):
| Model | Resolved | Cost per issue | Total |
|---|---|---|---|
| Cheap tier | 22 | $0.32 | $7.04 | Premium tier | 26 | $0.61 | $15.86 |
Premium costs 2.3x the total for 18% more resolved issues. On a 40-issue backlog that is $9 more. The same trade at 4,000 issues/month is $880 — and still only buys 18% more throughput. The cheaper tier is not obviously wrong; you are trading resolution rate for volume, and the right answer depends on whether a missed issue costs more than a dollar.
What to instrument
- Cost per resolved issue, not per attempt. The denominator is merged PRs, not requests.
- Retry multiplier per model. A model at 1.4x is a third of its list price before you count anything else.
- Tokens per accepted diff, including plan and self-review tokens the diff never shows.
- Human review minutes. A 12% cheaper model that takes 20% longer to review is more expensive.
Bottom line
Public benchmarks rank capability on tasks that do not resemble your backlog. Real-SWE is an early attempt at the honest version. Until you run it on your own repos, treat every model comparison as a cost-per-resolved-issue hypothesis and price the hypothesis with your own token counts.
Related
- Real cost per task: frontier models
- Cost per successful task
- Heavy agentic coding cost: Claude vs Qwen
- Evals cost control
- Pricing an open-weight model
Want this applied to your own LLM spend? FinOps LLM runs a free audit of your AI costs and shows where the savings are. Book free audit →