Real-SWE cost per resolved issue

Updated September 27, 2026 · first published September 27, 2026

Real-SWE landed on 12 September 2026 and picked up 275 points and 158 comments on Hacker News. The premise is the part most cost models get wrong: benchmark suites are public, memorised, and unrepresentative of the code your agents actually touch. Real-SWE runs models against private, real-world enterprise repositories.

Why the cost number matters more than the score

A public benchmark gives you a pass rate. It does not give you a cost. Cost per resolved issue needs three things at once:

Multiply those and you get the only figure that maps to a line item. A model that scores 4 points below the leader but resolves issues in one attempt at a third of the tokens is cheaper per resolved issue, and on a real backlog that difference compounds faster than the score gap suggests.

Why public benchmarks overstate quality

Enterprise codebases have the properties that break public benchmark assumptions: long internal conventions, undocumented invariants, tests that encode historical accidents, and build systems nobody fully understands. A model that pattern-matches a public repo cannot pattern-match a 400-file monorepo with three years of accreted decisions. That is why a widely discussed result this month was a 27B open-weights creative-writing model performing at Fable 5 level at a reported 40x cheaper price — open weights start from a local-repository prior that API models lack.

Running the calculation

Take a real backend repo where 40 issues are in scope. Say the cheaper model resolves 22 in one attempt at 60K input / 8K output, and the premium model resolves 26 with a 1.4x retry multiplier at 90K input / 14K output. At Opus 5.5 rates ($4/$20, cache reads $0.20):

ModelResolvedCost per issueTotal
Cheap tier22$0.32$7.04
Premium tier26$0.61$15.86

Premium costs 2.3x the total for 18% more resolved issues. On a 40-issue backlog that is $9 more. The same trade at 4,000 issues/month is $880 — and still only buys 18% more throughput. The cheaper tier is not obviously wrong; you are trading resolution rate for volume, and the right answer depends on whether a missed issue costs more than a dollar.

What to instrument

  1. Cost per resolved issue, not per attempt. The denominator is merged PRs, not requests.
  2. Retry multiplier per model. A model at 1.4x is a third of its list price before you count anything else.
  3. Tokens per accepted diff, including plan and self-review tokens the diff never shows.
  4. Human review minutes. A 12% cheaper model that takes 20% longer to review is more expensive.

Bottom line

Public benchmarks rank capability on tasks that do not resemble your backlog. Real-SWE is an early attempt at the honest version. Until you run it on your own repos, treat every model comparison as a cost-per-resolved-issue hypothesis and price the hypothesis with your own token counts.

Related


Want this applied to your own LLM spend? FinOps LLM runs a free audit of your AI costs and shows where the savings are. Book free audit →

Back to research