Quick answer: As of 27 September 2026, GPT-6 Sol has exactly one independent long-context score: 84% on Artificial Analysis AA-LCR v1.1, tied with GPT-5.6 Sol. The test’s documents average about 100K tokens....

GPT-6 Sol long-context benchmarks: what has actually been measured, and what is GPT-5.6 Sol

Updated September 27, 2026 · first published September 27, 2026

As of 27 September 2026, GPT-6 Sol has exactly one independent long-context score: 84% on Artificial Analysis AA-LCR v1.1, tied with GPT-5.6 Sol. The test’s documents average about 100K tokens. No one has yet published how GPT-6 Sol holds up at 256K, 512K or 1M tokens. Most “Sol” long-context numbers online, including the 91.5% MRCR score, are for GPT-5.6 Sol. GPT-6 Sol came out on 22 September.

This matters for cost because long context is where GPT-6 prices jump. Above 272K input tokens the whole request is billed at long-context rates. Deciding how much context an agent keeps means weighing that price against quality data that doesn’t exist yet.

What long-context scores does GPT-6 Sol have?

One aggregate score, at around 100K tokens.

ModelAA-LCR v1.1Rank
GPT-6 Sol (max)84.0%7
GPT-5.6 Sol (max)84.0%7
GPT-6 Astra80.7%27
Kimi K3 (leader, 24 Sep 2026)88.7%1

AA-LCR has 100 hand-written questions over document sets that average 99,325 tokens. The answers have to be put together from several passages rather than found in one. It tests reasoning over a long input. It doesn’t show how accuracy changes as the input grows, because every question is about the same length.

Which long-context numbers are GPT-5.6 Sol, not GPT-6 Sol?

The MRCR scores. Launch coverage of GPT-6 Astra compared it with “Sol” on OpenAI’s MRCR v2 8-needle test. That Sol was GPT-5.6 Sol, and Vellum’s write-up confirms it.

MRCR v2, 8-needle256K–512K512K–1M
GPT-6 Astra100%96.3%
GPT-5.6 Sol91.5%73.8%
GPT-6 Solnot published

OrcaRouter’s GPT-6 Sol review makes the same point: GPT-5.6 Sol has a published long-context retrieval result, and GPT-6 Sol has nothing equivalent.

Where have we looked?

The places that would publish it, all checked on 27 September 2026:

What should teams do until per-length data exists?

Keep GPT-6 Sol requests under 272K input tokens and compact early. The evidence so far shows GPT-6 Sol matching GPT-5.6 Sol at about 100K. GPT-5.6 Sol held 91.5% up to 512K on MRCR but fell to 73.8% beyond that. It’s reasonable to expect GPT-6 Sol to be usable up to about 250K, but that is an inference, not a measurement. Above 272K you pay double for input on quality nobody has measured.

In Codex this is already the default. It gives GPT-6 Sol a 272K window and compacts at 244,800 tokens. For your own agents, alert on requests that cross 272K and treat any increase as something to test. Run your own retrieval check on your own documents before relying on context beyond 256K.

We will update this page when Context Arena or another independent evaluator publishes GPT-6 Sol results by context length.

Sources

Related


Want this applied to your own LLM spend? FinOps LLM runs a free audit of your AI costs and shows where the savings are. Book free audit →

Back to research