Found first: a primary source the press has not covered yet.
A new paper introduces DISCO, a distributed inference framework that separates long-context grounding from reasoning to prevent accuracy collapse at million-token scale. On the RULER-QA benchmark at one million tokens, DISCO achieves 78.4% accuracy where standard baselines collapse, while cutting inference costs by over 80% compared to full-context baseline approaches.
What the source says
The paper, by Guanzheng Chen, Viet Dac Lai, Subhojyoti Mukherjee, Branislav Kveton, Seunghyun Yoon, Franck Dernoncourt, Qizhe Xie, and Trung Bui (affiliations not listed in the preprint), frames the core problem as "context rot": the collapse of LLM reasoning quality as input length grows, attributed to a structural entanglement between grounding and reasoning in monolithic architectures. DISCO, inspired by Apache Spark, assigns parallel localized grounding to a fleet of Worker LLMs and gives a central Driver LLM, trained via reinforcement learning (GRPO), the job of planning extraction tasks and synthesizing evidence into a final answer. On LongBench v2, DISCO outperforms full-context models by up to 9.8 points and matches Gemini-3-Pro-Preview.
Why it matters
Million-token context windows are a standard frontier-model claim, but accuracy collapse at that scale is a documented deployment problem. Production systems that need reliable retrieval and reasoning over long documents, codebases, or transcripts cannot depend on advertised window sizes if reasoning degrades past a practical threshold. The over-80% cost reduction means the disaggregated approach is more accurate at long range and cheaper to run, which directly changes the economics of long-context inference in production. Matching Gemini-3-Pro-Preview from an open, distributed architecture is the result that warrants watching if the numbers hold under independent replication.