paper web signal

DISCO paper: split grounding from reasoning to fix context rot

TL;DR

  • DISCO reports 78.4% accuracy on RULER-QA at 1 million tokens using a Qwen3-8B worker configuration, versus 10.9% for a standard RAG baseline.
  • On LongBench v2's Long subset, the paper claims 48.7% versus 38.9% for a full-context baseline, a 9.8-point gap that widens to 50.9% with RL training.
  • The authors report matching Gemini-3-Pro-Preview on their benchmarks while cutting inference cost from $126.7 to $20.4, a reduction of over 80%.

A new preprint on arXiv, DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation, reports 78.4% accuracy on RULER-QA at one million tokens using Qwen3-8B workers, against 10.9% for a standard RAG baseline. The authors call the failure mode they are attacking "context rot": reasoning quality collapsing as inputs grow even inside advertised million-token windows.

The system splits the job in two. A fleet of Worker LLMs does parallel, localized grounding over partitioned context; a single Driver LLM, trained with reinforcement learning, orchestrates the reasoning. The paper compares the design to distributed computing frameworks, writing that DISCO "partitions long context across a fleet of Worker LLMs dedicated exclusively to parallel, localized grounding" and, by doing so, "effectively eliminates context rot, maintaining robust fidelity up to one million tokens where monolithic architectures collapse."

On LongBench v2's Long subset, the authors report 48.7% for DISCO with a Qwen3-14B worker versus 38.9% for a full-context baseline, a 9.8-point gap that rises to 50.9% once the Driver is trained with RL. On ∞Bench they report a 63.68% average across En.MC and En.QA. The cost claim is the sharpest: the paper says DISCO matches Gemini-3-Pro-Preview on their benchmarks while cutting inference cost from $126.7 to $20.4, a reduction of over 80%.

The results are on long-context benchmarks the community already knows are partly synthetic, and the abstract does not quantify orchestration latency or Worker disagreement rates. The Driver-model identity behind the $20.4 figure is also not spelled out in the passages surfaced here.