paper web signal

Tiny Rank-8 LoRA Lifts Qwen3-8B Chain Accuracy From 15.5% to 99%

TL;DR

  • Qwen3-8B jumped from 15.5% to 99% exact accuracy on 24-line reference chains with a rank-8 LoRA at a single early layer.
  • Thirteen pretrained base models reliably follow only 1.4-3.6 lines of in-context reference chains in their default configuration.
  • Ouro-1.4B reached 60-line chains after four loops and at least 160 after eight; task-specific LoRAs also lifted MuSiQue scores.

Qwen3-8B's exact accuracy on 24-line reference chains climbed from 15.5% to 99% after researchers trained a rank-8 LoRA at a single early layer and left every other weight frozen, according to a new preprint from Zehao Jin, Ruixuan Deng, and Junran Wang.

The authors tested thirteen pretrained base models and found a shared ceiling: in standard operation each 'reliably follow only 1.4-3.6 lines' of in-context reference chains, and 'extra pretrained loops add little.' The small adapter changes that without touching the frozen weights. On Ouro-1.4B, the paper reports chains stretching to 60 lines after four loops and at least 160 after eight.

The authors frame the fix as reorganizing information flow rather than injecting new capability. 'The LoRA starts a relay: program lines pass on their chain identity through a short range of middle layers,' the paper writes, while frozen heads read progressively further up the chain. Removing parent-line attention breaks the relay, which the authors read as evidence the adapter is unlocking existing attention circuitry rather than learning the task from scratch.

The same technique also lifts scores on MuSiQue, a multi-hop question-answering benchmark, suggesting the depth-underuse pattern is not confined to synthetic chain tasks.