huggingface.co web signal

UNREAL turns a frozen LLM into its own retriever with <500K params

TL;DR

  • UNREAL adds fewer than 500K trainable parameters to a frozen LLM, under 0.005% of each backbone's parameter count.
  • On a 3B-token, 21M-chunk Wikipedia index, recall@10 rises from 49.1% to 73.2% on HotpotQA and 31.7% to 60.1% on 2WikiMultiHopQA.
  • At 128K tokens, NoLiMa accuracy climbs from 1.0% to 24.83%; the method cuts FLOPs and time-to-first-token from roughly 32K tokens onward.

A new paper on Hugging Face Papers argues that corpus retrieval and long-context inference are the same operation at different scales, and shows a single frozen LLM can do both. UNREAL — short for Unifying REtrieval And Long-Context with a Single Model — adds "fewer than 500K trainable parameters" on top of a frozen backbone, which the authors note is "under 0.005% of each tested backbone's parameter count," and derives retrieval queries directly from the model's residual stream rather than a separate encoder.

The headline numbers come from a 3B-token, 21M-chunk Wikipedia index. The paper reports complete-evidence recall@10 rising "from 49.1% to 73.2% on HotpotQA, from 31.7% to 60.1% on 2WikiMultiHopQA, and from 8.8% to 14.4% on MuSiQue." Applied as a long-context selector, the same module raises NoLiMa accuracy "from 1.0% to 24.83% at its maximum context length of 128K tokens," and LV-Eval F1 "from 49.97% to 54.66% at 256K." The authors say UNREAL "reduces FLOPs and time-to-first-token relative to full-context inference from roughly 32K tokens onward, with larger gains as context grows."

Four backbones were tested — the dense Qwen3.5-4B and Muse-Glimmer-30B, plus the hybrid Qwen3.5-35B-A3B and Nemotron-3.5-Lightning-30B-A3B — and "all four tested backbones outperform the strongest dedicated retriever-reranker systems," including BM25, Qwen3-Embedding, BGE, a Jina cross-encoder reranker, and a LightOn multi-vector retriever. It is the tenth RAG story we have tracked in the last 90 days, and the second this week to argue that the retriever-plus-generator split is becoming an architectural accident rather than a necessity.