arxiv.org web signal

Naver Labs distills transformer memory into recurrent agents

TL;DR

  • Naver Labs Europe researchers propose distilling full-history transformers into recurrent variants for long-horizon streaming vision and robotics.
  • A teacher model compresses its observation history into a fixed-size bottleneck representation that directly supervises the student's recurrent memory.
  • The approach achieves linear-time complexity on the Mem-RPE task and is also validated on streaming visual question answering.

A new preprint from Naver Labs Europe argues that recurrent transformers lag full-history transformers not because of architecture, but because of how they learn to compress the past.

Philippe Weinzaepfel, Christian Wolf, Mert Bülent Sariyildiz, Guillaume Bono and Gianluca Monaci state the problem flatly. "Without access to an observation history, recurrent models must explicitly decide what to retain in memory at each step, a significantly harder learning problem," they write.

Their fix is a distillation setup. A teacher transformer compresses its full observation history into a fixed-size bottleneck representation, and that bottleneck directly supervises the recurrent student's memory. The authors report the method "allows to train a recurrent latent robotic memory with linear-time complexity on the Mem-RPE task while substantially narrowing the performance gap to full-history transformers." They apply the same principle to streaming visual question answering and describe "improved recurrent predictions thanks to memory distillation."

The abstract publishes no accuracy figures and no comparison against baselines outside the authors' own full-history teacher.

Shared on Bluesky by 2 AI experts