arxiv.org web signal

Tan et al. survey 59 attention variants across 14 LLM lineages

TL;DR

  • The survey reads 59 release-level records across 14 major model lineages through one five-dimensional lens covering memory representation, update, access, readout, and integration.
  • The paper names the core bottleneck up front: dense token interactions carry quadratic prefill cost and a key-value cache that grows with context length.
  • The authors argue explicit-memory and recurrent-state architectures retain distinct interfaces while increasingly handling overlapping memory functions, blurring the attention-versus-recurrence split.

The paper reads 59 release-level records from 14 major model lineages through one five-dimensional lens: Memory Representation, Memory Update, Access, Readout, and Integration.

The driver, as the arxiv preprint by Zhentao Tan and co-authors states up front, is cost. "Dense token interactions incur quadratic prefill cost and a key-value cache growing with context length," they write. That is what any team scaling context length is paying for.

The survey's main architectural claim is that the old divide between transformer-style and recurrent-style methods is thinner than it looks. "Explicit-memory and recurrent-state methods retain distinct interfaces but increasingly control overlapping memory functions," the authors write, drawing the point out of the 59 records they catalogue.

Depth itself is treated as a design dimension. Layer-wise composition distributes complementary memory processing across network depth rather than stacking identical blocks, and persistent memory, in the authors' framing, is "organized across temporal scope, network depth, substrate type, and representation granularity."

The abstract names neither the 14 lineages nor the 59 records the survey draws on.

Shared on Bluesky by 2 AI experts