paper web signal

LOCI cuts world-model peak memory ~30% with hybrid cache

TL;DR

  • LOCI reports about 30% lower peak memory than a full-softmax baseline at equal sequence length.
  • On the public MIND memory benchmark, the 5B-parameter LOCI scored 14.36 PSNR versus 12.54 for the 8B-parameter HY-WorldPlay baseline.
  • The architecture pairs a KV cache for past observations with a recurrent linear-attention memory conditioned on projective camera geometry.

A new preprint introduces LOCI, a hybrid memory design that trims peak memory on streaming video world models by about 30% at equal sequence length versus a full-softmax baseline, while reproducing revisited scenes more faithfully than a battery of larger systems.

The architecture is split down the middle of the transformer. "In half of the transformer blocks, main attention keeps a key-value cache of past observations; in the other half, it is restricted to the current chunk and complemented by a recurrent linear-attention memory whose reads and writes are conditioned on projective camera geometry, so viewpoint enters both memory addressing and stored content," the authors write in the arXiv paper.

The pitch: KV caches preserve detail but blow up with video length, while recurrent state is compact but loses access to individual past observations. LOCI tries to keep both.

On the public MIND memory benchmark, LOCI (a 5B-parameter model built on the Wan2.2-TI2V backbone) scores 14.36 PSNR against HY-WorldPlay (8B) at 12.54, with reported comparisons also against Matrix-Game 3.0, AlayaWorld (15B), Alaya-EVOKE (14B), LingBot-World (28B), CaR, and a reference run of GIM-World at 1.3B parameters. Against a same-recipe full-softmax model under an identical bounded KV budget, the authors report "+0.89dB PSNR over a same-recipe full-softmax model under an identical bounded KV budget". With a bounded bank of retained observations, the project page claims a run of "300s at a constant 23.6 GiB with bounded sparse access".

Code is on GitHub at xiaji2021/LOCI and a revisit dataset is on Hugging Face; checkpoints are listed as "coming soon."

The abstract shows no matched-compute comparison against the larger baselines, and no per-benchmark memory numbers beyond the ~30% headline.