huggingface.co web signal

HLA-WM Fixes Long-Range Forgetting in Video World Models

TL;DR

  • On 60-second SANA-WM-Bench, HLA-WM lifts the base autoregressive generator by 0.74 dB PSNR and cuts rotation error by 28.5% without any retraining.
  • The method uses roughly 12x less historical-state storage than full KV caching while costing at most a 1.6% inference-throughput hit.
  • On MBench-A, all three revisit-consistency metrics improve across all four subsets and every evaluated inference mode.

A team from Monash University, the University of Adelaide, and Zhejiang University is proposing a training-free fix for a specific failure mode in long video world models: when the camera leaves a scene and later returns, the model generates something plausible but spatially inconsistent with what it saw before. Their paper on Hugging Face, dated October 5, 2026, blames Gated DeltaNet's single accumulated recurrent state, which quietly attenuates early scene information as every intermediate chunk's transition matrix is applied on top of it.

The authors quantify the decay on a representative SANA-WM trajectory: a camera leaves the region observed at Chunk 3, loops around, and returns near Chunk 35, by which point the retained influence of the original memory has fallen to 0.0416. "Although the early scene becomes relevant again near the end of the rollout, its retained influence continuously decreases as intermediate chunks are generated," the paper reports. Their diagnosis is a mismatch between how recurrent memory is updated and what actually matters: "historical information is updated according to temporal progression, whereas its relevance in world-model rollouts is often determined by spatial and geometric proximity."

HLA-WM's answer is to cache compact chunk-wise summaries of GDN's state transitions, address them with a camera-frustum overlap proxy built from poses, intrinsics, and a scene depth estimate, and recompose the selected summaries into a query-specific recurrent state. Pretrained weights stay frozen. On the 60-second SANA-WM-Bench, the authors report a 0.74 dB PSNR gain and a 28.5% rotation-error reduction over the base generator. On MBench-A, all three revisit-consistency metrics improve across all four subsets and every inference mode they test.

The efficiency claim is the other half of the pitch: at a 60-second context, HLA-WM "requires 12× less historical-state storage than full KV caching while incurring at most a 1.6% reduction in inference throughput across the evaluated pipelines." The paper also notes that "downstream refinement introduces mode-dependent trade-offs," without breaking those out in the abstract. It lands in a busy week for video-model research we are tracking on the AI video beat, alongside Kling AI's HK IPO filing and the Ego2Act benchmark release.