HF Paper HLA-WM Uses Hybrid Linear Attention to Stabilize Long-Horizon Video World Models
Summary
A Monash/Adelaide/Zhejiang team published HLA-WM on October 5, pairing hybrid linear attention with KV caching so video world models can recover previously-seen scenes after they leave recent context without unbounded memory growth. The paper argues long-horizon video simulation breaks when models rely purely on full-history KV caches, and shows the hybrid attention stack preserves spatial consistency when the camera returns to earlier views.
Originally reported by huggingface.co
Read the original article →Original headline: HF Paper HLA-WM Uses Hybrid Linear Attention to Stabilize Long-Horizon Video World Models