NVIDIA Long-WAM lifts RoboCasa GR-1 to 78.7% with 19.2s context
TL;DR
- Long-WAM raises RoboCasa GR-1 success from 63.3% at zero context to 78.7% at 19.2 seconds of visual history.
- A bidirectionally pretrained initialization shows no net gain over the same window, isolating the win to autoregressive video pretraining.
- On a Unitree G1 the policy reports 95% success on dynamic cup stacking versus 0 of 20 trials for π₀.₅ and Fast-WAM, at 107.4 ms per action chunk on RTX 5090.
On RoboCasa GR-1, extending a world-action model's visual context from 0 to 19.2 seconds lifts task success from 63.3% to 78.7%, but only when the underlying video foundation was pretrained autoregressively. A bidirectionally pretrained initialization shows no net gain over the same window, according to Long-WAM, a paper from NVIDIA, MIT, HKU and UCSD posted to Hugging Face.
"Access to history is not the same as using it: longer histories pay off far more when the video foundation is pretrained autoregressively," the abstract states. The team pretrains a model called LongLive2.0-Robot on roughly 10,000 window-equivalent hours of robot and egocentric video without action labels, then adapts it for world-action prediction while preserving the history-to-future structure.
On a Unitree G1 the policy reports 95% success on dynamic cup stacking, where "π₀.₅ and Fast-WAM succeed in none of 20 trials," per the abstract. Each action chunk on an RTX 5090, "including future-video latent prediction, takes 107.4 ms."
Long-WAM also claims best-of-compared results on LIBERO-Long, RoboTwin 2.0 and DOMINO, and the authors position it as a "memory-informed executor" that complements higher-level planning in composite tasks. It joins a steady run of NVIDIA robotics work in our tracker this quarter.
Originally reported by huggingface.co
Read the original article →Original headline: NVIDIA's Long-WAM Shows AR Pretraining Lets Robots Use 19.2s of Visual Context, Lifts RoboCasa GR-1 to 78.7%