RealtimeWAM distills world-action robot models to 12.2ms on H100
TL;DR
- RealtimeWAM cuts Fast-WAM inference from 299.7ms to 12.2ms on H100, a 24.55x speedup, with under 1% success-rate drop across LIBERO, LIBERO-Plus and RoboTwin 2.0.
- Teacher-Anchored Consistency Distillation fixes a local-global error gap by supervising one-step students against the frozen teacher's multi-step rollout endpoint.
- Cross-Expert Wavefront Pipelining overlaps video and action experts via block-wise sharing of the video KV cache, synchronizing only before action attention consumes it.
A new paper posted to Hugging Face by researchers at Nanyang Technological University, Beihang University, SenseTime and Continental Automotive Singapore reports pushing a Mixture-of-Transformers World Action Model from 299.7 ms per inference down to 12.2 ms on an NVIDIA H100. That is a 24.55x end-to-end speedup, with less than 1% task-success drop across LIBERO, LIBERO-Plus and RoboTwin 2.0.
The abstract frames the bottlenecks plainly: "intra-expert iteration (i.e., multi-step action denoising) and inter-expert waiting (i.e., sequential execution of the video and action experts) still limit inference efficiency." Two methods address them. Teacher-Anchored Consistency Distillation "supplements local consistency with explicit supervision from the frozen teacher's multi-step rollout endpoint, enabling accurate one-step action generation." Cross-Expert Wavefront Pipelining "overlaps the two experts through block-wise sharing of the video KV cache, synchronizing only immediately before the corresponding action attention consumes it."
On RoboTwin 2.0 the one-step student scores 90.84% overall against the Fast-WAM teacher's 91.51%. On LIBERO it matches the baseline exactly at 97.0%. On LIBERO-Plus, the harder out-of-distribution suite with camera, robot, language, lighting, background, noise and layout perturbations, overall success slips from 73.6% to 73.0%.
At 12.2 ms per step the model clears a 30 Hz control loop's 33.3 ms budget with headroom. Code and checkpoints are released under the ModelTC LightX2V inference framework. The paper joins a visible run of robotics-inference work our tracker has logged (158 robotics alerts in the last 90 days), including yesterday's RACE framework on scaling VLA action chunks.
Originally reported by huggingface.co
Read the original article →Original headline: RealtimeWAM Cuts MoT Robot World-Action Model Inference 24.5x With Teacher-Anchored Distillation