huggingface.co web signal

RealtimeWAM distills world-action robot models to 12.2ms on H100

Robotics ai-business

TL;DR

  • RealtimeWAM cuts Fast-WAM inference from 299.7ms to 12.2ms on H100, a 24.55x speedup, with under 1% success-rate drop across LIBERO, LIBERO-Plus and RoboTwin 2.0.
  • Teacher-Anchored Consistency Distillation fixes a local-global error gap by supervising one-step students against the frozen teacher's multi-step rollout endpoint.
  • Cross-Expert Wavefront Pipelining overlaps video and action experts via block-wise sharing of the video KV cache, synchronizing only before action attention consumes it.

A new paper posted to Hugging Face by researchers at Nanyang Technological University, Beihang University, SenseTime and Continental Automotive Singapore reports pushing a Mixture-of-Transformers World Action Model from 299.7 ms per inference down to 12.2 ms on an NVIDIA H100. That is a 24.55x end-to-end speedup, with less than 1% task-success drop across LIBERO, LIBERO-Plus and RoboTwin 2.0.

The abstract frames the bottlenecks plainly: "intra-expert iteration (i.e., multi-step action denoising) and inter-expert waiting (i.e., sequential execution of the video and action experts) still limit inference efficiency." Two methods address them. Teacher-Anchored Consistency Distillation "supplements local consistency with explicit supervision from the frozen teacher's multi-step rollout endpoint, enabling accurate one-step action generation." Cross-Expert Wavefront Pipelining "overlaps the two experts through block-wise sharing of the video KV cache, synchronizing only immediately before the corresponding action attention consumes it."

On RoboTwin 2.0 the one-step student scores 90.84% overall against the Fast-WAM teacher's 91.51%. On LIBERO it matches the baseline exactly at 97.0%. On LIBERO-Plus, the harder out-of-distribution suite with camera, robot, language, lighting, background, noise and layout perturbations, overall success slips from 73.6% to 73.0%.

At 12.2 ms per step the model clears a 30 Hz control loop's 33.3 ms budget with headroom. Code and checkpoints are released under the ModelTC LightX2V inference framework. The paper joins a visible run of robotics-inference work our tracker has logged (158 robotics alerts in the last 90 days), including yesterday's RACE framework on scaling VLA action chunks.