paper web signal

InternW0-Δ pretrains on 20K-hour open robot corpus

TL;DR

  • The paper claims a training corpus of over 20K hours, which the authors call the largest open-source corpus of its kind.
  • The corpus fuses robot demonstrations, UMI data, egocentric human demonstrations and Ego2Robot data under a common state-action representation.
  • Authors pledge to release training code, model weights, infrastructure, data-processing pipeline and processed data where licenses permit.

InternW0-Δ, a new preprint from the InternRobotics group, pretrains a unified world-action model on what the authors describe as "over 20K hours of processed training data, to our knowledge the largest open-source corpus of its kind." The paper landed on arXiv on September 25.

The model combines "pretrained visual dynamics, scene-level semantics, 4D geometric and motion priors, and action generation within a Mixture-of-Transformers (MoT) framework." A frozen VLM provides semantic guidance to a video expert and an action expert, and a mechanism the authors call "Causal Imprint" supplies predictive representations "directly to the action expert without future-video rollout at inference."

The corpus fuses robot demonstrations, UMI data, egocentric human demonstrations, and Ego2Robot data, "curated and aligned under a common state-action representation." The abstract reports "strong performance across simulation benchmarks and real-robot platforms" but publishes no per-benchmark numbers.

The release plan is broad and conditional. The paper says the team "will open source the training code, model weights, infrastructure, data-processing pipeline, and processed data where licenses permit." The companion GitHub repository already ships code under an MIT license along with task-specific checkpoints for LIBERO, RoboTwin, and RoboDojo, and lists Hugging Face hosting of full model weights as coming soon.