huggingface.co web signal

Alibaba Amap, CASIA DreamTrue Cuts Interaction Defects to 6.25%

Alibaba Robotics Research ai-research

TL;DR

  • DreamTrue cuts the human-rated interaction defect rate on AgiBot from 48.12% to 6.25% and the object defect rate from 31.88% to 3.12%.
  • The system took first place in the AgiBot World Challenge 2026 world-model track with an nDTW of 0.8772 for action following.
  • The authors release calibration data for 153,666 episodes spanning 1,660 hours and a reward-model corpus of 44.9K annotated videos with 30.4K defect labels.

The headline figure: on the AgiBot benchmark, DreamTrue drops the human-rated interaction defect rate from 48.12% to 6.25%. The object defect rate falls from 31.88% to 3.12%. Those numbers come from three independent annotators voting on 960 videos, with a Fleiss' κ of 0.7817 on any-defect labels.

DreamTrue is a multi-view, cross-embodiment robot world model from the Institute of Automation at the Chinese Academy of Sciences (CASIA) and Amap, Alibaba Group, described in a preprint on Hugging Face. It took first place in the AgiBot World Challenge 2026 world-model track, posting the highest overall score, the highest action-following score and the highest visual quality score. On AgiBot it reaches an nDTW of 0.8772, and 0.8831 under counterfactual actions, which are trajectories deliberately perturbed with SE(3) rotations and translations.

The method has two ingredients. An offline geometric calibration step renders action trajectories into image-space conditions; the authors report mean IoU improvements of +0.228 on AgiBot (97.1% of episodes improved) and +0.212 on DROID (91.6% improved), without needing dedicated calibration sequences or depth sensors. A counterfactual post-training stage then generates synthetic videos under modified trajectories, scored by an embodied video reward model trained on 44.9K annotated videos with 30.4K defect labels across three dimensions: embodiment, object, interaction.

Alongside the paper the authors release calibration data for 153,666 retained episodes spanning 1,660 hours of video, drawn from AgiBotWorld-Beta, DROID and RoboMIND 2.0. The generator is built on Wan2.1-VACE-14B with roughly 0.7B trainable parameters and LoRA rank 128.

The authors flag the obvious failure mode themselves: "The action representation, generator, and VLM reward model all operate in image space; under occlusion or limited views, the framework may generate or reward visually plausible but physically incorrect interactions." The paper joins a visible run of robotics world-model work in our robotics feed, including Meta FAIR's RoboJEPA release earlier this week.