huggingface.co web signal

MBZUAI's WorldGuide Hits 47.69% on Video-CraftBench Goal-Only

TL;DR

  • WorldGuide reports 47.69% task success on Video-CraftBench given only a starting image and task goal, against 32.73% for MiniMax-H3.
  • Ablating the closed loop drops task success from 33.33% to 11.71% on WorldGuide Bench, a 21.62-point swing.
  • The system pairs a Qwen2.5-VL-7B-Instruct planner with a HunyuanVideo-1.5 executor and trains on 58,679 step-level demonstrations across 245 tasks.

A video world model from the Mohamed bin Zayed University of Artificial Intelligence reports 47.69% task success on Video-CraftBench given only a starting image and a task goal, compared with 32.73% for MiniMax-H3 under the same conditioning. On the authors' own WorldGuide Bench, it clears 33.33% against MiniMax-H3's 29.90%, and MiniMax-H3 was given reference action plans while WorldGuide was not.

WorldGuide runs a closed loop. 'Given only an initial image and task goal, it predicts atomic actions, generates corresponding video clips, and uses generated results to select the next action or terminate,' the abstract states. A Qwen2.5-VL-7B-Instruct planner picks the next atomic action; a HunyuanVideo-1.5 executor renders the clip; the planner then reads the rendered frames to decide the next action or emit a <|Task Completed|> token.

The ablation is the headline. Strip out the loop and plan everything up front, with no visual feedback, and task success drops from 33.33% to 11.71%, a 21.62-point swing on WorldGuide Bench. A hierarchical memory that keeps recent latent frames at full spatial resolution and older ones down to 1/64 scale, capped at 1,366 latent frames, adds another 17.91 points on Video-CraftBench.

Training used 58,679 step-level demonstrations across 245 tasks in 27 procedural categories, including origami, cooking, assembly, repair, and knotting. A human audit of the dataset pegged task-label correctness at 83.7% and essential-action coverage at 84.5%, which bounds how clean the ground truth is before any model touches it. The paper lands in a busy stretch for the segment on our tracker, with 46 video generation stories in the past 90 days, including OuroWorld's static-scene cinemagraph loop posted the same day.