huggingface.co web signal

WorldGuide Video World Model Clears 47.69% of VideoCraft-Bench via Closed-Loop Planner-Executor Loop

Summary

WorldGuide chains a ContextPlanner that predicts the next atomic action from visual progress with an Executor that renders the clip, then re-reads the generated frames to pick the next step or terminate. On the authors' 59K-clip WorldGuide Bench it hits 33.33% task success vs 29.90% for MiniMax-H3 given reference actions, and reaches 47.69% on VideoCraft-Bench vs 32.73% goal-only. Closed-loop execution contributes +21.62% over an open-loop ablation.