huggingface.co web signal

Ego2Act Benchmark Pins 2,640 Videos on Hand Manipulation

TL;DR

  • Ego2Act bundles 2,640 videos across 110 real-world tasks that test whether video models can simulate egocentric hand manipulation from a scene image and goal.
  • The paper reports current models often skip or partially execute steps, leaving later steps without their dependent states and failing to complete the goal.
  • The authors ship Ego2ActJudge, a reference-free evaluator they claim aligns with human consensus on task completion and physics plausibility better than baselines.

A new benchmark lands squarely on what video generation models can and cannot simulate when asked to accomplish something concrete. Ego2Act, posted on Hugging Face, is "a goal-directed benchmark featuring 2,640 videos from 110 real-world tasks across day-to-day settings, varying object clutter and multi-step complexity." Given an initial scene image and a high-level goal, it asks whether a model can generate a realistic egocentric video of a hand actually doing the task.

The verdict is unflattering. The paper reports that "models' generated simulations often skip or partially execute steps, leaving later steps missing dependent states, which leads to unfulfilled goal." The authors add that models "consistently fail at fine-grained physical dynamics, particularly during complex object manipulation and persistent world modeling."

Alongside the dataset the team releases Ego2ActJudge, described as a "reference-free evaluation pipeline that achieves better task completion and physics plausibility evaluation alignment with human consensus compared to relevant baselines." No per-model scores or ranked results appear in the abstract. It is the second Hugging Face video-benchmark paper we have tracked this week, after VTR-Bench measured on-screen text rendering.