huggingface.co web signal

RobotWorld: Frontier VLMs Leave 63 of 84 Robot Tasks Unsolved

TL;DR

  • GPT-6 Astra led at 16/84 tasks (19.0%); Claude Opus 5.5 followed at 13/84 (15.5%); Kimi K3, DeepSeek V4.1 Flash, and Gemini 3.8 Flash together added only 1 further unique task.
  • Across five frontier models, 63 of 84 tasks went unsolved by anyone — the five domains span manipulation, mobile manipulation, locomotion, driving, and aerial control.
  • Astra was stronger on spatial and constrained-contact goals; Opus 5.5 was stronger on continuous-balance and timed-interaction goals and cleared 3 of 4 aerial tasks.

Frontier models cleared only 21 of 84 physical-world tasks in RobotWorld, a new simulation testbed posted to Hugging Face that spans manipulation, mobile manipulation, locomotion, driving, and aerial control. The remaining 63 tasks went unsolved by any of the five models tested.

GPT-6 Astra led at 16/84 (19.0%), followed by Claude Opus 5.5 at 13/84 (15.5%). Kimi K3 managed 2/84, and DeepSeek V4.1 Flash and Gemini 3.8 Flash each cleared a single task. The two leaders overlapped on 8 tasks, with Astra adding 8 unique wins and Opus 5.5 adding 5.

The capability profiles split cleanly along the kind of control the task demands. The paper reports that "Astra succeeds more often on spatial and constrained-contact goals, whereas Opus 5.5 succeeds more often on continuous-balance and timed-interaction goals." Opus 5.5 cleared 3 of 4 aerial-control tasks (75.0%) and the one quadruped locomotion task either leader solved; Astra was stronger across the 38-task manipulation set, scoring 23.7% there against Opus 5.5's 18.4%.

The authors argue the gap is one of composition, not perception. Current agents "can construct sophisticated perception and control workflows, including image segmentation, camera calibration, spatial estimation, and dynamics-based computation," but those building blocks "do not consistently compose into successful behaviour": agents "lose task-relevant object states despite reaching commanded poses, fail to correct ineffective actions, recover too late, or mistake unfinished tasks for completion."

Each model-task pair was run once over September 30 to October 6, 2026. Astra's evaluation cost an estimated $9,913 against 941M tokens; Opus 5.5 cost $2,916 against 1.63B tokens. RobotWorld is the third robotics evaluation we have logged in two days, alongside RoboQuest and SWE-Game.