huggingface.co web signal

RobotWorld Benchmarks VLMs on 84 Robot Tasks, 63 Go Unsolved by Any Frontier Model

Summary

RobotWorld tests five frontier models on 84 simulated manipulation, mobile manipulation, locomotion, driving and aerial tasks across embodiments from fixed-base arms to drones. GPT-6 Astra led at 19% success (16/84), followed by Claude Opus 5.5 at 15.5% — only 21 of 84 tasks were solved by any model. Astra was stronger on spatial and contact goals; Opus 5.5 better on continuous-balance and timed tasks like drone juggling.