RobotWorld Benchmarks VLMs on 84 Robot Tasks, 63 Go Unsolved by Any Frontier Model
Summary
RobotWorld tests five frontier models on 84 simulated manipulation, mobile manipulation, locomotion, driving and aerial tasks across embodiments from fixed-base arms to drones. GPT-6 Astra led at 19% success (16/84), followed by Claude Opus 5.5 at 15.5% — only 21 of 84 tasks were solved by any model. Astra was stronger on spatial and contact goals; Opus 5.5 better on continuous-balance and timed tasks like drone juggling.
Originally reported by huggingface.co
Read the original article →Original headline: RobotWorld Benchmarks VLMs on 84 Robot Tasks, 63 Go Unsolved by Any Frontier Model