huggingface.co web signal

RoboQuest Benchmark: Best Frontier Model GPT-6 Astra Clears Just 23.2% of Goal-Directed Exploration Tasks

Robotics OpenAI Anthropic ai-research

Summary

RoboQuest puts five frontier VLMs on 10 MuJoCo tasks that require actively acquiring hidden information: GPT-6 Astra tops out at 23.2% success ($11.29/episode), Claude Opus 5.5 at 13.8%, GPT-6.1 Sol at 12.2%, Claude Fable 5.1 at 11.4%, Gemini 3.8 Flash at 2.0%. Attribution shows 43% of GPT-6 Astra's 1,143 failures were 'missing evidence' — the robot stopped exploring with target compartments never opened — not execution error, since isolated skills hit 80-81%.