RoboQuest Benchmark: Best Frontier Model GPT-6 Astra Clears Just 23.2% of Goal-Directed Exploration Tasks
Summary
RoboQuest puts five frontier VLMs on 10 MuJoCo tasks that require actively acquiring hidden information: GPT-6 Astra tops out at 23.2% success ($11.29/episode), Claude Opus 5.5 at 13.8%, GPT-6.1 Sol at 12.2%, Claude Fable 5.1 at 11.4%, Gemini 3.8 Flash at 2.0%. Attribution shows 43% of GPT-6 Astra's 1,143 failures were 'missing evidence' — the robot stopped exploring with target compartments never opened — not execution error, since isolated skills hit 80-81%.
Originally reported by huggingface.co
Read the original article →Original headline: RoboQuest Benchmark: Best Frontier Model GPT-6 Astra Clears Just 23.2% of Goal-Directed Exploration Tasks