GPT-6 Astra clears 2.8% of RecreationWorld agent tasks
TL;DR
- GPT-6 Astra scores 58.1% overall on RecreationBench but passes all programmatic tests on only 2.8% of its 250 tasks.
- The benchmark spans five platforms — Ubuntu, macOS, Windows, Android and Web — and asks agents to rebuild a reference app from scratch.
- The authors find agents copy static interface structure far more reliably than the interactions and computed outputs behind it.
GPT-6 Astra scores 58.1% overall on RecreationBench but passes every programmatic test on just 2.8% of tasks, according to the RecreationWorld paper posted to arXiv by a 33-author team led by Shuai Bai.
The benchmark asks a hybrid computer-use agent — one that mixes graphical interaction with code — to observe a running reference application on Ubuntu, macOS, Windows, Android or Web, then build a faithful implementation of it from scratch. RecreationBench ships 250 tasks across those five platforms.
The headline gap between the two numbers is the paper's main claim. "Agents reproduce static interface structure more reliably than interactions and computed outputs," the authors write: the buttons and layouts come back, the behavior behind them does not.
The abstract does not break the 58.1% or the 2.8% down by platform, and it does not report figures for other frontier models on the same 250 tasks.
Originally reported by paper
Read the original article →Original headline: RecreationWorld: GPT-6 Astra Scores 58% on Hybrid Agent Bench but Passes All Programmatic Tests on Just 2.8%