SWE-Game: Opus5 Leads 247 Godot Agent Tasks, Still Below 60
TL;DR
- SWE-Game grounds 247 tasks in 41 executable Godot reference games spanning 13 gameplay categories in 2D and 3D.
- Across six models tested, Opus5 topped all five task types; brief-to-game reached 50.38 and construction tasks stayed below 60/100.
- Executable checks hit 92.59% balanced accuracy on human-labeled behaviors, versus 78.41% for a video-based VLM judge.
On SWE-Game, a new benchmark of 247 Godot tasks, the top coding agent tested, Opus5, reached only 50.38 out of 100 on building a game from a short brief. Across the three construction task types, no model cleared 60/100.
The paper, SWE-Game: Can Coding Agents Build the Games We Want?, covers 41 executable reference Godot games across 13 gameplay categories in 2D and 3D. Five task types test agents on development from a brief, implementation from a game design document, skeleton completion, repair of 83 injected-fault cases, and Godot-to-Unity porting. "Analysis of reviewed submissions identifies requirement omissions and gameplay logic errors as predominant implementation problems," the authors write.
On the evaluation side, executable checks reached 92.59% balanced accuracy on human-labeled behaviors from 100 agent-built games, versus 78.41% for a video-based VLM judge. Rubric-based visual scores hit a 0.829 Spearman correlation with human ratings on 200 gameplay clips. It joins a cluster of agent-benchmark papers moving through our agents feed this week, among them RobotWorld on robotics tasks.
Originally reported by huggingface.co
Read the original article →Original headline: SWE-Game Benchmarks 247 Godot Build Tasks, Claude Opus 5 Tops Scores but Stays Below 60/100