huggingface.co web signal

SWE-Game: Opus5 Leads 247 Godot Agent Tasks, Still Below 60

TL;DR

  • SWE-Game grounds 247 tasks in 41 executable Godot reference games spanning 13 gameplay categories in 2D and 3D.
  • Across six models tested, Opus5 topped all five task types; brief-to-game reached 50.38 and construction tasks stayed below 60/100.
  • Executable checks hit 92.59% balanced accuracy on human-labeled behaviors, versus 78.41% for a video-based VLM judge.

On SWE-Game, a new benchmark of 247 Godot tasks, the top coding agent tested, Opus5, reached only 50.38 out of 100 on building a game from a short brief. Across the three construction task types, no model cleared 60/100.

The paper, SWE-Game: Can Coding Agents Build the Games We Want?, covers 41 executable reference Godot games across 13 gameplay categories in 2D and 3D. Five task types test agents on development from a brief, implementation from a game design document, skeleton completion, repair of 83 injected-fault cases, and Godot-to-Unity porting. "Analysis of reviewed submissions identifies requirement omissions and gameplay logic errors as predominant implementation problems," the authors write.

On the evaluation side, executable checks reached 92.59% balanced accuracy on human-labeled behaviors from 100 agent-built games, versus 78.41% for a video-based VLM judge. Rubric-based visual scores hit a 0.829 Spearman correlation with human ratings on 200 gameplay clips. It joins a cluster of agent-benchmark papers moving through our agents feed this week, among them RobotWorld on robotics tasks.