Station Agents Rediscover 62.7% of ICLR Oral Paper Findings
TL;DR
- Station, a multi-agent open-world environment, rediscovered 62.7% of the findings criteria across three ICLR oral papers on average.
- Baselines trailed well behind: Codex Multiagent-v2 scored 15.4% and AI Scientist-v2 landed between 14.4% and 20.6% on the same tasks.
- Agents received only each paper's main research question; the authors withheld the paper's results and disabled web access during the runs.
Station, a multi-agent system, rediscovered 62.7% of the findings criteria from three ICLR oral papers after being handed only each paper's main research question, with the papers' own results withheld and web access disabled. The number comes from an arXiv preprint by Wenyu Du and Stephen Chung.
The baselines sit far below. Codex Multiagent-v2 scored 15.4% on the same tasks. AI Scientist-v2 scored between 14.4% and 20.6%.
The paper describes Station as 'an open-world environment in which multiple agents simulate a scientific ecosystem.' Two added mechanisms, a Supervisor and periodic Meta Reflection, 'encourage persistent exploration even when intermediate metrics are lacking,' the authors write.
Du and Chung also ran Station on 'two open-ended tasks without oracle papers' and report that some agent discoveries 'closely match discoveries reported by researchers after the knowledge cutoff date.' The abstract names neither task and gives no match figure.
The authors' own summary: the results 'indicate that a suitable environment can enable agents to autonomously make meaningful progress in open-ended scientific discovery.'
Originally reported by paper
Read the original article →Original headline: Station Multi-Agent System Autonomously Rediscovers 63% of ICLR Findings, 4x Rival Agents