Sakana AI's AI Scientist lands in Nature after peer review
TL;DR
- Sakana AI's AI Scientist-v2 had an AI-generated paper pass blind peer review at ICLR 2025's ICBINB workshop, scoring 6, 7, 6.
- The system's Automated Reviewer reaches 69% balanced accuracy against human judgment, with F1 above NeurIPS 2021 inter-human agreement.
- Authors concede the pipeline still produces naive ideas, struggles with methodological rigor, and remains prone to hallucination.
An AI-generated paper passed blind human peer review at an ICLR 2025 workshop, with three anonymous reviewers scoring it 6, 7, and 6. Sakana AI announced that the system behind it, The AI Scientist-v2, has now been written up in Nature in a paper titled "Towards end-to-end automation of AI research."
The authors are Chris Lu, Cong Lu, Robert Tjarko Lange, Yutaro Yamada, Shengran Hu, Jakob Foerster, David Ha, and Jeff Clune, working across Sakana AI, the University of British Columbia, the Vector Institute, and the University of Oxford. The workshop in question was ICLR 2025's "I Can't Believe It's Not Better" (ICBINB). According to the Nature article, the submission was made with the full cooperation of ICLR 2025 leadership and the ICBINB organizers, and the paper was pre-committed to withdrawal if accepted. The accepted score of 6.33 "scored higher than 55% of human-authored papers," the authors report.
The system runs the full loop: generating ideas, surveying literature, designing experiments through what the team calls "parallelized agentic tree search," running them, and writing the manuscript in LaTeX.
It also reviews itself. The paper reports its Automated Reviewer hits a "balanced accuracy of 69%" against human judgment, with an F1-score the authors say exceeds the inter-human agreement measured in the NeurIPS 2021 consistency experiment.
The team is explicit about what the system still gets wrong. It "occasionally produces naive or underdeveloped ideas" and "can struggle with deep methodological rigor," they write, and it remains prone to hallucination and citation errors. One acceptance at one workshop, under disclosed conditions, is not a general claim about machine-written science.
Two of the researchers we follow had posted the Sakana writeup by the time it hit our radar.
Shared on Bluesky by 2 AI experts
-
The new Sakana AI website is live. 🐡 We have updated our digital home to reflect our products, enterprise solutions, and latest research. Sakana AIの新サイトを公開しました。エンタープライズ事例やApplied Teamの採用情報も更新しています。ぜひご覧ください! sakana.ai
View on Bluesky →
Originally reported by sakana.ai
Read the original article →Original headline: Sakana AI