AI agents finish the engineering of AI research, miss the science
TL;DR
- Frontier agents got six days and thousands of dollars of compute to tackle the central research question of two unpublished NeurIPS 2026 papers.
- The agents completed all the engineering unaided but the papers' original authors unambiguously rejected the output as research.
- The 24-author team names five recurring failure modes and releases the expert reviews, agent repositories, and logs for replication.
The most rigorous public attempt yet to test whether AI agents can actually do AI research just came back with a clear 'not yet.' Twenty-four researchers, including Peter Kirgis, Sayash Kapoor, Arvind Narayanan, Helen Toner, Gillian Hadfield, and Rishi Bommasani, took two unpublished NeurIPS 2026 submissions, handed frontier agents the central open-ended research question of each, gave them six days and 'thousands of dollars of compute,' and asked the papers' original authors to grade the output. The preprint on arXiv calls this setup 'shadow evaluations,' pitched as a middle ground between narrow verifiable benchmarks and dumping AI-generated papers into blind peer review.
The finding is sharp. The agents 'completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions.' Both papers were, in the authors' phrasing, 'unambiguously rejected.' A robustness check with a second model and scaffold reproduced the same failures.
Why this matters: forecasts of explosive AI progress typically assume agents will soon be doing AI research themselves, compressing the R&D loop. This is one of the first controlled looks at that claim where the graders are the people whose own unpublished work is being reproduced, not a benchmark script. The paper enumerates five recurring failure modes, 'poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift,' which reads more like a taxonomy of research taste than a list of missing tools.
The honest caveat is scale. Two papers is a small sample, and the abstract does not name which frontier agents were tested or say how the compute budget compared with what the human authors used. Take the specifics as reported, not settled. The team is releasing the expert reviews, survey responses, agent repositories, and logs, so the methodology can be rerun as new models ship.
For anyone building on the R&D-automation thesis, whether that is investors pricing in fast timelines, safety teams sizing their runway, or leaders sequencing agent bets, the direction of the evidence is worth internalizing. The engineering layer is arriving. The judgment layer is not, at least not on this evidence and not yet.
Originally reported by paper
Read the original article →Original headline: Frontier Agents Ace Research Engineering, Fail at the Actual Science—24-Author Study