On the one hand, it may not ever be great at truly open ended forms of discovery, which aren't readily testable and aren't obviously amenable to some kind of reinforcement learning type feedback. So "geniuses in a lab" may never turn out to be generalizable. arxiv.org/abs/2607…
- Frontier agents got six days and thousands of dollars of compute to tackle the central research question of two unpublished NeurIPS 2026 papers.
- The agents completed all the engineering unaided but the papers' original authors unambiguously rejected the output as research.
- The 24-author team names five recurring failure modes and releases the expert reviews, agent repositories, and logs for replication.