AutoResearchEval finds 45 failure modes across all agent setups
TL;DR
- AutoResearchEval evaluated eight harness-model combinations on 100 frontier-science tasks and found the same 45 failure patterns in every combination.
- The authors attribute the pattern to a missing metacognitive loop and place the deficit at the model level, not the scaffold.
- The benchmark spans seven scientific domains and the full research lifecycle from ideation through review, yielding 800 annotated agent trajectories.
AutoResearchEval, an evaluation framework posted to arxiv, tested eight harness-model combinations against 100 tasks drawn from published frontier science and produced 800 agent trajectories with process-level annotation. The 45 failure patterns the authors compile into the AutoResearch Failure Taxonomy, or ARFT, recur across every combination tested, "including the strongest models tested."
The paper attributes the convergence to a single missing capability. "Current agents lack a metacognitive loop, which entails the ability to check what they produced against what they found, revise when it does not hold up, and question whether the path they took was sound," the abstract states. The 100 tasks span seven scientific domains and the full research lifecycle: ideation, retrieval, execution, analysis, writing, and review.
Because the same failures appear regardless of scaffold, the authors place the deficit "at the model level rather than in any particular scaffold." Whether orchestration-level interventions can close it, they write, "is an open question this work does not test." AutoResearchEval and ARFT are being publicly released.
Originally reported by paper
Read the original article →Original headline: 45 Failure Modes Recur Across Every AI Research Agent Tested—Including Strongest Models—on 100 Frontier Science Tasks