ASI-Bench halves 18 frontier agents' scores without guidance
TL;DR
- On ASI-Bench, the average score across 18 frontier agent-model configurations falls from 50.91 with full guidance to 26.62 when the agent picks the method.
- The benchmark covers 60 project-level research tasks across 11 scientific domains, built by more than 40 experts and over 31,000 human hours.
- Every task goes through expert review, AI-assisted auditing, sandbox execution, and scorer validation before it is admitted.
The average score of 18 frontier agent–model configurations on ASI-Bench drops from 50.91 with full human methodological guidance to 29.10 when only the method is specified, and to 26.62 when the agents must determine the method themselves, according to the paper posted to arxiv on August 18, 2026.
The benchmark contains 60 project-level research tasks across 11 scientific domains, built by over 40 experts at a cost of 31,000+ human hours. Every task, the authors write, goes through "expert review, AI-assisted auditing, sandbox execution, and scorer validation."
The authors describe ASI-Bench as "the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own."
Their reading of the numbers is direct: "This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research."
The abstract does not name the 18 configurations, break scores down by domain, or specify the scoring ceiling.
Originally reported by paper
Read the original article →Original headline: ASI-Bench: 18 Frontier Agents Score 50.91 With Full Guidance, 26.62 When Working Alone