arxiv.org web signal

SciConBench: Top AI research agent scores just 0.337 F1

TL;DR

  • SciConBench scores 8 frontier models and deep research agents against 9.11K expert-written conclusions from systematic reviews; the best factual F1 is 0.337.
  • A new clean-room harness called SciConHarness consistently drops agent scores versus unconstrained evaluation, suggesting prior benchmarks were inflated by data leakage.
  • Audits of consumer tools including Google AI Overview and OpenEvidence found incomplete and sometimes contradictory conclusions, even when ground truth was retrievable.

The best scientific AI agent tested scored only 0.337 on factual F1, a measure of whether its synthesized conclusions matched expert-written ones, according to a new benchmark posted on arXiv.

SciConBench, as the authors call it, holds 9.11K questions paired with expert-written conclusions drawn from systematic reviews. An automated pipeline decomposes each answer into atomic facts and scores it on factual precision and recall. The paper evaluates 8 frontier models and deep research agents, running them through a companion clean-room harness, SciConHarness, that gives each agent controlled web access to block the evaluation set from leaking in through retrieval.

That design choice is where the numbers get uncomfortable. "Our clean-room setting consistently reduces performance relative to unconstrained evaluation," the paper reports, "suggesting that leakage inflates estimates of models' true synthesis capabilities." In plain terms, scores reported without a clean-room, which is most published numbers, are probably overstated.

The paper reserves its sharpest finding for consumer-facing systems. In audits of Google AI Overview and OpenEvidence, the authors write, the tools "frequently generate incomplete and sometimes contradictory conclusions, even when the ground-truth answer is available."

The authors frame the stakes in health, where "scientific AI agents increasingly retrieve evidence, reason across sources, and synthesize conclusions used in consequential decisions," and conclude that reliable synthesis of such conclusions "remains an open challenge, and that clean-room evaluation is essential for assessing open-domain AI agents."

Shared on Bluesky by 2 AI experts