Faraday 27B beats Opus 4.8 and GPT-5.5 on paper replication
TL;DR
- Faraday, a 27B-parameter AI agent, outperformed Claude Opus 4.8 and GPT-5.5 on held-out paper replication tasks.
- The team built Replica, a scalable task space for replication, plus an auto-generated rubric-based judge said to track human assessment.
- Authors frame Faraday as a stepping stone toward AI agents doing long-horizon scientific work without complex harnesses.
A 27B-parameter model called Faraday, post-trained specifically to reproduce published papers, beat both Claude Opus 4.8 and GPT-5.5 on a held-out set of replication tasks, according to a new arxiv paper from Damon Falck and collaborators. Two experts in our Who's Who directory have already shared the paper.
The paper introduces Replica, described as "a scalable task space for paper replication," alongside "an auto-generated rubric-based judge that has low noise and agrees with human assessment of replication quality." Faraday uses coding agents as tools and, in the authors' words, is a "stepping stone towards AI agents capable of long-horizon scientific innovation without requiring complex harnesses."
The authors frame replication itself as research-like. "The act of replication typically illuminates details that were previously underspecified, and thus requires similar hypothesis-driven exploration to open-ended research," they write.
The abstract publishes no per-task accuracy figures, no harness details for the Claude and GPT baselines, and no numeric agreement rate between the auto-rubric and human graders.
Shared on Bluesky by 2 AI experts
Originally reported by arxiv.org
Read the original article →Original headline: Training AI Scientists to Replicate Research