Stanford's Paper2Agent turns research papers into AI agents
TL;DR
- Paper2Agent converts a research paper and its code into a working AI agent in about 45 minutes for roughly $14 on a personal laptop.
- In a 100-paper sweep of computational biology, 74 papers yielded working agents and 593 of 599 auto-generated tools passed validation.
- The AlphaGenome-based agent flagged GPR137 as the probable causal gene for the psoriasis-associated variant rs887314.
Paper2Agent converted 74 of 100 computational biology papers into working AI agents in an evaluation published in Nature on 16 September by a Stanford team led by Jiacheng Miao and James Zou. Of the 599 tools the framework proposed across those papers, 593 passed validation.
The system takes a paper together with its codebase and builds a set of model context protocol (MCP) servers that expose the paper's methods as callable tools, generating its own tests as it goes. Users then query the resulting agent from a host like Claude Code. The authors demonstrate agents built from AlphaGenome (genomic variant interpretation), Scanpy (single-cell analysis) and TISSUE (spatial transcriptomics).
"Paper2Agent reimagines research dissemination by turning static papers into active AI agents," the paper argues. "Each agent serves as an interactive expert on the corresponding paper, capable of demonstrating, applying and adapting its methods."
The AlphaGenome agent scored "98.7 ± 1.3% accuracy on 15 tutorial-derived queries and 100.0 ± 0.0% accuracy on 15 novel queries", against 82.7 ± 3.4% and 78.7 ± 4.4% for a Claude plus repository baseline, at 1.9× lower median runtime. Generating the 22 AlphaGenome tools took about 45 minutes and cost roughly $14 on a personal laptop.
In one case study, the AlphaGenome-based agent flagged GPR137 as the probable causal gene for the psoriasis-associated variant rs887314, reporting Spearman correlation = 0.613, P = 3.79 × 10−3 between predicted regulatory signatures and expression data. Five researchers in our Who's Who tracker had posted the link by the time we picked it up.
Twenty-six of the hundred candidate papers did not yield working agents, and the demonstrations here sit inside a Claude Code host.
Shared on Bluesky by 5 AI experts
-
It would be fun to listen to a future episode of the Mystery AI Hype Theater 3000 podcast (that uses *ridicule as praxis") about this article 🤣 "Reimagining research papers as interactive and reliable AI agents" www.nat…
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Reimagining research papers as interactive and reliable AI agents - Nature