Paper2Agent turns 74 biology papers into runnable AI agents
TL;DR
- Paper2Agent, published in Nature on September 16, converts research papers into interactive AI agents served through Model Context Protocol servers.
- The framework agentified 74 of 100 computational biology papers into 593 validated tools, scoring 91.2% on 300 benchmark questions.
- An AlphaGenome-based agent hit 98.7% accuracy on tutorial queries and 100% on novel ones, generated in 45 minutes for about $14.
Paper2Agent, published in Nature on September 16, converts research papers into interactive AI agents by wrapping their manuscript, code, data and workflows behind a Model Context Protocol server that a chat agent can call. The framework agentified 74 of 100 computational biology papers, producing 593 validated tools that answered 91.2% of 300 benchmark questions correctly.
The authors, Jiacheng Miao, Joe R. Davis, Yaohui Zhang, Jonathan K. Pritchard and James Zou, describe a six-step pipeline: locate the paper's codebase, set up an isolated environment, discover tutorials, run them end-to-end, extract and validate callable tools, then assemble them into an MCP server. Claude Code is the agent framework doing the work.
"Each agent serves as an interactive expert on the corresponding paper, capable of demonstrating, applying and adapting its methods to new projects," the paper states. On the AlphaGenome case study, the resulting agent generated 22 tools in about 45 minutes for roughly $14, hit 98.7 ± 1.3% accuracy on tutorial-derived queries and 100.0 ± 0.0% on novel ones, against 82.7 ± 3.4% and 78.7 ± 4.4% for a Claude-plus-repository baseline. Median runtime dropped 1.9× and 3.1×.
The GitHub release puts the code in front of anyone who wants to reproduce it, and five of the researchers tracked in our directory had already shared the article by the time it landed on our radar.
The authors frame the shift bluntly: "Paper2Agent shifts the research output from a document or codebase encoding knowledge to a knowledgeable entity capable of execution and dialogue." The paper reports a 100% correct-rejection rate for out-of-scope queries. It does not report how well the pipeline generalises past the 100 computational biology papers it was tested on.
Shared on Bluesky by 5 AI experts
-
It would be fun to listen to a future episode of the Mystery AI Hype Theater 3000 podcast (that uses *ridicule as praxis") about this article 🤣 "Reimagining research papers as interactive and reliable AI agents" www.nat…
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Reimagining research papers as interactive and reliable AI agents - Nature