Stanford's Paper2Agent turns 74 of 100 papers into agents
TL;DR
- Paper2Agent, published in Nature, converted 74 of 100 computational biology papers into working AI agents wrapped as Model Context Protocol servers.
- The AlphaGenome-derived agent produced 22 validated tools in about 45 minutes for $14 and scored 98.7% on tutorial queries.
- Two paired paper-agents flagged an MPHOSPH9 variant tied to ADHD risk that Zou says had not previously been reported.
A team at Stanford has published a framework in Nature that automatically converts research papers into interactive AI agents, and reports that 74 of 100 computational biology papers were successfully agentified in testing.
The tool, called Paper2Agent, wraps each paper's manuscript, code, data and workflows into a Model Context Protocol server that a chat agent can call as tools. For the AlphaGenome paper, the pipeline generated 22 MCP tools in about 45 minutes at a cost of $14, and all 22 passed validation without human intervention; for Scanpy it produced 7 tools in a similar time for $13. Across the full benchmark, the resulting agents scored 91.2 ± 1.6% accuracy on tutorial questions, against 80.3 ± 2.3% for a Claude + Repo baseline. On AlphaGenome the tutorial accuracy was 98.7 ± 1.3%, rising to 100% on novel queries.
The paper reports a downstream biological finding: two paper-derived agents, when connected, flagged a molecular variant near the MPHOSPH9 gene as associated with increased ADHD risk, a connection that, according to Zou, had not been reported before.
"For essentially all of human history, the way that we represent knowledge is in the form of these very passive artifacts," senior author James Zou said. "This is an opportunity to fundamentally reimagine what knowledge looks like."
The authors are careful about what their headline numbers mean. "Accuracy measures whether Paper2Agent faithfully executes a paper's workflows and matches expected outputs," the paper states. "It does not establish that the underlying scientific conclusion is correct." The 26% of papers that failed conversion did so because of "incomplete code, missing documentation and software environments that could not be resolved."
The team proposes journals introduce an "agent availability" section alongside existing data and code availability statements.
Shared on Bluesky by 5 AI experts
-
It would be fun to listen to a future episode of the Mystery AI Hype Theater 3000 podcast (that uses *ridicule as praxis") about this article 🤣 "Reimagining research papers as interactive and reliable AI agents" www.nat…
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Reimagining research papers as interactive and reliable AI agents - Nature