nature.com web signal

Paper2Agent turns 74 of 100 biology papers into AI agents

TL;DR

  • Paper2Agent turned 74 of 100 computational biology papers into working AI agents; 593 of 599 proposed tools passed validation.
  • Agent-mediated queries hit 91.2% accuracy on 300 tutorial questions versus 80.3% for Claude reading the raw repository.
  • Three paper agents collaborated to prioritize GPR137 as a probable causal gene for psoriasis at the rs887314 locus.

A paper is normally something you read. The framework laid out in a Nature paper from Jiacheng Miao, James Zou and colleagues makes it something you can ask questions of.

Paper2Agent is an automated pipeline that ingests a research paper and its codebase and produces a Model Context Protocol server exposing the paper's tools, datasets and workflows to a chat agent like Claude Code. The authors ran it on 100 computational biology papers; 74 came out as working agents. Of 599 tools those agents proposed, 593 validated. On a benchmark of 300 tutorial-derived questions, the agent-mediated setup scored 91.2% accuracy against 80.3% for Claude reading the raw repository, at $0.20 per query versus $0.38 and 1.6 minutes versus 4.3.

The framing the authors reach for is "lowering the barriers between publication and practical application." A single-paper agent built from AlphaGenome hit 98.7% on tutorial queries and 82.7% on open-ended ones, and generation itself took about 45 minutes and roughly $13 to $14 per paper.

The concrete pitch is a discovery, not just a benchmark. Three agents built from separate papers, one on AlphaGenome, an MPRA-coupled scCRISPRi study, and a CD4+ T cell Perturb-seq study, were pointed at psoriasis genetics and prioritized GPR137 at the rs887314 locus using what the paper calls "a new data integration approach...not proposed in the source papers." The correlation held only under stimulated T cell conditions (Spearman ρ = 0.613, P = 3.79 × 10⁻³). The link came in from five researchers on our Who's Who list.

The failure mode is stated plainly: 26% of the biology papers could not be agentified because of incomplete codebases, missing documentation or environment failures, which the authors read as a signal about which papers were actually reproducible to begin with. Open-ended scientific reasoning stays "human-in-the-loop"; the agents propose hypotheses, researchers pick directions.

Shared on Bluesky by 5 AI experts