Google AMIE team: only prospective trials earn clinical AI trust
TL;DR
- The AMIE authors argue in Nature Medicine that leaderboard scores cannot substitute for prospective real-world clinical trials of conversational medical AI.
- Their companion feasibility trial ran 100 adult patients at Beth Israel Deaconess from April to November 2025 with live physician oversight and zero safety stops.
- AMIE reached 90% top-seven diagnostic accuracy but was outperformed by primary-care physicians on treatment practicality (p=0.003) and cost-effectiveness (p=0.004).
The team behind Google's AMIE conversational diagnostic system argues in Nature Medicine that leaderboard scores alone cannot certify medical AI for the clinic. "Trust in clinical artificial intelligence (AI) cannot be benchmarked into existence," the authors write. "It must be earned through rigorous prospective studies in real-world clinical settings, where the hardest lessons often concern the humans and systems around the AI, not the technology itself."
The commentary carries fourteen authors across Google Research, Google DeepMind, Harvard Medical School and Stanford, including Mike Schaekermann, Anil Palepu and Adam Rodman. It sits alongside the group's own prospective feasibility trial of AMIE, run at Beth Israel Deaconess Medical Center's Healthcare Associates clinic in Boston from April 2025 to November 2025. One hundred adult patients text-chatted with the system up to five days before their primary-care appointments while a supervising physician watched every session on live video with screen sharing.
Zero safety stops were required by the human AI supervisors across four pre-specified safety criteria.
The clinical numbers are not uniformly flattering. AMIE's differential contained the eventual diagnosis within its top seven answers in roughly 90% of cases, with 75% top-three accuracy, and blinded reviewers found no significant difference in differential quality, management or safety between AMIE and the primary-care providers. Those same providers, however, outperformed the model on treatment practicality (p=0.003) and cost-effectiveness (p=0.004), a gap the group flags as the kind of finding a benchmark score would not have surfaced.
The work was funded by Alphabet Inc.
Shared on Bluesky by 2 AI experts
Originally reported by nature.com
Read the original article →Original headline: Prospective evidence for conversational medical AI is hard, but non-negotiable - Nature Medicine