AMIE team: benchmarks alone can't earn medical AI's trust
TL;DR
- A Nature Medicine commentary from Google's AMIE team argues clinical AI trust must be earned in prospective real-world trials, not benchmark scores.
- The accompanying AMIE feasibility trial ran 100 patient chats at Beth Israel Deaconess from April to November 2025 with zero safety stops.
- AMIE's differential included the final diagnosis in 90% of cases; physicians still beat it on treatment practicality (p=0.003) and cost (p=0.004).
Trust in clinical AI cannot be benchmarked into existence. That is the flat opening line of a new Nature Medicine commentary from the team behind AMIE, Google's conversational diagnostic system.
"It must be earned through rigorous prospective studies in real-world clinical settings," the authors write, "where the hardest lessons often concern the humans and systems around the AI, not the technology itself." Lead authors Mike Schaekermann, Anil Palepu, Adam Rodman, and Ami Parekh sit across Google Research, Google DeepMind, Harvard Medical School, and Stanford.
The trial the commentary discusses ran at Beth Israel Deaconess Medical Center's Healthcare Associates clinic in Boston from April 2025 to November 2025. One hundred adult patients text-chatted with AMIE up to five days before their primary-care appointments. Seven board-certified internal medicine physicians monitored every session in real time against four pre-specified stop criteria: self-harm concern, significant emotional distress, physician-identified clinical harm, or an explicit patient request to end.
Zero safety stops were triggered. AMIE's differential diagnosis included the final diagnosis in 90% of cases, with 75% top-3 and 56% top-1 accuracy. On blinded rater review, AMIE was indistinguishable from primary-care physicians on differential quality (p=0.6), management-plan appropriateness (p=0.1), and management-plan safety (p=1.0). Physicians still beat AMIE on treatment practicality (p=0.003) and cost-effectiveness (p=0.004).
The cohort skewed young and English-speaking. 51% of patients were under 50 and 86% reported English as their home language. The pre-registration is NCT06911398. The model was Gemini 2.5 Pro, swapped to Flash after 50 encounters for latency.
The commentary's argument is that a leaderboard score is not what earns an LLM its clinic keys; a pre-registered study with humans watching every session is. Two of the AI researchers on our radar shared it the day it went live.
Shared on Bluesky by 2 AI experts
Originally reported by nature.com
Read the original article →Original headline: Prospective evidence for conversational medical AI is hard, but non-negotiable - Nature Medicine