AMIE team says clinical AI trust demands real-world trials
TL;DR
- In a Beth Israel Deaconess feasibility trial, AMIE included the final diagnosis inside its top seven possibilities in 90% of 100 adult cases.
- Human supervisors watched every AMIE text-chat visit on video and triggered zero safety stops.
- Physicians still beat AMIE on management practicality (p = 0.003) and cost-effectiveness (p = 0.004).
The team behind Google's AMIE conversational medical AI argues in a new Nature Medicine commentary that benchmark scores are not enough. "Trust in clinical artificial intelligence (AI) cannot be benchmarked into existence," the authors write. "It must be earned through rigorous prospective studies in real-world clinical settings, where the hardest lessons often concern the humans and systems around the AI, not the technology itself."
The piece, led by Google Research's Mike Schaekermann with Anil Palepu, Adam Rodman and colleagues at Google DeepMind, Beth Israel Deaconess Medical Center, Stanford and Harvard Medical School, is paired with a feasibility trial run at Beth Israel Deaconess's Healthcare Associates clinic in Boston from April to November 2025. One hundred adult patients completed a pre-visit text chat with AMIE. The system's differential diagnosis included the final diagnosis inside its top seven possibilities in 90% of cases, hit the top three in 75%, and named the single most likely diagnosis in 56%. Physician supervisors watched every session over video and triggered zero safety stops.
Even so, AMIE did not eclipse the humans across the board. Clinicians beat the model on management practicality (p = 0.003) and cost-effectiveness (p = 0.004). The authors, who built the system they are studying, point to the humans and workflow around the AI as the harder frontier.
Two of the researchers we follow shared the paper the day it appeared. The trial is pre-registered as NCT06911398, and the commentary pitches it as the shape prospective evidence should take: pre-registered, IRB-approved, single-center.
Shared on Bluesky by 2 AI experts
Originally reported by nature.com
Read the original article →Original headline: Prospective evidence for conversational medical AI is hard, but non-negotiable - Nature Medicine