AMIE team: benchmarks alone can't validate clinical AI
TL;DR
- Nature Medicine commentary from Google's AMIE team argues clinical AI trust must be earned through prospective real-world trials, not benchmark scores alone.
- The companion feasibility study at Beth Israel Deaconess ran 100 AMIE patient chats from April to November 2025 with zero safety stops required.
- AMIE hit 90% top-7 and 75% top-3 diagnostic accuracy, but PCPs beat the system on practicality (p=0.003) and cost-effectiveness (p=0.004).
"Trust in clinical artificial intelligence (AI) cannot be benchmarked into existence," Mike Schaekermann and colleagues write in a Nature Medicine commentary published alongside a companion feasibility trial of Google's AMIE conversational diagnostic system. Trust, they argue, "must be earned through rigorous prospective studies in real-world clinical settings, where the hardest lessons often concern the humans and systems around the AI, not the technology itself."
The companion study ran at Beth Israel Deaconess Medical Center's Healthcare Associates clinic in Boston from April 2025 to November 2025. One hundred adult patients text-chatted with AMIE up to five days before their primary-care appointments; seven board-certified internal medicine physicians supervised every session in real time against four pre-specified safety criteria (self-harm concern, significant emotional distress, physician-identified clinical harm, or explicit patient request to end). Zero safety stops were required across all 100 interactions.
At an 8-week chart review, AMIE's differential diagnosis list included the final diagnosis in 90% of cases (top-7), with 75% top-3 and 56% top-1 accuracy. Blinded comparative ratings showed no significant difference between AMIE and the clinic's PCPs on differential quality, management plan appropriateness, or plan safety. Primary-care providers still beat the system on practicality (p=0.003) and cost-effectiveness (p=0.004).
The author list runs across Google Research, Google DeepMind, Harvard Medical School and Beth Israel Deaconess, and Stanford. Two of the clinicians we track in our Who's Who circulated the piece the day it appeared.
Shared on Bluesky by 2 AI experts
Originally reported by nature.com
Read the original article →Original headline: Prospective evidence for conversational medical AI is hard, but non-negotiable - Nature Medicine