Google's AMIE team: benchmarks alone can't validate medical AI
TL;DR
- A Nature Medicine commentary by Google Research, DeepMind, Beth Israel Deaconess, Stanford and Included Health authors argues clinical AI trust must be earned through prospective real-world studies, not benchmarks.
- In an AMIE feasibility trial at Beth Israel Deaconess, 100 adult patients chatted with the AI before primary care visits, with zero safety stops required across the run.
- AMIE's differential diagnosis included the final diagnosis in 90% of cases, hit top-three accuracy 75% of the time, and named the correct top pick in 56%.
"Trust in clinical artificial intelligence (AI) cannot be benchmarked into existence." That is the opening argument of a Nature Medicine commentary published September 14 by 14 authors drawn from Google Research, Google DeepMind, Beth Israel Deaconess Medical Center, Stanford, and Included Health.
The authors continue that trust "must be earned through rigorous prospective studies in real-world clinical settings, where the hardest lessons often concern the humans and systems around the AI, not the technology itself."
The reference case is the group's own prospective, pre-registered, IRB-approved feasibility trial of Google's AMIE conversational diagnostic system, run at Beth Israel Deaconess. Google Research documented that 100 adult patients completed pre-visit text-chats with AMIE while a physician supervisor watched by live video, with clear stop criteria for harm, distress, or a patient request to end. Zero safety stops were required across the run.
On accuracy, AMIE's differential diagnoses included the final diagnosis in 90% of cases, reached top-three accuracy 75% of the time, and named the single most likely diagnosis correctly in 56%. Primary care providers still outperformed the system on cost-effectiveness and practicality of the management plan. The commentary link was moving through the medical-AI accounts on our radar within a day of publication.
Shared on Bluesky by 2 AI experts
Originally reported by nature.com
Read the original article →Original headline: Prospective evidence for conversational medical AI is hard, but non-negotiable - Nature Medicine