AMIE team: prospective trials non-negotiable for clinical AI
TL;DR
- Google's AMIE authors argue in Nature Medicine that trust in clinical AI must come from prospective trials, not benchmark scores.
- Their Beth Israel Deaconess feasibility study of 100 adult patients ran April to November 2025 with zero safety stops.
- AMIE's differential diagnosis included the final diagnosis in 90% of cases, but physicians beat it on management practicality and cost-effectiveness.
The Google-led team behind AMIE, the company's conversational diagnostic AI, argues in Nature Medicine that leaderboard scores prove almost nothing about whether such a system should be trusted with a patient. "Trust in clinical artificial intelligence (AI) cannot be benchmarked into existence," write Mike Schaekermann, Adam Rodman and colleagues from Google Research, Google DeepMind, Harvard Medical School and Stanford.
The commentary sits alongside the group's own prospective feasibility study at Beth Israel Deaconess Medical Center's Healthcare Associates clinic in Boston, which enrolled 100 adult patients between April and November 2025. Ninety-eight completed appointments after chatting with AMIE via secure text; a physician supervisor watched by live video, with pre-defined criteria for harm, distress, or a patient request to stop. Zero safety stops were triggered.
The Google Research write-up of the trial reports that AMIE's differential diagnosis included the final diagnosis in 90% of cases, with 75% top-3 accuracy. Primary care physicians and AMIE were comparable on differential-diagnosis and management-plan quality, but the PCPs still beat the model on management practicality (p = 0.003) and cost-effectiveness (p = 0.004).
The authors' takeaway is the harder part. Prospective validation, they say, surfaces problems benchmarks miss because "the hardest lessons often concern the humans and systems around the AI, not the technology itself." Two of the health-AI researchers we track flagged the paper the day it went live.
Shared on Bluesky by 2 AI experts
Originally reported by nature.com
Read the original article →Original headline: Prospective evidence for conversational medical AI is hard, but non-negotiable - Nature Medicine