AMIE team says clinical AI trust needs prospective trials
TL;DR
- A Nature Medicine commentary from Google's AMIE team argues clinical AI trust must be earned through prospective real-world studies, not benchmarks.
- The AMIE feasibility trial at Beth Israel Deaconess enrolled 100 patients under pre-registered protocol NCT06911398 with zero safety stops.
- AMIE's top-7 differential included the final diagnosis in 90% of cases, but primary care physicians beat it on management practicality and cost.
"Trust in clinical artificial intelligence (AI) cannot be benchmarked into existence." That is the opening of a commentary in Nature Medicine from the team behind AMIE, Google's conversational diagnostic system. The full sentence continues: "It must be earned through rigorous prospective studies in real-world clinical settings, where the hardest lessons often concern the humans and systems around the AI, not the technology itself."
The 14-author byline, led by Mike Schaekermann, draws from Google Research, Google DeepMind, Harvard Medical School, Beth Israel Deaconess Medical Center, Stanford and Included Health.
The argument leans on the group's own feasibility study at Beth Israel Deaconess, run under pre-registered protocol NCT06911398. The top-7 differential AMIE produced included the final diagnosis in 90% of the 100 patients enrolled, and named it as the single most likely option in 56%. No interaction triggered a safety stop.
Then the harder result. Blinded assessors judged diagnostic quality between AMIE and the primary care physicians similar overall. But, in the trial team's phrasing, "PCPs outperformed AMIE in the practicality and cost effectiveness of Mx plans." Two researchers we follow shared the commentary the day it posted.
Shared on Bluesky by 2 AI experts
Originally reported by nature.com
Read the original article →Original headline: Prospective evidence for conversational medical AI is hard, but non-negotiable - Nature Medicine