Google AMIE team calls prospective trials 'non-negotiable'
TL;DR
- A Nature Medicine commentary from Google's AMIE team argues clinical AI trust cannot be benchmarked into existence and requires prospective real-world studies.
- The paper it accompanies is a pre-registered feasibility trial of AMIE at Beth Israel Deaconess Medical Center with 100 adult patients.
- Under human supervision, zero safety stops fired; AMIE's top guess matched the final diagnosis in 56% of cases, its top seven in 90%.
"Trust in clinical artificial intelligence cannot be benchmarked into existence." That is how a group of Google researchers, DeepMind engineers and academic clinicians open a commentary in Nature Medicine, arguing that scoring well on medical exam datasets is not the same as being safe in a clinic.
The piece, from Mike Schaekermann, Anil Palepu, Adam Rodman and colleagues at Google Research, Google DeepMind, Beth Israel Deaconess Medical Center and Stanford, says prospective real-world evidence is now the non-negotiable step for conversational medical AI. "It must be earned through rigorous prospective studies in real-world clinical settings," they write, "where the hardest lessons often concern the humans and systems around the AI, not the technology itself."
Their own case study is a pre-registered, IRB-approved feasibility trial of AMIE at Beth Israel Deaconess Medical Center. One hundred adult patients text-chatted with the system before a scheduled primary-care visit. Ninety-eight kept the appointment. Human supervisors watched every session and could halt it under four pre-defined criteria, including immediate concern for self-harm or any potential clinical harm the supervisor identified.
None of those stops fired. AMIE's top single guess matched the eventual diagnosis in 56 percent of the 100 cases, and one of its top seven suggestions matched in 90 percent. Clinicians reported that the pre-visit chat "helped shift the visit dynamic from simple data gathering to data verification."
The commentary is written by the AMIE team itself, which is worth holding in mind when reading its message about benchmarks. The paper the authors are effectively defending is a 100-patient, single-site feasibility study, funded by Alphabet.
Shared on Bluesky by 2 AI experts
Originally reported by nature.com
Read the original article →Original headline: Prospective evidence for conversational medical AI is hard, but non-negotiable - Nature Medicine