AMIE authors call prospective trials for medical AI non-negotiable
TL;DR
- A Nature Medicine comment led by Google Research argues clinical AI trust cannot be earned through benchmarks alone, only through prospective real-world studies.
- The piece draws on a feasibility trial in which 100 adult patients chatted with Google's AMIE system before appointments at Beth Israel Deaconess, under live physician oversight.
- In that trial, zero safety stops were required, and AMIE's differential included the final diagnosis in 90% of cases with 75% top-3 accuracy.
Trust in clinical AI 'cannot be benchmarked into existence,' the authors of a new Nature Medicine comment write. Most of the listed authors work at Google Research or Google DeepMind, alongside clinicians including Adam Rodman of Beth Israel Deaconess Medical Center.
The comment, led by Mike Schaekermann, Anil Palepu and Po-Hsuan Cameron Chen of Google Research, states that trust 'must be earned through rigorous prospective studies in real-world clinical settings, where the hardest lessons often concern the humans and systems around the AI, not the technology itself.'
The piece draws on the same team's prospective feasibility study of AMIE, Google's conversational diagnostic system, at Beth Israel Deaconess. One hundred adult patients chatted with AMIE by text up to five days before their appointments while physician supervisors watched every session live via screen-share. Zero safety stops were required, and AMIE's differential diagnosis included the final diagnosis in 90% of cases, with 75% top-3 accuracy.
Disclosures sit alongside the argument: authors report Alphabet Inc. funding and potential equity, and one author is affiliated with Included Health.
Shared on Bluesky by 2 AI experts
Originally reported by nature.com
Read the original article →Original headline: Prospective evidence for conversational medical AI is hard, but non-negotiable - Nature Medicine