nature.com web signal

AMIE team: prospective trials non-negotiable for medical AI

TL;DR

  • A Nature Medicine commentary from the AMIE team argues clinical AI trust cannot be benchmarked into existence and must come from real-world prospective studies.
  • In their own Beth Israel Deaconess feasibility trial, AMIE's differential diagnosis included the final diagnosis in 90% of 100 adult cases, with 75% top-3 accuracy.
  • Physicians still beat AMIE on management practicality (p = 0.003) and cost-effectiveness (p = 0.004); the authors say the hardest lessons are human and system factors.

A commentary in Nature Medicine from the Google, DeepMind, Harvard and Stanford group behind the AMIE conversational diagnostic system argues that clinical trust in these tools has to be earned in real hospitals, not on benchmark leaderboards.

"Trust in clinical artificial intelligence cannot be benchmarked into existence," the authors write in Nature Medicine. "It must be earned through rigorous prospective studies in real-world clinical settings, where the hardest lessons often concern the humans and systems around the AI, not the technology itself."

The comment, published September 14, is signed by Mike Schaekermann, Anil Palepu, Adam Rodman, Ami Parekh, Ethan Goh and colleagues at Google Research, Google DeepMind, Harvard Medical School and Stanford. It draws on the same group's prospective feasibility study of AMIE at Beth Israel Deaconess Medical Center, where 100 adults chatted with the system before urgent-care visits under live physician oversight. In that trial, "AMIE's differential diagnosis (DDx) included the final diagnosis in 90% of cases, with 75% top-3 accuracy," and no human safety supervisor had to halt a consultation.

Physicians still outperformed AMIE on management practicality (p = 0.003) and cost-effectiveness (p = 0.004). Two of the researchers we follow in our Who's Who directory shared the piece within days of publication.

Shared on Bluesky by 2 AI experts