nature.com web signal

AMIE team: benchmarks alone can't earn clinical AI trust

TL;DR

  • A new Nature Medicine comment from the Google AMIE team argues clinical AI trust cannot be benchmarked into existence and must be earned prospectively.
  • The authors write that the hardest lessons in real-world deployment concern the humans and systems around the AI, not the technology itself.
  • In the companion AMIE feasibility trial at BIDMC, 100 adult patients completed pre-visit interactions; safety supervisors did not need to intervene and AMIE's differential included the final diagnosis in 90% of cases.

Trust in clinical AI, the authors of a new Nature Medicine comment write, "cannot be benchmarked into existence. It must be earned through rigorous prospective studies in real-world clinical settings."

The piece is by Mike Schaekermann of Google Research, Anil Palepu, Adam Rodman and colleagues at Google DeepMind, Harvard Medical School, Stanford and Beth Israel Deaconess Medical Center. It sits alongside the team's own companion feasibility trial of AMIE, the Articulate Medical Intelligence Explorer.

That companion study ran 100 adult patients through AMIE for pre-visit history taking at BIDMC. Human safety supervisors monitored every patient-AMIE session in real time and, in the trial team's words, "did not need to intervene." AMIE's differential included each patient's final diagnosis in roughly 90% of cases.

The comment's sharper claim is that what these studies expose is rarely the model. "The hardest lessons often concern the humans and systems around the AI, not the technology itself," the authors write. Two researchers we follow had it in circulation the day it went live.

Shared on Bluesky by 2 AI experts