nature.com web signal

AMIE authors: benchmarks can't earn medical AI trust in clinics

TL;DR

  • Google Research, DeepMind, Harvard and Stanford authors argue in Nature Medicine that clinical AI trust must come from prospective real-world studies, not benchmarks.
  • AMIE's feasibility trial saw zero safety-supervisor interventions across 100 adult patients text-chatting up to 5 days before their appointment at Beth Israel Deaconess.
  • AMIE matched primary-care providers on differential quality (p = 0.6) but lost on management-plan practicality (p = 0.003) and cost effectiveness (p = 0.004).

Trust in clinical AI cannot be benchmarked into existence, a group of Google Research, Google DeepMind, Harvard and Stanford authors argue in a new Nature Medicine Comment. It must be earned, they write, "through rigorous prospective studies in real-world clinical settings, where the hardest lessons often concern the humans and systems around the AI, not the technology itself."

The argument does not stand alone. It arrives with a prospective feasibility study of AMIE, the Articulate Medical Intelligence Explorer, at Beth Israel Deaconess Medical Center. According to Google's writeup of the trial, 100 adult patients text-chatted with the system up to 5 days before their appointment. Human safety supervisors monitored every consultation and "did not need to intervene to stop any consultations based on pre-defined criteria."

The diagnostic scores land in AMIE's favor. Per chart review 8 weeks post-encounter, AMIE's differential diagnosis included the final diagnosis in 90% of cases, with 75% top-3 accuracy. Blinded assessment showed no significant difference from primary-care providers on overall differential quality (p = 0.6) or management-plan safety (p = 1.0).

Where AMIE lost was practical. "PCPs outperformed AMIE in the practicality (p = 0.003) and cost effectiveness (p = 0.004) of Mx," the abstract reports.

Mike Schaekermann, Research Lead at Google Research, is among the authors; two researchers on our Who's Who radar shared the piece within a day of publication. The study itself was single-center, text-only, with no controlled comparison group.

Shared on Bluesky by 2 AI experts