nature.com web signal

AMIE authors: benchmarks alone can't earn clinical AI trust

TL;DR

  • The AMIE team ran a prospective feasibility trial at Beth Israel Deaconess Medical Center's Healthcare Associates clinic from April 2025 to November 2025 with 100 adult patients.
  • AMIE's differential included the final diagnosis in 90% of cases and hit 75% top-3 accuracy across the cohort, with zero safety stops flagged by live physician supervisors.
  • Primary-care providers still beat AMIE on treatment practicality (p = 0.003) and cost-effectiveness (p = 0.004), which the authors flag as the human-and-systems side of the evaluation.

The team behind Google's AMIE diagnostic chatbot ran their system past 100 adult patients at Beth Israel Deaconess Medical Center's Healthcare Associates clinic between April 2025 and November 2025, with a physician on a video call monitoring every text-chat session for four pre-specified safety triggers. The trial was pre-registered as NCT06911398. Zero safety stops were called.

That feasibility study is the evidence base under a new comment in Nature Medicine by Mike Schaekermann, Anil Palepu, Adam Rodman, Ami Parekh, Ethan Goh and co-authors from Google Research, Google DeepMind, Beth Israel Deaconess, Harvard Medical School, Stanford and Included Health. Their line is flat: "Trust in clinical artificial intelligence cannot be benchmarked into existence. It must be earned through rigorous prospective studies in real-world clinical settings, where the hardest lessons often concern the humans and systems around the AI, not the technology itself."

The numbers, from the companion Google Research write-up of the same trial: AMIE's differential included the final diagnosis in 90% of cases, top-3 accuracy was 75%, and AMIE was accurate as the single most likely diagnosis in 56%. Ninety-eight of the 100 patients actually kept their appointment. Primary-care providers matched AMIE on overall differential quality and management-plan safety, but beat it on management practicality (p = 0.003) and cost-effectiveness (p = 0.004). The authors note AMIE had no EHR access, no physical exam, and no multimodal input.

The comment is short on triumphalism about the accuracy numbers and long on the argument that this kind of trial — pre-registered, live-supervised, in an actual clinic — is the floor, not a nice-to-have.

Shared on Bluesky by 2 AI experts