nature.com web signal

AMIE team: prospective trials non-negotiable for clinical AI

TL;DR

  • Google's AMIE team argues in Nature Medicine that trust in clinical AI cannot be earned through benchmark scores alone.
  • The accompanying AMIE feasibility trial at Beth Israel Deaconess ran April to November 2025 with 100 adult patients and zero safety stops.
  • AMIE hit 90% diagnostic accuracy and 75% top-3, but primary-care providers still beat it on treatment practicality and cost.

The team behind Google's AMIE conversational diagnostic system argues in Nature Medicine that trust in clinical AI cannot be built out of benchmark leaderboards, only out of prospective trials in real clinics.

The commentary, signed by Mike Schaekermann, Anil Palepu, Adam Rodman, Ami Parekh, Ethan Goh and colleagues from Google Research, Google DeepMind, Harvard Medical School and Stanford, states the case plainly: "Trust in clinical artificial intelligence (AI) cannot be benchmarked into existence. It must be earned through rigorous prospective studies in real-world clinical settings, where the hardest lessons often concern the humans and systems around the AI, not the technology itself."

It lands alongside a feasibility trial of AMIE at Beth Israel Deaconess Medical Center's Healthcare Associates clinic in Boston, which ran from April 2025 to November 2025. One hundred adult patients text-chatted with the system up to five days before their appointments, with human safety supervisors monitoring every session in real time. AMIE's differential diagnosis included the final diagnosis in 90% of cases at 8-week chart review, with 75% top-3 accuracy and zero safety stops across four pre-specified safety criteria.

The caveat is on the same page. Primary-care providers still outperformed AMIE on treatment practicality (p=0.003) and cost-effectiveness (p=0.004), and the study is single-site and industry-authored. Two of the researchers we track in medical AI shared the piece the day it appeared.

Shared on Bluesky by 2 AI experts