nature.com web signal

Google's AMIE team says medical AI trust needs prospective trials

TL;DR

  • Google's AMIE authors argue in Nature Medicine that clinical AI trust must be earned through prospective real-world studies, not benchmarks.
  • A 100-patient feasibility trial at Beth Israel Deaconess ran text-chat visits under live physician safety supervision, with zero safety stops triggered.
  • AMIE's differential included the final diagnosis in 90% of cases; PCPs still beat it on practicality and cost effectiveness of management plans.

Google's AMIE team, writing in a Nature Medicine comment, argues that trust in clinical AI has to be earned in real clinics rather than on leaderboards. "Trust in clinical artificial intelligence (AI) cannot be benchmarked into existence. It must be earned through rigorous prospective studies in real-world clinical settings," the authors write, adding that "the hardest lessons often concern the humans and systems around the AI, not the technology itself."

The commentary accompanies a prospective feasibility trial of AMIE, Google's conversational diagnostic system, run at Beth Israel Deaconess Medical Center's Healthcare Associates clinic in Boston from April to November 2025. One hundred adult patients text-chatted with the model up to five days before their primary care appointments. Physicians served as real-time safety supervisors on every session under four pre-specified stop criteria, and the accompanying Google Research write-up reports that "zero safety stops were required by the human AI supervisors."

AMIE's differential included the final diagnosis in 90% of cases when the list ran to seven possibilities, and its single most likely diagnosis was correct in 56%. Clinicians rated the model and their own diagnostic reasoning similarly on differential quality. But primary care providers still beat AMIE on the "practicality and cost effectiveness" of the management plan. The chatbot ran with no electronic health record access and could not perform a physical exam.

Lead authors Mike Schaekermann, Anil Palepu and Adam Rodman, joined by colleagues across Google Research, Google DeepMind, Stanford and Harvard Medical School, pitch the study (NCT06911398) as a template for what prospective evidence should look like: pre-registered, IRB-approved, single-center by design. Neither the comment nor the companion write-up publishes false-positive rates for the safety triggers, or how the numbers hold outside a supervised primary-care setting.

Shared on Bluesky by 2 AI experts