nature.com web signal

AMIE team: prospective trials, not benchmarks, earn AI trust

TL;DR

  • Google's AMIE team argues in Nature Medicine that clinical AI trust must come from prospective real-world trials, not benchmark scores.
  • A companion feasibility trial ran 100 adult patients through pre-visit chats at Beth Israel Deaconess from April to November 2025.
  • AMIE's differential included the final diagnosis in 90% of cases at 8-week chart review, with zero safety stops required.

Trust in clinical AI cannot be benchmarked into existence. That is the opening claim of a Nature Medicine comment from the team behind AMIE, Google's conversational diagnostic system, published September 14 alongside its first prospective feasibility trial.

"Trust in clinical artificial intelligence (AI) cannot be benchmarked into existence," Mike Schaekermann of Google Research and co-authors write. "It must be earned through rigorous prospective studies in real-world clinical settings, where the hardest lessons often concern the humans and systems around the AI, not the technology itself."

The companion trial ran at Beth Israel Deaconess Medical Center's Healthcare Associates clinic in Boston from April to November 2025. One hundred adult patients text-chatted with AMIE up to five days before their primary-care appointments. Board-certified internal medicine physicians watched every session in real time against four pre-specified safety criteria. "Zero safety stops were required by the human AI supervisors."

At 8-week chart review, AMIE's differential diagnosis included the final diagnosis in 90% of cases, 75% within the top three, and 56% as the single most likely pick. Primary-care providers still outperformed the model on the practicality (p=0.003) and cost-effectiveness (p=0.004) of management plans, and AMIE had no EHR access and no physical exam.

The comment is co-signed across Google Research, Google DeepMind, Harvard Medical School, Stanford and Included Health. Two researchers we follow posted the paper to their networks the day it went up. It reads less like an editorial and more like a public evidence standard: pre-registered, IRB-approved, single-center by design, and shadowed by physicians who could halt the conversation at any moment.

Shared on Bluesky by 2 AI experts