nature.com web signal

AMIE team: benchmarks can't earn conversational medical AI trust

TL;DR

  • A Nature Medicine commentary from Google's AMIE team argues benchmark scores cannot substitute for prospective real-world trials of conversational medical AI.
  • In a Beth Israel feasibility trial, AMIE matched the final diagnosis 90% at top-7, 75% at top-3, and 56% at top-1 on chart review.
  • A blinded review found AMIE and primary care physicians produced similar-quality plans; the physicians beat the model on practicality and cost.

"Trust in clinical artificial intelligence cannot be benchmarked into existence." That is the opening line of a Nature Medicine commentary from the 14-author team behind AMIE, Google's conversational diagnostic AI. The piece continues that trust "must be earned through rigorous prospective studies in real-world clinical settings, where the hardest lessons often concern the humans and systems around the AI, not the technology itself."

The commentary lands with an accompanying feasibility trial. Between April and November 2025 at Beth Israel Deaconess Medical Center's Healthcare Associates clinic in Boston, 100 adults chatted with AMIE by text-chat up to five days before their scheduled appointments. Ninety-eight attended. Human safety supervisors monitored every session in real time. Zero safety stops were required.

On chart review eight weeks after each visit, AMIE's differential diagnosis contained the final diagnosis 90% of the time within its top seven, 75% within its top three, and named it most likely in 56% of cases. A blinded assessment "suggested similar overall quality between AMIE and PCPs, without significant differences." The primary care physicians still beat the model on practicality and cost of management plans.

The commentary is signed by the AMIE team itself: Adam Rodman at Beth Israel Deaconess Medical Center and Harvard Medical School and Ami Parekh at Included Health alongside twelve Google Research and DeepMind authors including Mike Schaekermann, Shekoofeh Azizi, Joëlle Barral and Yossi Matias. A perspective from that particular team, arguing benchmark scores are not enough, is the news.

The commentary is a call, not new evidence on its own. The feasibility trial reports diagnostic overlap and safety at eight weeks; it does not report downstream treatment decisions or clinical outcomes. Two researchers we track had the piece circulating within a day of publication.

Shared on Bluesky by 2 AI experts