nature.com web signal

Google's AMIE team: benchmarks alone can't earn clinical AI trust

TL;DR

  • Google, DeepMind, Harvard and Stanford authors argue in Nature Medicine that trust in clinical AI must come from prospective real-world trials, not benchmarks.
  • Companion AMIE feasibility trial at Beth Israel Deaconess with 100 patients from April to November 2025 required zero safety stops under live physician oversight.
  • AMIE hit 90% top-7 and 75% top-3 diagnostic accuracy, but primary care providers beat it on practicality (p=0.003) and cost-effectiveness (p=0.004).

"Trust in clinical artificial intelligence (AI) cannot be benchmarked into existence," authors from Google Research, Google DeepMind, Harvard Medical School and Stanford write in a Nature Medicine commentary. It must be earned, they argue, through rigorous prospective studies in real-world clinical settings.

The argument sits alongside a companion feasibility trial of AMIE, Google's conversational diagnostic model, run at Beth Israel Deaconess Medical Center's Healthcare Associates clinic in Boston from April 2025 to November 2025. One hundred adult patients interacted with AMIE by text-chat up to five days before their primary care appointment, and dedicated physicians monitored every session in real time via video. Zero safety stops were required. The system ran on Gemini 2.5 Pro, later switched to Gemini 2.5 Flash after 50 encounters, with no domain-specific fine-tuning.

AMIE's differential diagnosis included the final diagnosis in 90% of cases, with 75% top-3 accuracy. Blinded clinical evaluators rated AMIE and primary care providers similarly for differential diagnosis quality (p=0.6) and management plan safety (p=1.0). But PCPs beat AMIE on practicality of management (p=0.003) and cost-effectiveness (p=0.004).

The commentary's framing is unambiguous: "the hardest lessons often concern the humans and systems around the AI, not the technology itself."

Shared on Bluesky by 2 AI experts