nature.com web signal

Google's AMIE team: clinical AI trust can't be benchmarked in

TL;DR

  • A Nature Medicine commentary from Google's AMIE team argues clinical AI trust must be earned through prospective real-world trials, not benchmark scores.
  • The companion feasibility study ran 100 patients through AMIE at Beth Israel Deaconess between April and November 2025, with physician supervisors watching every chat.
  • AMIE's differential diagnosis included the final diagnosis in 90% of top-7 lists and triggered zero safety stops across the trial.

The Google and Google DeepMind team behind AMIE, the company's conversational diagnostic AI, has published a commentary in Nature Medicine arguing that leaderboard scores are no longer enough to justify putting a chat-based doctor in front of real patients.

"Trust in clinical artificial intelligence (AI) cannot be benchmarked into existence," Mike Schaekermann and thirteen co-authors write. "It must be earned through rigorous prospective studies in real-world clinical settings, where the hardest lessons often concern the humans and systems around the AI, not the technology itself."

The argument is not abstract. The same authors, working with Beth Israel Deaconess Medical Center and Harvard Medical School, ran a prospective single-arm feasibility study of AMIE at BIDMC's Healthcare Associates primary care clinic in Boston between April 2025 and November 2025. One hundred adult patients were enrolled and 98 completed both an AMIE text-chat session, held up to five days before their scheduled appointment, and the appointment itself. Every AMIE session was monitored in real time by board-certified physicians designated as "AI supervisors" with the authority to intervene.

Across all 100 interactions, "zero safety stops were required by the group of AI supervisors overseeing these interactions." AMIE's differential diagnosis included the eventual final diagnosis 90% of the time in its top-7 list, 75% in the top-3 and 56% in the top-1. Blinded evaluators found "no significant difference in the overall quality of the differential diagnosis and management plan from AMIE versus PCPs," though the primary care physicians scored significantly better on cost-effectiveness (p=0.004) and practicality (p=0.003) of the plans they wrote. Of the 44 PCPs who reviewed AMIE transcripts, one said the notes read like "a third-year medical student."

The commentary landed with two AI researchers we track posting it within a day. It is co-signed by Adam Rodman of Beth Israel Deaconess and Harvard alongside the Google DeepMind names, giving the methodological argument an academic voice from outside the vendor.

Shared on Bluesky by 2 AI experts