nature.com web signal

AMIE authors: clinical AI trust demands prospective trials

TL;DR

  • A Nature Medicine commentary from the Google AMIE team argues clinical AI credibility must come from prospective real-world studies, not leaderboard scores.
  • In their Beth Israel Deaconess feasibility trial, AMIE's differential covered the final diagnosis 90% at top-7, 75% at top-3 and 56% at top-1.
  • Primary-care physicians rated comparably on diagnosis quality but outperformed AMIE on the practicality and cost-effectiveness of management plans.

"Trust in clinical artificial intelligence (AI) cannot be benchmarked into existence. It must be earned through rigorous prospective studies in real-world clinical settings." That is the opening premise of a commentary in Nature Medicine from the Google Research, Google DeepMind, Harvard Medical School, Beth Israel Deaconess and Stanford team behind AMIE, the conversational diagnostic system.

The authors — Mike Schaekermann, Anil Palepu, Adam Rodman, Ami Parekh, Ethan Goh and colleagues — draw the argument from their own feasibility trial at Beth Israel Deaconess. One hundred adult patients chatted with AMIE before a primary-care visit; 98 attended their scheduled appointment. The model's differential included the eventual final diagnosis in 90% of cases at top-7, 75% at top-3 and 56% at top-1. Zero safety interventions were required.

But the same trial found something more instructive. Clinical evaluators rated AMIE and the primary-care physicians comparably on differential diagnosis quality and management-plan safety. The PCPs still beat AMIE on the practicality and cost-effectiveness of management plans. AMIE had no access to the electronic health record, no physical exam, no visual patient assessment.

That, the commentary argues, is the point: "the hardest lessons often concern the humans and systems around the AI, not the technology itself."

Shared on Bluesky by 2 AI experts