nature.com web signal

Google AMIE team: benchmarks alone can't earn clinical AI trust

TL;DR

  • A Nature Medicine commentary from the Google, DeepMind, Harvard and Stanford group behind AMIE argues clinical AI trust must be earned through prospective real-world studies, not benchmarks.
  • The companion feasibility trial at Beth Israel Deaconess enrolled 100 adults; AMIE's differential included the final diagnosis in 90% of cases, with 75% top-3 accuracy.
  • Human supervisors triggered zero safety stops, but primary care physicians still beat AMIE on management plan practicality and cost-effectiveness.

The Google Research and DeepMind team behind AMIE, the conversational diagnostic model, has used a Nature Medicine commentary published September 14 to argue that leaderboard scores cannot earn clinical trust. "Trust in clinical artificial intelligence (AI) cannot be benchmarked into existence," the authors, led by Mike Schaekermann, write. "It must be earned through rigorous prospective studies in real-world clinical settings, where the hardest lessons often concern the humans and systems around the AI, not the technology itself."

The comment lands alongside a feasibility trial of AMIE at Beth Israel Deaconess Medical Center in Boston, described in Google Research's own write-up. One hundred adults were enrolled and 98 completed text-chat visits with the model under live physician supervision. Human supervisors triggered no safety stops. AMIE's differential diagnosis contained the eventual final diagnosis in 90% of cases, with 75% top-3 accuracy and 56% for the single most likely condition. Two researchers we track shared the paper the week it appeared.

Primary care physicians were rated similarly to AMIE on differential quality and on the appropriateness and safety of their management plans. The PCPs, however, produced plans judged more practical and more cost-effective. The trial was single-center and text-only, limits the authors flag themselves.

Shared on Bluesky by 2 AI experts