nature.com web signal

AMIE team: benchmarks won't earn clinical AI trust, only trials

TL;DR

  • The Google-led AMIE team argues in Nature Medicine that trust in clinical AI must be earned in real-world prospective trials, not on benchmarks.
  • AMIE's first prospective feasibility trial at Beth Israel Deaconess ran 100 adult patients through pre-visit text chats with a physician AI supervisor watching every session live.
  • AMIE placed the final diagnosis in its top 7 in 90% of cases, top 3 in 75%, and as the single most likely pick in 56%.

"Trust in clinical artificial intelligence (AI) cannot be benchmarked into existence." That is the opening line of a commentary in Nature Medicine co-authored by the Google team behind AMIE with outside clinicians including Adam Rodman of Beth Israel Deaconess.

"It must be earned through rigorous prospective studies in real-world clinical settings," the authors write, "where the hardest lessons often concern the humans and systems around the AI, not the technology itself."

The piece runs alongside AMIE's first prospective feasibility trial, described in a companion Google research post: 100 adult patients at Beth Israel Deaconess Medical Center in Boston text-chatted with AMIE before primary-care visits, pre-registered at ClinicalTrials.gov as NCT06911398. A physician "AI supervisor" watched every session live over video and screen share, trained to break in on four pre-defined safety criteria. Zero safety stops fired across all interactions. AMIE placed the eventual final diagnosis in its top 7 in 90% of cases, top 3 in 75%, and as the single most likely pick in 56%.

The interface was text only. There was no controlled comparison group.

Shared on Bluesky by 2 AI experts