Google AMIE team: benchmarks can't earn clinical AI trust
TL;DR
- A Nature Medicine comment from Google's AMIE team argues clinical AI trust must be earned through prospective real-world studies, not benchmark leaderboards.
- In their Beth Israel Deaconess feasibility trial (NCT06911398), AMIE included the final diagnosis in 90% of 100 adult cases and hit 75% top-3.
- Physicians still beat AMIE on management practicality (p = 0.003) and cost-effectiveness (p = 0.004), with zero safety stops triggered.
"Trust in clinical artificial intelligence (AI) cannot be benchmarked into existence." That is the opening line of a comment in Nature Medicine from the team behind Google's AMIE conversational diagnostic system, arguing that leaderboards and retrospective evaluations do not add up to evidence a physician can act on.
The authors, from Google Research, Google DeepMind, Beth Israel Deaconess Medical Center, Harvard Medical School and Stanford, anchor the piece to their own prospective feasibility study (NCT06911398) at Beth Israel Deaconess. AMIE's differential diagnosis included the final diagnosis in 90% of 100 adult cases, with 75% top-3 accuracy and 56% single-diagnosis accuracy. Zero safety stops were required. Physicians still beat AMIE on management practicality (p = 0.003) and cost-effectiveness (p = 0.004).
The authors' own reading of that split: the hardest lessons "concern the humans and systems around the AI, not the technology itself."
That gap is the entire argument for prospective work. A model that names the right diagnosis can still recommend a plan the clinic cannot run or the patient cannot afford, and no benchmark score catches that. Two experts on our tracker have already shared the piece.
The commentary reads as a template for what prospective conversational-AI evidence ought to look like, and the authors describe the trial as pre-registered, IRB-approved and single-center by design.
Shared on Bluesky by 2 AI experts
Originally reported by nature.com
Read the original article →Original headline: Prospective evidence for conversational medical AI is hard, but non-negotiable - Nature Medicine