AMIE team: benchmarks alone can't earn medical AI trust
TL;DR
- AMIE's differential diagnosis contained the eventual diagnosis in 90% of cases and placed it in the top three in 75%, per the accompanying feasibility trial.
- One hundred adult patients used AMIE by text-chat up to five days before their appointments at a Boston clinic between April and November 2025.
- Human AI supervisors triggered zero safety stops across all interactions, despite four pre-specified safety criteria being in place.
Trust in clinical AI cannot be benchmarked into existence.
That is the flat opening of a Nature Medicine commentary from the Google team behind AMIE, the company's conversational diagnostic system. "It must be earned through rigorous prospective studies in real-world clinical settings," the authors write, "where the hardest lessons often concern the humans and systems around the AI, not the technology itself." Two of the medical-AI researchers we track posted the piece within days of its release.
The comment accompanies a feasibility trial of AMIE at Beth Israel Deaconess Medical Center's Healthcare Associates clinic in Boston. Between April and November 2025, one hundred adult patients interacted with AMIE by text-chat up to five days before their scheduled appointments. AMIE's differential diagnosis contained the eventual diagnosis in 90% of cases and placed it in the top three in 75%. Human AI supervisors triggered zero safety stops across all interactions, despite four pre-specified safety criteria being in place.
The trial team described AMIE as having "helped shift the visit dynamic from simple data gathering to data verification, allowing for more collaborative conversations and shared decision-making." Named authors on the commentary include Mike Schaekermann of Google Research, alongside co-authors at Google DeepMind, Stanford, and Beth Israel Deaconess Medical Center.
One clinic, one hundred patients.
Shared on Bluesky by 2 AI experts
Originally reported by nature.com
Read the original article →Original headline: Prospective evidence for conversational medical AI is hard, but non-negotiable - Nature Medicine