AMIE team pushes prospective trials for medical chatbots
TL;DR
- A Nature Medicine comment from the AMIE team argues trust in clinical AI cannot be benchmarked into existence and must come from real-world trials.
- The pre-registered AMIE feasibility study at Beth Israel Deaconess enrolled 100 patients from April to November 2025 with human safety supervisors on every session.
- AMIE's differential included the final chart-review diagnosis in 90% of cases, with 75% top-3 and 56% top-1 accuracy.
A team from Google Research, Google DeepMind, Harvard Medical School and Stanford, writing in Nature Medicine, argues that "Trust in clinical artificial intelligence (AI) cannot be benchmarked into existence." It has to come, they write, from prospective studies in real clinics.
The comment leans on the authors' own AMIE feasibility trial, a pre-registered study run at Beth Israel Deaconess Medical Center's Healthcare Associates primary-care clinic. One hundred adult patients were enrolled from April to November 2025, with 98 completing the study, interacting with the conversational diagnostic AI by text chat up to five days before their appointments. Human safety supervisors monitored every session in real time. "Across all AMIE-patient interactions in this study, zero safety stops were required by the human AI supervisors," the authors report. Two researchers we track posted the paper within days.
On accuracy, AMIE's differential diagnosis included the final chart-review diagnosis at 8 weeks in 90% of cases, with 75% top-3 accuracy and 56% as the single most likely diagnosis. "AMIE and PCPs were rated to be on par with respect to their overall management plan and differential diagnosis quality," the authors write, though the human primary-care providers were rated higher on cost-effectiveness and practicality.
The framing in the abstract carries the argument: real-world trials, not leaderboard scores, are what earns clinical trust, and "the hardest lessons often concern the humans and systems around the AI, not the technology itself."
Shared on Bluesky by 2 AI experts
Originally reported by nature.com
Read the original article →Original headline: Prospective evidence for conversational medical AI is hard, but non-negotiable - Nature Medicine