AMIE team calls prospective medical AI trials non-negotiable
TL;DR
- AMIE's differential diagnosis included the final diagnosis in 90% of cases at 8-week chart review, with 75% top-3 accuracy.
- Across 100 pre-visit chats at Beth Israel Deaconess, physician supervisors did not have to trigger a single safety stop.
- The commentary argues trust in clinical AI cannot be benchmarked into existence and must come from prospective studies in real-world settings.
One hundred adult patients spent up to five days chatting with a conversational diagnostic AI before their primary care appointment. Across every session, the physicians watching by live video call with screen sharing did not need to trigger a single safety stop.
That trial, run at Beth Israel Deaconess Medical Center's Healthcare Associates clinic in Boston from April 2025 to November 2025, is the empirical anchor for a companion commentary in Nature Medicine arguing that trials like it are the only way medical AI earns clinical trust. 'Trust in clinical artificial intelligence (AI) cannot be benchmarked into existence,' write Mike Schaekermann, Anil Palepu, Adam Rodman and colleagues from Google Research, Google DeepMind, Harvard Medical School and Stanford. 'It must be earned through rigorous prospective studies in real-world clinical settings, where the hardest lessons often concern the humans and systems around the AI, not the technology itself.'
AMIE, Google's conversational diagnostic system, produced a differential that included the final diagnosis in 90% of cases at 8-week chart review, with 75% top-3 accuracy. Blinded reviewers rated its differential and management plans as similar overall in quality to those of the primary care physicians who saw the same patients, per the Google Research write-up. Patient attitudes shifted more positive after the AMIE chat and stayed elevated after the in-person visit. Of the 100 pre-visit interactions, 98 patients attended the appointment.
Two of the researchers on our tracker posted the commentary the day it dropped.
A hundred patients at one Boston clinic is a feasibility signal, and the authors are explicit that the hardest work sits with the humans and systems around the model, not the model itself.
Shared on Bluesky by 2 AI experts
Originally reported by nature.com
Read the original article →Original headline: Prospective evidence for conversational medical AI is hard, but non-negotiable - Nature Medicine