nature.com web signal

AMIE team: benchmarks alone can't earn clinical AI's trust

TL;DR

  • Nature Medicine commentary from AMIE authors argues clinical AI trust must be earned through prospective real-world studies, not benchmark scores.
  • Companion feasibility trial had 100 adult patients chat with AMIE before primary care visits at Beth Israel Deaconess Medical Center.
  • AMIE placed the eventual diagnosis in its top seven possibilities for 90% of patients, with zero safety stops required.

"Trust in clinical artificial intelligence (AI) cannot be benchmarked into existence." That is the opening line of a Nature Medicine commentary published September 14, 2026, co-signed by researchers from Google Research, Google DeepMind, Harvard Medical School and Stanford, and funded by Alphabet Inc. Trust "must be earned through rigorous prospective studies in real-world clinical settings, where the hardest lessons often concern the humans and systems around the AI, not the technology itself," the authors write in Nature Medicine.

The signatories include Mike Schaekermann and Anil Palepu of Google Research, Shekoofeh Azizi of Google DeepMind, Adam Rodman of Beth Israel Deaconess Medical Center and Harvard Medical School, Ethan Goh of Google Research and Stanford, and Ami Parekh of Included Health. It is a lineup that pairs Google's model builders with the clinicians who would run the systems in front of patients.

The commentary arrives six months after the Google Research team behind AMIE published a prospective, single-center feasibility trial of the same system. Per a Google Research write-up on that study, "100 adult patients completed pre-visit interactions with AMIE" before ambulatory primary care appointments at Beth Israel Deaconess Medical Center; 98 kept their appointments, physicians supervised every session live over video-call with screen-sharing, and "zero safety stops were required." AMIE "achieved 90% accuracy in including the final diagnosis within its top 7 diagnostic possibilities."

Two of the researchers we follow had already surfaced the commentary by the time we opened it, an early signal that the argument is landing with the audience it is written for.

Shared on Bluesky by 2 AI experts