nature.com web signal

Google DeepMind team: benchmarks won't earn clinical AI trust

TL;DR

  • A Nature Medicine commentary from Google Research and DeepMind argues trust in clinical AI cannot be established by benchmark scores alone.
  • The authors say the hardest lessons in real-world deployment concern the humans and systems around the AI, not the model.
  • Their accompanying AMIE feasibility study with 100 patients at Beth Israel Deaconess found zero safety stops but PCPs beat the model on practicality.

Trust in clinical AI 'cannot be benchmarked into existence' and must be earned through prospective, real-world studies where 'the hardest lessons often concern the humans and systems around the AI, not the technology itself,' argue Mike Schaekermann, Anil Palepu and Po-Hsuan Cameron Chen in a Nature Medicine commentary published September 14. The corresponding authors write from Google Research and Google DeepMind, with named affiliations at Harvard Medical School, Beth Israel Deaconess Medical Center and Stanford University.

The piece lands alongside the group's own prospective feasibility study of AMIE at Beth Israel Deaconess. According to a Google Research write-up, 100 adult patients text-chatted with the model before their primary care appointments under live physician safety supervision, and no safety stops were triggered. AMIE placed the correct diagnosis inside its top 7 possibilities in 90% of cases and as the single most likely diagnosis in 56%. Blinded evaluator panels rated AMIE and primary care providers comparable on differential and plan quality overall, but the PCPs outperformed the model on practicality and cost-effectiveness of plans.

The framing is notable because the authors sit inside the same team publishing the AMIE results; here they are arguing that those numbers, on their own, do not establish clinical trust. Two of the researchers we follow shared the commentary the day it went up.

Shared on Bluesky by 2 AI experts