nature.com web signal

Nature Medicine: benchmarks alone can't earn clinical AI trust

TL;DR

  • A Nature Medicine commentary from the AMIE team argues clinical AI trust must be won through prospective studies, not benchmark scores.
  • The authors span Google Research, Google DeepMind, Harvard Medical School, Beth Israel Deaconess and Stanford.
  • The hardest lessons in clinical AI, they write, come from the humans and systems around the model, not the algorithms themselves.

A new Nature Medicine commentary from the team behind Google's AMIE conversational diagnostic system argues that leaderboard scores alone cannot certify medical AI for the clinic.

"Trust in clinical artificial intelligence cannot be benchmarked into existence," write Mike Schaekermann, Anil Palepu, Adam Rodman and co-authors in Nature Medicine. It has to be earned, the authors argue, through rigorous prospective studies in real-world clinical settings.

The byline stretches across Google Research, Google DeepMind, Harvard Medical School, Beth Israel Deaconess Medical Center and Stanford, the same coalition behind the AMIE feasibility work published earlier in the year. Their weight in this piece falls less on the model and more on the deployment surface: the hardest lessons in clinical AI, they say, tend to come from the humans and systems surrounding the model rather than the algorithms themselves.

The commentary has been moving through the medical-AI research circles on our radar; two researchers we follow surfaced it into our feed since publication.

Shared on Bluesky by 2 AI experts