nature.com web signal

AMIE team: benchmarks can't earn conversational medical AI trust

TL;DR

  • Google's AMIE team argues in Nature Medicine that leaderboard benchmarks cannot substitute for prospective real-world studies of conversational medical AI.
  • An accompanying feasibility trial ran at BIDMC's Healthcare Associates clinic in Boston from April to November 2025, enrolling 100 adults with 98 attending the appointment.
  • AMIE placed the final diagnosis in its top seven in 90% of cases and named it most likely in 56%, but PCPs won on practicality and cost.

Trust in clinical AI cannot be benchmarked into existence. That is the opening line of a Comment published in Nature Medicine by the team behind Google's AMIE, arguing that leaderboard scores are the wrong instrument for deciding whether a conversational medical AI belongs anywhere near real patients.

"Trust in clinical artificial intelligence (AI) cannot be benchmarked into existence," the authors write. "It must be earned through rigorous prospective studies in real-world clinical settings, where the hardest lessons often concern the humans and systems around the AI, not the technology itself." The byline spans Google Research, Google DeepMind, Harvard Medical School, Beth Israel Deaconess Medical Center, Stanford and Included Health.

The Comment accompanies a prospective feasibility trial run at BIDMC's Healthcare Associates clinic in Boston from April 2025 to November 2025. In a companion Google Research post, Mike Schaekermann and Alan Karthikesalingam report that 100 adult patients completed a pre-visit chat with AMIE and 98 kept their scheduled primary-care appointment. AMIE placed the final diagnosis inside its top seven possibilities in 90% of cases and named it as the single most likely diagnosis in 56%.

Blinded clinicians rated AMIE's differential and management-plan quality as comparable to the primary care providers on most axes. On one, they were not: "PCPs outperformed AMIE in the practicality and cost effectiveness of Mx plans," the post says. Human supervisors watching every session over video never triggered a safety stop.

The study was single-center, text-only, and had no controlled comparison arm. The authors' own framing is that prospective work like this is the price of admission, not the finish line.

Shared on Bluesky by 2 AI experts