nature.com web signal

Dresden on-prem clinical AI matches GPT-5.2 in Nature test

TL;DR

  • Best on-premise model reached about 90% accuracy on a seven-condition benchmark, versus 90.7% for a GPT-5.2 cloud baseline on the same test.
  • Behavioral consistency across repeated stochastic runs was the strongest signal of diagnostic correctness, with an AUC of 0.860.
  • At a 0.90 consistency threshold, the agent retained 49.4% of cases at 98.9% accuracy, with only three autonomous errors.

An on-premise diagnostic agent developed by Jakob N. Kather's group in Dresden reached roughly 90% accuracy on a seven-condition benchmark and 84% on a four-condition benchmark, close to the 90.7% Nature Medicine reports for a GPT-5.2 cloud baseline on the same seven-condition test, while running entirely inside the hospital.

The paper frames clinical translation as blocked by two things: "institutionally governed deployment and reliable decision-time uncertainty estimation." The team's answer to the second half is a consistency score, whether the model returns the same diagnosis across repeated stochastic runs of the same case. That signal "provided the strongest discrimination of correctness (area under the curve (AUC) = 0.860)," the paper reports. At a consistency threshold of 0.90, the agent retained 49.4% of cases at 98.9% diagnostic accuracy, with only three autonomous errors.

"Our goal is an AI agent with selective autonomy," Kather said in a Dresden press release. "These systems should support clinicians in decision-making, but never take over completely." The evaluation runs a Physician Agent against a Patient Agent in simulated dialogue rather than live encounters, and the study is anchored on MIMIC-IV-derived cases rather than a hospital's own EHR. Two of the clinicians we track shared the paper within days of publication.

Shared on Bluesky by 2 AI experts