nature.com web signal

TU Dresden's on-premise medical AI hits 90% on MIMIC-IV

TL;DR

  • A fully on-premise clinical AI agent from TU Dresden scored 90.04% on a seven-disease MIMIC-IV task and 83.8% on a four-disease task.
  • Diagnostic behavioral consistency across repeated runs (AUC 0.860) outperformed the model's own internal probability as a correctness signal.
  • At a 0.90 consistency threshold the agent retained 49.4% of cases at 98.9% accuracy; the rest are deferred to clinicians.

An on-premise diagnostic AI agent developed at TU Dresden reached 90.04% accuracy on a seven-disease MIMIC-IV task and 83.8% on a four-disease task, according to a paper published in Nature Medicine on September 15, 2026.

The system, led by Prof. Jakob N. Kather with first author Li Zhang, was evaluated in a simulation environment where two AI agents interact, one taking the role of the physician and the other the role of the patient. Conditions tested included appendicitis, cholecystitis, pneumonia, pulmonary embolism, and urinary tract infections. Automated scoring agreed with physician consensus on more than 90% of 181 randomly selected cases.

"Our goal is an AI agent with selective autonomy," Kather said. "These systems should support clinicians in decision-making, but never take over completely."

The interesting part is where the reliability signal came from. The strongest predictor of a correct diagnosis was not the model's internal probability score. It was consistency, whether the agent produced the same diagnosis across repeated runs of the same case. The paper reports that this diagnostic behavioral consistency provided the strongest discrimination of correctness, with an AUC of 0.860. At a 0.90 consistency threshold, 49.4% of cases were retained at 98.9% diagnostic accuracy; the rest would be routed to a clinician for review.

Kather framed the on-premise choice as a European posture, in a release from EurekAlert: "We need to move faster, not slower. What matters is that we develop these systems in a way that is safe, transparent and preserves data sovereignty." The work was done jointly with the National Center for Tumor Diseases in Heidelberg.

The evaluation is retrospective and simulation-based rather than tested during routine clinical care. Two researchers on our radar shared the paper the day it landed.

Shared on Bluesky by 2 AI experts