nature.com web signal

Dresden on-premise medical AI hits 90% diagnostic accuracy

TL;DR

  • Dresden team's on-premise clinical agent scored 90.04% on a seven-disease diagnostic task and 83.8% on a four-disease task.
  • The strongest reliability signal was diagnostic consistency across repeated runs of the same case, with AUC 0.860.
  • Gated at 0.90 consistency, the system kept 49.4% of cases at 98.9% accuracy and deferred the rest to clinicians.

The Dresden team reports 90.04% diagnostic accuracy on a seven-disease task and 83.8% on a four-disease task, from a clinical AI agent that runs entirely inside a hospital's own infrastructure. The paper is in Nature Medicine, led by Jakob N. Kather at the Else Kröner Fresenius Center for Digital Health at TU Dresden, with first author Li Zhang and collaborators at the National Center for Tumor Diseases in Heidelberg.

The evaluation setup is two AI agents in conversation. One plays the physician and can request follow-up questions, clinical findings, and lab results. The other plays the patient. The conditions tested were appendicitis, cholecystitis, pneumonia, pulmonary embolism, and urinary tract infections.

The more interesting finding is the reliability signal. The best predictor of a correct diagnosis was not any confidence score the model self-reports, but whether the model produced the same answer when the same case was fed to it repeatedly. In a press release from the university, Kather said the goal is "an AI agent with selective autonomy. These systems should support clinicians in decision-making, but never take over completely."

Applied as a gate: at a consistency threshold of 0.90, the system kept 49.4% of cases and hit 98.9% accuracy on that retained pool. Everything else would be routed to a clinician. Automated evaluation matched physician consensus in over 90% of 181 reviewed cases. A couple of the researchers on our radar were passing the paper around over the weekend.

One gap the authors flag but do not close: results were poorer on simulated older patients, and Medical Xpress reports that "it remains unclear why simulated cases involving older patients in particular showed poorer results." Testing so far is agent-versus-agent, not real clinicians on real workflows.

Shared on Bluesky by 2 AI experts