nature.com web signal

TU Dresden's on-premise medical AI nears cloud accuracy

TL;DR

  • TU Dresden's on-premise diagnostic AI hit around 90% accuracy versus 90.7% for a cloud baseline running the same agent architecture.
  • The strongest reliability signal was consistency: diagnoses that stayed stable across repeated runs were most likely to be correct.
  • Physicians reviewed 181 randomly selected cases and their consensus agreed with the automated evaluation more than 90% of the time.

An on-premise medical diagnostic system developed at TU Dresden reached roughly 90% accuracy on one clinical benchmark and 84% on a second, closing most of the gap to a cloud-hosted baseline that scored 90.7% on the same task, according to a study published in Nature Medicine. The best-performing local model was Qwen-3.5, running entirely inside the hospital's own infrastructure.

The setup uses two AI agents in a simulation environment, one playing the physician and the other the patient, tested across five conditions: appendicitis, cholecystitis, pneumonia, pulmonary embolism and urinary tract infections. Physicians then hand-reviewed 181 randomly selected cases, and their consensus aligned with the automated evaluation more than 90% of the time, EurekAlert reported.

The more novel piece is a measure that flags which diagnoses the system is likely to have gotten right. The strongest signal turned out to be behavioral. "The more stable the diagnosis remained across multiple runs, the more likely it was to be correct," the researchers found. That lets the system defer uncertain cases to a clinician rather than running fully autonomously.

Senior author Jakob N. Kather, Professor of Clinical Artificial Intelligence at TU Dresden, framed the design as deliberately partial. "Our goal is an AI agent with selective autonomy," he said, one that would "never take over completely." He also pushed back on the pace of clinical AI deployment: "We need to move faster, not slower," because "patients have so far seen very little benefit." Simulations showed weaker performance on older patients, a limitation the authors flag for follow-up.

Shared on Bluesky by 2 AI experts