Dresden on-premise medical AI nears cloud diagnostic accuracy
TL;DR
- An open-weight model running entirely on hospital hardware hit 90.0% diagnostic accuracy versus 90.7% for cloud baseline GPT-5.2 on a 551-case benchmark.
- The system reached 98.9% accuracy on the 49.4% of cases it flagged as high-consistency, deferring the rest to human clinicians.
- Behavioral consistency across repeated runs was the strongest single indicator of a correct diagnosis, with AUC 0.860 and 0.875 under stress testing.
An open-weight model running entirely inside hospital infrastructure matched a frontier cloud model on a 551-case diagnostic benchmark, according to a study published September 15, 2026 in Nature Medicine by a team at TU Dresden's Else Kröner Fresenius Center for Digital Health.
On the MIRA-v2 benchmark, drawn from MIMIC-IV data and covering seven diagnostic conditions, the local model Qwen-3.5 reached 90.0% accuracy against 90.7% for cloud baseline GPT-5.2. The paper's abstract reports 90.04% on the seven-disease task and 83.8% on a second benchmark of 2,400 acute abdominal cases across four conditions. External evaluation used 990 cases from VivaBench spanning ten specialty groups.
The central claim is not the raw accuracy but the reliability layer wrapped around it. The authors report that diagnostic behavioral consistency, meaning whether the model returns the same answer across repeated runs, is the strongest single indicator of when it is right, with AUC of 0.860 and 0.875 under stress testing. "At 0.90 consistency threshold, 49.4% of cases were retained at 98.9% diagnostic accuracy, supporting autonomous handling with human review for uncertain cases."
"Our goal is an AI agent with selective autonomy," corresponding author Jakob Nikolas Kather said in a press release. "These systems should support clinicians in decision-making, but never take over completely."
Two researchers we track shared the paper the day it appeared. First author Li Zhang framed the work as "an important step towards more reliable medical AI agents." Per-condition breakdowns are missing from the public abstract, and a physician-consensus figure of over 90% alignment across 181 randomly selected cases appears in the press release rather than the paper's summary.
Shared on Bluesky by 2 AI experts
Originally reported by nature.com
Read the original article →Original headline: On-premise medical AI agents for reliable clinical decision-making - Nature Medicine