On-prem medical AI matches cloud accuracy in Nature Medicine study
TL;DR
- An on-premise agent running open-weight Qwen-3.5 hit 90.04% on the MIRA-v2 seven-disease benchmark, within 0.7 percentage points of cloud GPT-5.2 at 90.7%.
- Behavioral consistency across repeated runs beat token-level probability as a signal for correct diagnoses, with AUC 0.860 versus 0.747.
- Filtering to a consistency score of at least 0.90 kept 49.4% of cases at 98.9% accuracy, making 3 errors across 272 autonomously handled decisions.
A fully on-premise clinical agent running the open-weight Qwen-3.5 model reached 90.04% diagnostic accuracy on the MIRA-v2 seven-disease benchmark, within 0.7 percentage points of a cloud GPT-5.2 baseline at 90.7%, according to a study published in Nature Medicine on 15 September 2026 with Jakob Nikolas Kather as senior author.
The more consequential finding is not the accuracy tie but how the system decides when to trust its own output. The authors tested three families of confidence signal: token-level likelihood, hedging-based linguistic certainty, and behavioral stability across repeated runs. "Behavioral consistency provided the strongest indicator of diagnostic correctness," the paper reports, with a receiver-operating AUC of 0.860 versus 0.747 for the probability score.
At a consistency threshold of at least 0.90, the agent retained 49.4% of cases and got 98.9% of them right, with 3 errors across 272 autonomously handled decisions; the remaining 49 errors concentrated in the deferred stream routed for clinician review. Two of the AI researchers we track in our Who's Who directory shared the paper's link the same week it landed.
Under a stress test the authors call induced information scarcity, using ungrounded patient testimony, overall accuracy fell from 90.6% to 70.2%. The consistency metric remained discriminative at AUC 0.875, while the token-probability and linguistic-certainty scores moved only modestly despite the accuracy collapse. That gap is the authors' case for consistency as the more robust routing signal.
Limits are explicit. Both primary benchmarks derive from MIMIC-IV, a single institution. The system reasons over text only with no native image interpretation. Five-run consistency estimation raises token use roughly five-fold versus single-pass inference. An observed age-gradient in accuracy, lower in older patients, is flagged for bias auditing. All evaluations are retrospective; prospective validation is not yet done.
"Safety, therefore, depends on reliable decision-time signals indicating when to trust, verify or defer," the paper states.
Shared on Bluesky by 2 AI experts
Originally reported by nature.com
Read the original article →Original headline: On-premise medical AI agents for reliable clinical decision-making - Nature Medicine