On-premise clinical AI matches cloud GPT-5.2 in Nature study
TL;DR
- An on-premise clinical agent hit 90.04% accuracy on a seven-disease task, within 0.7 points of the GPT-5.2 cloud baseline at 90.7%.
- Behavioral consistency across five stochastic runs was the strongest reliability signal, with AUC 0.860 for flagging correct vs incorrect answers.
- At a 0.90 consistency threshold, the agent auto-handled 49.4% of cases at 98.9% accuracy and deferred the rest for physician review.
An on-premise clinical AI agent reached 90.04% accuracy on a seven-disease diagnostic task and 83.8% on a four-disease task, according to a paper this month in Nature Medicine. The best local model tested, Qwen-3.5, landed within 0.7 percentage points of the GPT-5.2 cloud baseline at 90.7%.
The work is by Li Zhang, Georg Wölflein, Dyke Ferber and colleagues, benchmarked on MIMIC-IV-derived sets MIRA-v2 (n=551) and CDM (n=2,400). Alongside raw accuracy the team built what they call a multi-perspective reliability framework: three families of decision-time signals meant to flag when the model's answer should be trusted.
Behavioral consistency across five stochastic runs was the strongest of those signals, with an AUC of 0.860 for discriminating correct from incorrect answers. The authors write that their reliability signals "identify a lower-risk subset for autonomous handling and defer the remainder for review." Set the consistency threshold at 0.90 and the agent keeps 49.4% of cases at 98.9% diagnostic accuracy; the rest go to a clinician.
The stress test is where the ceiling shows up. Under input perturbation, accuracy fell from 90.6% to 70.2%, a 20.3-point drop, though the consistency signal itself held its ground at AUC 0.875. And even inside the high-confidence bucket, the paper notes that "high-confidence errors reflected premature closure on a coherent" diagnosis, meaning the model was wrong because it locked in early.
Two clinical AI researchers we track shared the paper on release day.
Shared on Bluesky by 2 AI experts
Originally reported by nature.com
Read the original article →Original headline: On-premise medical AI agents for reliable clinical decision-making - Nature Medicine