On-premise clinical AI agent hits 90% in Nature Medicine
TL;DR
- An on-premise agent using Qwen-3.5 hit 90.04% on a seven-disease benchmark, essentially matching a cloud GPT-5.2 baseline of 90.7%.
- Behavioral consistency across five stochastic runs was the strongest predictor of correctness (AUC 0.860), beating token-probability scores at 0.747.
- At a 0.90 consistency threshold, the system autonomously handled 49.4% of cases at 98.9% accuracy and deferred the rest to clinicians.
An on-premise clinical AI agent built on open-weight models hit 90.04% diagnostic accuracy on a seven-disease benchmark, essentially matching a cloud GPT-5.2 baseline of 90.7%, according to a study published in Nature Medicine.
The system, from a group led by Jakob Nikolas Kather, pairs a multi-turn diagnostic dialogue between a physician agent and a patient agent with a confidence framework designed to flag when the model should defer to a human. On the MIRA-v2 benchmark of 551 MIMIC-IV cases spanning appendicitis, cholecystitis, pneumonia, pulmonary embolism and three other acute conditions, Qwen-3.5 scored 90.0%, GLM-5 scored 89.7%, and GPT-OSS scored 85.3%. On a separate four-condition CDM benchmark of 2,400 cases, Qwen-3.5 reached 83.8%, above the previous open-weight best of 70.5% from Gemma-3.
The reliability layer is the paper's more distinctive contribution. The authors compared five candidate confidence signals — token-level probabilities, hedging language, clinical concept density, and behavioral consistency across five stochastic runs — and found that behavioral consistency on the final diagnosis was the strongest correctness discriminator, with AUC = 0.860, beating the token-probability score at 0.747. "Behavioral consistency provided the strongest indicator of diagnostic correctness," the authors write.
That signal drives what the paper calls selective autonomy. At a consistency threshold of 0.90, the system kept 49.4% of cases — 272 of 551 — and got 98.9% of them right, with just 3 autonomous errors; the remaining cases were flagged for clinician review. Under a stress test in which the patient agent was stripped of chief complaint and clinical history and forced to confabulate symptoms, accuracy dropped from 90.6% to 70.2%, but the consistency metric continued to discriminate correct from incorrect answers (AUC = 0.875).
The design is not free. Five-run consistency analysis increased token usage roughly 5-fold relative to single-pass inference, and threshold calibration varied with the choice of sentence encoder used to measure semantic similarity across runs. The authors position the setup as governance-first: a framework in which "institutionally governed deploy is paired with decision-time reliability estimation."
Shared on Bluesky by 2 AI experts
Originally reported by nature.com
Read the original article →Original headline: On-premise medical AI agents for reliable clinical decision-making - Nature Medicine