nature.com web signal

Nature study: on-prem clinical AI knows which cases to defer

TL;DR

  • An on-premise medical AI agent reached 90.04% accuracy on a seven-disease task and 83.8% on a four-disease task.
  • Behavioral consistency across repeated runs was the strongest reliability signal, with an AUC of 0.860.
  • At a consistency threshold of 0.90, 49.4% of cases were handled autonomously at 98.9% diagnostic accuracy.

An on-premise medical AI agent, running open-weight models inside the hospital, hit 90.04% on a seven-disease benchmark and 83.8% on a four-disease task. The more interesting result was what it did with cases it wasn't sure about.

The Nature Medicine paper reports that "diagnostic behavioral consistency provided the strongest discrimination of correctness (area under the curve (AUC) = 0.860)," a plainer signal than the token-level probability measures the field usually leans on. At a consistency threshold of 0.90, the system retained 49.4% of cases for autonomous handling at 98.9% accuracy, sending the rest to a clinician.

On the model side, open-weight systems running locally held up against the frontier cloud baseline. Qwen-3.5 topped the on-premise field at 90.0%, ahead of GLM-5 (89.7%), GLM-4.5-Air (88.4%) and GPT-OSS (85.3%); the cloud reference GPT-5.2 came in at 90.7%. External validation on VivaBench, a 990-case set spanning ten specialties, dropped to 72.22%.

The authors' framing is careful. Behavioral consistency, they write, "helps determine when autonomous handling may be reasonable and when uncertainty should remain with the clinician." Two of the researchers we follow shared the paper the day it appeared.

Shared on Bluesky by 2 AI experts