nature.com web signal

On-premise clinical AI agent hits 90.04% on MIMIC-IV task

TL;DR

  • A fully on-premise clinical AI agent reached 90.04% accuracy on a seven-disease MIMIC-IV task and 83.8% on a four-disease task, approaching a cloud baseline.
  • Behavioral consistency of the diagnosis was the strongest reliability signal, with AUC = 0.860 for identifying correct answers and 0.875 under stress testing.
  • At a consistency threshold of 0.90, 49.4% of cases were retained at 98.9% diagnostic accuracy; the remainder are deferred for clinician review.

A fully on-premise clinical AI agent reached 90.04% accuracy on a seven-disease diagnosis task drawn from MIMIC-IV, according to a Nature Medicine paper led by Jakob Nikolas Kather's group.

On a four-disease task the same system hit 83.8%. The authors report it 'approaching a cloud baseline on the primary benchmark' while running entirely inside the institution.

The more interesting result is the selective-autonomy layer wrapped around it. The system scores its own outputs on 'internal-likelihood, language-based and behavioral-stability measures,' and behavioral consistency of the diagnosis carried the strongest signal about whether the agent was actually right, with an AUC of 0.860 that held up at 0.875 under stress testing. 'At a consistency threshold of 0.90, 49.4% of cases were retained at 98.9% diagnostic accuracy,' the paper reports. Roughly half the caseload gets triaged for autonomous handling at near-perfect correctness; the rest is deferred for clinician review.

Two researchers we track shared it shortly after publication.

Both benchmarks derive from MIMIC-IV, and the abstract does not name the local model powering the on-premise stack or specify which system serves as the cloud baseline.

Shared on Bluesky by 2 AI experts