nature.com web signal

On-premise clinical AI agent hits 90% on MIMIC-IV benchmark

TL;DR

  • The on-premise agent reached 90.04% accuracy on a seven-disease diagnostic task and 83.8% on a four-disease task, approaching cloud-model baselines.
  • Diagnostic behavioral consistency separated correct from incorrect answers with AUC 0.860, holding at 0.875 under stress testing.
  • At a 0.90 consistency threshold, the agent covered 49.4% of cases at 98.9% accuracy, the paper's selective-autonomy triage rule.

An entirely on-premise clinical AI agent reached 90.04% accuracy on a seven-disease diagnostic task and 83.8% on a four-disease task, approaching cloud-model performance on benchmarks derived from MIMIC-IV. That is the headline result from a Nature Medicine paper published this month, with Jakob Nikolas Kather as corresponding author.

The point of the paper is not the accuracy figure but the reliability framework wrapped around it. The team quantified internal-likelihood, language-based, and behavioral-stability measures for both diagnosis and reasoning, aiming at what they call selective autonomy: letting the agent handle only the cases where its own confidence signals check out. Diagnostic behavioral consistency did most of the discriminating work, with an AUC of 0.860 for separating correct from incorrect answers, and 0.875 under stress testing.

The operational number is a triage rule. "At a consistency threshold of 0.90, 49.4% of cases were retained at 98.9% diagnostic accuracy," the paper reports. In plain terms, the agent could handle roughly half the cohort near-autonomously and refer the rest to a clinician.

Two researchers we track flagged the paper the day it appeared. The evaluation is confined to MIMIC-IV-derived benchmarks, with no external-cohort result reported.

Shared on Bluesky by 2 AI experts