nature.com web signal

Kather team's on-prem clinical AI nears cloud in Nature paper

TL;DR

  • On MIRA-v2 (n=551), a local Qwen-3.5 agent hit 90.0% versus GPT-5.2's cloud 90.7%, per Nature Medicine.
  • Behavioral consistency across five stochastic runs beat token-probability as a reliability signal (AUC 0.860 vs 0.747).
  • At a consistency threshold of 0.90, the system retained 49.4% of cases at 98.9% diagnostic accuracy.

A locally deployed medical AI agent hit 90.04% accuracy on a seven-disease diagnostic task, according to a paper in Nature Medicine by a team led by Jakob Nikolas Kather. The cloud baseline for the same benchmark, GPT-5.2, hit 90.7%.

The agent runs open-weight models on-premise via vLLM on NVIDIA H200 and RTX PRO 6000 GPUs. On the MIRA-v2 benchmark (n=551), Qwen-3.5 reached 90.0% and GLM-4.5-Air 88.4%. On a four-disease abdominal benchmark (n=2,400), Qwen-3.5 reached 83.8%. The prior best open-weight result on that task, Gemma-3, was 70.5%.

The paper's central finding is not the accuracy number but the reliability signal. Cross-run semantic agreement across five stochastic runs (what the authors call behavioral consistency) separated correct from incorrect diagnoses at AUC 0.860, versus 0.747 for token-level probability. At a consistency threshold of 0.90, the paper reports that "49.4% of cases were retained at 98.9% diagnostic accuracy," with the rest deferred for clinician review.

A blinded physician adjudication of 181 MIRA-v2 cases (33% of the benchmark) found 81.8% of reviewed cases clinically valid for both the agent's diagnosis and the reference label. Agreement between the LLM-based evaluator and the physician consensus was 92.3%, with Gwet's AC1 at 0.898.

The authors flag single-institution MIMIC-IV as the ecology for the primary benchmarks, text-only reasoning without native image interpretation, and per-deployment threshold calibration as constraints. "Prospective studies and dedicated bias audits will be required to determine how these signals affect clinician reliance, review burden, safety outcomes and fairness across patient subgroups," they write. Two researchers we track in our Who's Who directory circulated the paper.

Shared on Bluesky by 2 AI experts