nature.com web signal

On-premise Qwen-3.5 matches GPT-5.2 on diagnostic benchmark

TL;DR

  • Qwen-3.5 running on-premise hit 90.0% on the seven-disease MIRA-v2 benchmark, essentially matching the 90.7% GPT-5.2 cloud baseline.
  • Behavioral consistency across repeated runs was the strongest confidence signal (AUC 0.860); at threshold 0.90 the agent retained 49.4% of cases at 98.9% accuracy.
  • Physicians judged 81.8% of reviewed cases clinically valid; another 4.4% were valid despite the agent disagreeing with the EHR label.

The bottleneck for clinical LLMs has been legal, not technical: hospitals in most jurisdictions cannot send patient data to a hosted API. A Nature Medicine paper from a team with Jakob Nikolas Kather as senior author reports that a fully on-premise agent, running open-weight models, matched a cloud baseline on a benchmark of common acute diagnoses, and pairs that result with a runtime signal for when the agent should defer to a clinician.

On the seven-disease MIRA-v2 benchmark (551 cases covering appendicitis, cholecystitis, diverticulitis, pancreatitis, pneumonia, pulmonary embolism and UTI), the on-premise Qwen-3.5 configuration reached 90.0% accuracy against a GPT-5.2 cloud baseline at 90.7%. On the four-disease CDM benchmark of 2,400 abdominal cases, Qwen-3.5 hit 83.8%. On the external VivaBench, spanning ten specialty groups, it dropped to 72.22%.

The confidence framework is the part the authors lean on. Three signals were compared as gates on autonomy: token probability, linguistic hedging, and cross-run behavioural stability. "Behavioral consistency...helps determine when autonomous handling may be reasonable and when uncertainty should remain with the clinician," the paper reports. Consistency posted an AUC of 0.860, ahead of the probability score (0.747) and linguistic certainty (0.792), and held up (AUC 0.875) under stress tests with ungrounded patient inputs. Setting the consistency threshold at 0.90 kept 49.4% of cases inside the automated envelope, at 98.9% accuracy on those cases.

Physicians reviewed a subset: 81.8% of cases were judged clinically valid for both the agent and the EHR label, and a further 4.4% were valid for the agent even when it disagreed with the record. The paper frames the trade-off directly: "Clinical translation, however, remains limited by two unmet requirements: institutionally governed deployment and reliable decision-time uncertainty estimation." The framework does not describe how the deferred half of cases would route through a real ward, and every accuracy figure here is retrospective rather than prospective. Two of the AI researchers we track shared the paper this week.

Shared on Bluesky by 2 AI experts