arxiv.org web signal

PIHF lifts rare-disease Recall@1 32.7 points on GPT-5.4

TL;DR

  • Vanderbilt's PIHF method lifted rare-disease Recall@1 by 32.7 points for GPT-5.4 and 31.1 points for Qwen3.6-35B.
  • A clinical expert, not an RL loop, holds admit-and-rollback authority over each candidate revision to the natural-language policy.
  • The paper tested one proprietary and three open-weight executors spanning 3 to 49 billion active parameters.

A versioned natural-language policy, reviewed by a clinician with admit-and-rollback authority, lifted Recall@1 by 32.7 percentage points for GPT-5.4 and 31.1 points for Qwen3.6-35B on ultra-rare-disease benchmarks, a new Vanderbilt preprint on arXiv reports. The authors call the approach Policy Iteration with Human Feedback (PIHF).

Minh-Ha Nguyen of Vanderbilt's Department of Epidemiology and Cathy Shyr of Vanderbilt University Medical Center frame PIHF as a substitute for post-training reinforcement learning on in-context diagnostic work. A language-model critic and a clinical expert review complete-panel reasoning and tool-use trajectories, localize recurrent failures, and form candidate revisions. The expert "retains authority over admission and rollback," the paper states, while Recall@1 and Recall@5 validate outcomes after candidate execution.

The authors tested one proprietary executor and three open-weight executors spanning 3 to 49 billion active parameters. The 1.7-point gap between the GPT-5.4 and Qwen3.6-35B gains suggests the open-weight tier can land close to a frontier model when the policy artifact, not the weights, is carrying the task.

Two researchers on our Who's Who list shared the preprint on its first day. The abstract gives percentage-point deltas but no absolute Recall baselines, and no accounting of how much clinician time each revision loop required.

Shared on Bluesky by 2 AI experts