ICML 2026: RL-Only Medical LLM Beats SFT+RL; 27B Surpasses GPT-6 Astra on HealthBench

Found first: a primary source the press has not covered yet.

A team including Junying Chen and Benyou Wang has published a paper accepted at ICML 2026 introducing OnePO, a one-stage RL training method that eliminates supervised fine-tuning from medical LLM domain adaptation and outperforms the standard SFT+RL pipeline. Built on this approach, HuatuoGPT-3-27B scores 71.4 on HealthBench Professional, which the authors say surpasses frontier models including GPT-6 Astra.

What the source says

OnePO targets two failure modes the authors identify in applying RL directly to domain adaptation without SFT: Gradient Starvation, where easy samples dominate early training, and Teacher-Distribution Anchoring, where the reference model's prior constrains how far the policy can move. Two mechanisms address these: Adaptive Objective Evolution adjusts the training objective as learning progresses, and Teacher Retirement reduces the reference model's influence over time. With 20K training samples, OnePO scores 67.2 on HealthBench (Total), 2.7 points above SFT+RL and 7.4 points above pure RL. HuatuoGPT-3-27B, trained with this approach, scores 70.1 on HealthBench (Total) and 71.4 on HealthBench Professional. Code and models are available at FreedomIntelligence/HuatuoGPT-3 on GitHub.

Why it matters

The standard path to a specialist medical LLM has been two stages: supervised fine-tuning on labeled clinical examples, then RL. OnePO removes the first stage and still outperforms the two-stage approach, which lowers the barrier to domain adaptation and reduces dependence on labeled training data, a persistent bottleneck in medical AI. If the HealthBench Professional score holds under independent evaluation, open 27B models can now compete with closed frontier models on professional medical benchmarks.