paper web signal

HuatuoGPT-3's RL-only recipe tops GPT-6 Astra on HealthBench

TL;DR

  • HuatuoGPT-3's 27B open-source variant reports 70.1 on HealthBench Total and 71.4 on HealthBench Professional, surpassing frontier closed models including GPT-6 Astra.
  • The authors' one-stage RL method, OnePO, hits 67.2 on HealthBench Total with just 20K training samples, beating SFT+RL by 2.7 points and pure RL by 7.4.
  • The paper names two pure-RL failure modes, Gradient Starvation and Teacher-Distribution Anchoring, and treats teacher outputs as transient guidance rather than a fixed target.

HuatuoGPT-3, an open-source medical model described in a new paper on arxiv, reports a 27-billion-parameter variant scoring 71.4 on HealthBench Professional and 70.1 on HealthBench Total. Those numbers, the authors say, put it ahead of frontier closed models including GPT-6 Astra.

The method is the part the authors want cited. Their one-stage reinforcement-learning recipe, OnePO, replaces the standard SFT+RL two-stage pipeline. "OnePO achieves 67.2 on HealthBench (Total) with only 20K training samples, outperforming SFT+RL and pure RL by 2.7 and 7.4 points, respectively," the abstract reports.

The theoretical hook is a pair of named failure modes in pure-RL adaptation. "We identify these two failure modes as Gradient Starvation and Teacher-Distribution Anchoring," the paper writes. OnePO's answer is to treat teacher outputs as "transient guidance for policy improvement," strengthening learning on informative low-probability teacher tokens through Adaptive Objective Evolution, then retiring the teacher once the current policy can surpass it.

The abstract names neither the teacher model that produced the guidance nor the base model behind the 27B variant.