huggingface.co web signal

OPPD Trains Qwen2.5-Math-7B to Beat 64-Candidate Power Sampling

Fine-tuning Open Source ai-business

TL;DR

  • On-Policy Power Distillation lifts single-generation MATH500 accuracy by up to 23.0 points and GSM8K by 27.3 points over the untrained model.
  • One trained generation scores 2.4 and 3.5 points above published power sampling with 64 candidates, recovering 94% of the 16-candidate gain.
  • Against GRPO at the same budget and no reference answers, OPPD scores 3.8, 4.0 and 5.4 points higher on MATH500, GSM8K and AIME.

A language model can give the correct answer more probability than any single wrong answer and still sample a wrong one, because the wrong answers together hold more probability. That is the opening observation of a new paper on Hugging Face from a USC and Intel AI group, and it is the problem their method is built to erase at training time rather than at decode time.

The method, On-Policy Power Distillation (OPPD), trains a student to produce the answers that power sampling would have chosen, in one shot. On Qwen2.5-Math-7B the untrained baseline scores 66.8% on MATH500 at temperature 1; the paper reports OPPD "raises single-generation accuracy by up to 23.0 points on MATH500 and 27.3 on GSM8K over the untrained model at the same temperature," and that "one generation scores 2.4 and 3.5 points above published power sampling with 64 candidates, recovering 94% of the gain that 16 candidates give the untrained model."

Against GRPO, run at the same checkpoint and the same compute budget and without reference answers, OPPD scores 3.8, 4.0 and 5.4 points higher on MATH500, GSM8K and AIME. The two stack: "oppd applied after GRPO adds up to 9.3 points." Trained only on mathematics, OPPD still lifts HumanEval by up to 5.3 points, and on deepseek-math-7b-rl, a checkpoint already trained with verified rewards where "lowering the temperature gives nothing," it adds 4.4 points on MATH500.

The mechanism is a sequential Monte Carlo sampler: the student generates candidates and a frozen teacher's power distribution weights them, with those same weights driving a maximum-likelihood update. One loss coefficient moves the "sharpening exponent the model absorbs between 1.19 and 2.02, against 1.14 for ordinary on-policy distillation." The training recipe is small — 300 steps at 2 prompts per step, 16 candidates per generation — and the code is on GitHub.

The numbers are from the authors' own evaluation on 200 to 250 problems per benchmark with 4 samples each, no third-party replication yet. The arrival fits a run of training-time efficiency work our tracker has logged in the last day, alongside an activation-alignment paper and a prefill-cues result on Olmo-3-7B and Qwen3-14B matching RL math gains without the RL.