paper web signal

DiffGate lifts Qwen3-1.7B code pass@8 5.7 points over GRPO

TL;DR

  • DiffGate applies teacher distillation only to failed trajectories, scaled by group difficulty and smoothly bounded.
  • On code, Qwen3-1.7B gains 5.7 pass@8 points and Qwen3-0.6B gains 1.6 points over matched GRPO.
  • On math, avg@8 stays within 0.5 points of GRPO while pass@8 improves by 1.1 and 3.9 points.

A new post-training objective called DiffGate lifts code pass@8 of a Qwen3-1.7B student by 5.7 points over matched GRPO, according to a preprint from Karn Tiwari and collaborators.

The setup is a split of labor. On-policy distillation gives dense token-local guidance but is "weakly aligned with rollout correctness," while GRPO's group-relative rewards "capture task success but provide coarse token-level credit and vanish on all-failure groups." DiffGate wires them together: the verifier decides which rollouts get teacher guidance, and the teacher is applied "only to failed trajectories, scaled by group difficulty, and smoothly bounded to prevent extreme teacher-student discrepancies from dominating optimization."

The gain is uneven across settings. For the smaller Qwen3-0.6B student the code pass@8 lift is 1.6 points, not 5.7. Math avg@8 stays within 0.5 points of GRPO, and math pass@8 improves 1.1 and 3.9 points for the two student sizes. The abstract names no teacher model and reports no compute cost.