Alibaba: Top-16 Reverse KL Matches Full-Vocab OPD on 3 Pairs
TL;DR
- Alibaba Cloud researchers find restricting reverse KL to a student-selected top-16 subset of the shared vocabulary at strict positions matches full shared-vocabulary OPD on three teacher-student pairs.
- Adding mean squared error supervision on mismatched span groups lowered the full-average accuracy in every one of 18 tested positive-weight settings.
- Despite static vocabulary Jaccard overlap of 39.49-64.87%, strict student-token coverage ran 85.57-96.98% and the shared vocabulary held 99.69-99.90% of teacher probability mass at measured strict positions.
Restricting reverse KL to a student-selected top-16 subset of the shared vocabulary at strictly aligned positions matches full shared-vocabulary on-policy distillation across three heterogeneous teacher-student pairs, while adding mean squared error supervision on mismatched span groups lowers accuracy in every one of 18 positive-weight settings tested. That is the headline finding of a new preprint from Alibaba Cloud Computing, posted to Hugging Face papers, which argues the field has been chasing the wrong objective in cross-tokenizer distillation.
In the main results table, Strict top-16 posts a full average of 32.64 on Qwen→Llama, 42.26 on Granite→Phi, and 47.06 on Granite→Qwen, within roughly a quarter-point of the Strict full reference on each pair and ahead of the four evaluated cross-tokenizer baselines (ULD, Extended ULD, GOLD, SimCT) by 0.51 to 1.05 percentage points in the full average. The authors write that "with k=16, training preserves at least 96% of the full-average improvement achieved by full shared-vocabulary OPD over the undistilled student."
The case for narrowing supervision rests on a probability-mass measurement. Despite static vocabulary Jaccard overlap of 39.49-64.87% across the three pairs, strict student-token coverage runs 85.57-96.98%, and the shared vocabulary retains 99.69-99.90% of teacher mass and 98.99-99.81% of student mass at the measured strict positions. The vocabulary entries excluded from direct comparison, in other words, carry very little predictive weight on average.
Gradient diagnostics offer a mechanism for why extra span supervision hurts: at checkpoints from a strict-only run, the span gradients show weak or negative cosine agreement with the strict gradients and grow in relative magnitude across training. "Increasing coverage alone does not guarantee better distillation," the paper concludes, framing the shift as moving from alignment coverage to supervision reliability.
The study fits a run of fine-tuning work on this tracker in the past week, alongside the OPPD and TRACE papers, each narrowing where the training signal should actually land. Training ran for 100 on-policy iterations on a single node of 8 NVIDIA H20 GPUs with batch size 512, over 10,000 math prompts from DAPO-Math-17K and 10,000 code prompts from CodeForces.
Originally reported by huggingface.co
Read the original article →Original headline: Alibaba Paper Rethinks Cross-Tokenizer Distillation, Top-16 Reverse KL Beats Full-Vocab OPD