LinkedIn's Δ-MOPD Distills Teacher Shifts, Cuts H100-Hours 35%
TL;DR
- LinkedIn and Harvard's Δ-MOPD transfers each teacher's post-training shift re-anchored at the student's initialization instead of the raw endpoint policy.
- Removing inherited base pull cuts the teacher-term norm ratio from 5.2:1 to 1.44:1 and target–student KL from 0.152 to 0.031.
- With three composed teachers, the shift target beats endpoint composition by 4.11 Math and 1.95 five-benchmark points.
Multi-teacher on-policy distillation usually transfers each teacher's endpoint policy. A new preprint from LinkedIn and Harvard argues that is the wrong object: the endpoint carries along whatever preferences the teacher inherited from its own base model, and that inherited pull drowns out the post-training shift the student actually wants to learn.
The authors call their alternative Δ-MOPD. For each teacher, they subtract the base checkpoint from the endpoint logits and re-anchor the resulting shift at the student's own initialization, a frozen DeepSeek-R1-Distill-Qwen-1.5B. "Removing it reduces the teacher-term norm ratio from 5.2:1 to 1.44:1 and target–student KL from 0.152 to 0.031," the abstract reports. The closer target "reaches the endpoint arm's best evaluated Math score with 35% fewer allocated H100-hours."
The composition numbers are where the paper leans hardest. With three composed teachers, Δ-MOPD "exceeds endpoint composition by 4.11 Math and 1.95 five-benchmark points." With two, it only matches. An unpaired cross-tokenizer extension that projects a fourth shift onto the 151,665-token effective vocabulary improves all three Math benchmarks and adds a further 3.32 Math and 1.85 five-benchmark points. Under phased routing the observed order gap shrinks from 10.50 to 6.42 points.
Not every setting gains. Under interleaved routing, where each update sees a single teacher at a time, the two targets perform comparably. The authors are explicit about the scope of the claim: shift targets are useful "when teacher signals are combined at a state," not when they arrive one by one. The scaling run also carries a cost. Science/IF, which gets no in-domain prompts in that run, lands 1.29 points below the endpoint composite.
The paper is one of 58 fine-tuning items we've logged in the last 90 days on our fine-tuning tracker, following Alibaba's finding last week that Top-16 reverse KL matches full-vocab OPD on three pairs.
Originally reported by huggingface.co
Read the original article →Original headline: LinkedIn's Δ-MOPD Distills From Teacher-Minus-Base Logit Shifts, Hits Math SoTA With 35% Fewer H100-Hours