Meta and Princeton Paper: Mid-Training KL Distillation Boosts Reasoning But Slows Factual Recall
Summary
A team led by Jacqueline He and colleagues from Meta, Princeton, and Washington show that standard forward-KL distillation lifts both reasoning and factual recall in pre-training, but during mid-training it keeps improving reasoning while actively slowing factual-recall acquisition. They trace the split to teacher-confidence disparities across domains and propose Switch Distillation, which routes each token to KL or cross-entropy based on teacher uncertainty.
Originally reported by arxiv.org
Read the original article →Original headline: Meta and Princeton Paper: Mid-Training KL Distillation Boosts Reasoning But Slows Factual Recall