arxiv.org web signal

Meta and Princeton Paper: Mid-Training KL Distillation Boosts Reasoning But Slows Factual Recall

Meta Research Fine-tuning ai-research

Summary

A team led by Jacqueline He and colleagues from Meta, Princeton, and Washington show that standard forward-KL distillation lifts both reasoning and factual recall in pre-training, but during mid-training it keeps improving reasoning while actively slowing factual-recall acquisition. They trace the split to teacher-confidence disparities across domains and propose Switch Distillation, which routes each token to KL or cross-entropy based on teacher uncertainty.