HF Paper: Domain-Normalized Multi-Teacher On-Policy Distillation Beats Single-Teacher Baselines
Summary
A new Hugging Face daily paper introduces DN-MOPF (Domain-Normalized Multi-Teacher On-Policy Feedback), assimilating domain-agnostic feedback into single-student distillation across code, math, and instruction-following tasks at 8K-16K context. The write-up has drawn 22 upvotes since posting.
Originally reported by huggingface.co
Read the original article →Original headline: HF Paper: Domain-Normalized Multi-Teacher On-Policy Distillation Beats Single-Teacher Baselines