Michigan Paper: T5Gemma-2 Distill Beats GPT-2-M on Gen PPL
TL;DR
- Swapping T5-small for T5Gemma-2-270M embeddings cuts generative perplexity about 40% at matched entropy, but makes the latent space too sharp to diffuse through.
- Distilling T5Gemma-2 into a student encoder on the teacher's decoded probabilities yields 17.8 Gen. PPL versus a 15.4 real-text PPL on OpenWebText, outperforming GPT-2-M.
- The distillation trades discrimination for diffusibility: SST-2 probing on the student drops from 89.4 to 78.1.
Scaling the text encoder behind a continuous diffusion language model from T5-small to T5Gemma-2-270M cuts generative perplexity "by about 40% at the same entropy as T5-small," according to a new University of Michigan paper on Hugging Face. The stronger embeddings create their own problem: they are so sharply separated that even plausible alternative words sit in distinct neighborhoods, and continuous diffusion often misses all of them and ends up at an invalid embedding instead.
Zekai Zhang and co-authors distill T5Gemma-2 into a student encoder trained on the teacher's decoded probabilities as soft labels, which "pull the alternative embeddings closer while maintaining the encoding-decoding mechanism." Their medium-sized diffusion LM reaches 17.8 generative perplexity against a real-text PPL of 15.4 on OpenWebText, outperforming GPT-2-M on Gen. PPL.
The trade-off is explicit. SST-2 classification probing on the student drops from 89.4 to 78.1, because "the distillation trades discrimination for diffusibility." Code and project page are public.
Originally reported by huggingface.co
Read the original article →Original headline: HF Paper: Distilling T5Gemma-2 Embeddings Cuts Diffusion-LM Generation Perplexity ~40%