Yandex MMD Post-Training Cuts Diffusion LM Perplexity 17-21%
TL;DR
- Yandex Research's MMD post-training cuts MDLM-MMD generative perplexity 17-21% versus IDLM at matched entropy across 8, 16 and 32 OpenWebText sampling steps.
- On 16B DMax-LLaDA2.0 models, the method boosts tokens-per-forward 10.3-16.5% on math while lifting HumanEval-Instruct pass@1 by 2.4 points and MBPP-Instruct by 3.8.
- Training the 16B DMax variants takes roughly 13-19 minutes on 8 NVIDIA H100 GPUs, with no auxiliary models or full sampling trajectories.
Yandex Research has published a recipe for squeezing more out of diffusion language models after training, by matching generated and reference distributions in a frozen model's feature space. The abstract on Hugging Face describes a post-training method that minimizes "Maximum Mean Discrepancy (MMD) between generated and reference distributions in the feature space of a frozen pretrained DLM," and runs "without full sampling trajectories or jointly trained auxiliary models."
Across 8, 16 and 32 sampling steps on OpenWebText, the authors report MDLM-MMD cuts generative perplexity 17-21% versus IDLM at matched entropy. The 16B DMax-LLaDA2.0 variants with hybrid masked-uniform diffusion also pick up a 10.3-16.5% boost in tokens-per-forward on math benchmarks, which the paper says comes "with similar or higher accuracy on math and code benchmarks." HumanEval-Instruct pass@1 moves up 2.4 points, MBPP-Instruct 3.8.
Training is cheap by frontier-paper standards. Roughly 13 minutes for DMax-Math and 19 minutes for DMax-Coder on 8 NVIDIA H100 GPUs, a few hundred steps apiece.
The abstract reports only improvement deltas, not the baseline accuracies MDLM and IDLM were starting from. The paper lands amid a steady run of diffusion-LM research passing through our feed this week, including a separate study on targeted-bias attacks against the same model class. Code is at yandex-research/dlm-mmd.
Originally reported by huggingface.co
Read the original article →Original headline: Paper 'Representation-Space MMD' Post-Trains Diffusion LMs, Cuts Generative Perplexity 17–21% at Matched Entropy