huggingface.co web signal

Yandex MMD Post-Training Cuts Diffusion LM Perplexity 17-21%

TL;DR

  • Yandex Research's MMD post-training cuts MDLM-MMD generative perplexity 17-21% versus IDLM at matched entropy across 8, 16 and 32 OpenWebText sampling steps.
  • On 16B DMax-LLaDA2.0 models, the method boosts tokens-per-forward 10.3-16.5% on math while lifting HumanEval-Instruct pass@1 by 2.4 points and MBPP-Instruct by 3.8.
  • Training the 16B DMax variants takes roughly 13-19 minutes on 8 NVIDIA H100 GPUs, with no auxiliary models or full sampling trajectories.

Yandex Research has published a recipe for squeezing more out of diffusion language models after training, by matching generated and reference distributions in a frozen model's feature space. The abstract on Hugging Face describes a post-training method that minimizes "Maximum Mean Discrepancy (MMD) between generated and reference distributions in the feature space of a frozen pretrained DLM," and runs "without full sampling trajectories or jointly trained auxiliary models."

Across 8, 16 and 32 sampling steps on OpenWebText, the authors report MDLM-MMD cuts generative perplexity 17-21% versus IDLM at matched entropy. The 16B DMax-LLaDA2.0 variants with hybrid masked-uniform diffusion also pick up a 10.3-16.5% boost in tokens-per-forward on math benchmarks, which the paper says comes "with similar or higher accuracy on math and code benchmarks." HumanEval-Instruct pass@1 moves up 2.4 points, MBPP-Instruct 3.8.

Training is cheap by frontier-paper standards. Roughly 13 minutes for DMax-Math and 19 minutes for DMax-Coder on 8 NVIDIA H100 GPUs, a few hundred steps apiece.

The abstract reports only improvement deltas, not the baseline accuracies MDLM and IDLM were starting from. The paper lands amid a steady run of diffusion-LM research passing through our feed this week, including a separate study on targeted-bias attacks against the same model class. Code is at yandex-research/dlm-mmd.