UIUC HC-DLM Couples Discrete Tokens With a Continuous Latent
TL;DR
- UIUC's HC-DLM makes a continuous latent the only persistent generative state, with tokens read out at every denoising step and fed back as a scaffold.
- The training objective is derived from a variational bound on the token likelihood, coupling discrete and continuous generation in a single process.
- At matched model size, the authors report gains over discrete and continuous diffusion baselines on Sudoku and Countdown puzzle accuracy and LM1B perplexity.
A new paper from UIUC, posted to Hugging Face, goes after a specific failure mode in diffusion language models: when you decode tokens in parallel, each one is sampled independently from its own marginal, which severs the statistical dependencies between the tokens being produced in the same step.
The authors' fix, which they call Hierarchical Continuous Diffusion Language Models (HC-DLM), promotes a continuous latent to the only persistent generative state. "Tokens are read out from it at every step and feed back as a scaffold for the next latent update," the abstract says, with the whole thing trained against "a variational bound on the token likelihood." That is a deliberate contrast with recent methods that bolt continuous context onto an otherwise self-contained discrete chain.
On results, the paper is careful: HC-DLM "improves over discrete and continuous diffusion baselines at matched model size," measured as puzzle accuracy on Sudoku and Countdown and generative perplexity on LM1B. The abstract does not publish the headline percentages, the parameter counts, or the baselines it beats; those live in the full PDF.
It lands amid a visible run of diffusion- and latent-coupling papers in our open-source feed, including a separate Latent-Foresight write-up the same day.
Originally reported by huggingface.co
Read the original article →Original headline: HF Paper HC-DLM Couples Discrete Tokens With a Continuous Latent to Fix Parallel Decoding in Diffusion LMs