SpecFold claims up to 1.99x speedup on diffusion LLM decoding
TL;DR
- SpecFold reports up to 1.99x throughput over vanilla decoding and up to 1.64x over the Spiffy baseline on diffusion LLMs.
- The method reuses parent attention and FFN computation across draft branches during speculative verification via token-level residual gating.
- Evaluated across two DLLM families, five models and five benchmarks, with the authors describing it as orthogonal to temporal caching.
SpecFold, a new speculative-decoding method for diffusion large language models, reports up to 1.99x throughput over vanilla decoding and up to 1.64x over the prior Spiffy baseline, in a preprint posted this month on arXiv.
The authors, Chung-En Ho, Weiyu Sun, Cheng-Jhih Shih, He Li, Yong Liu and Yingyan Celine Lin, target what they call 'multi-branch computational redundancy' in diffusion LLMs that 'generate text through iterative block denoising' and are accelerated by verifying a main branch together with multiple draft branches in a single forward pass. SpecFold 'performs token-level residual gating and selectively reuses parent computation through folded attention and FFN while preserving residual hidden states,' the paper writes, backed by 'a Triton kernel implementation' for sparse multi-branch execution.
The method was evaluated 'across two DLLM families, five models, and five standard benchmarks,' and the authors position it as 'orthogonal to temporal caching and compatible with existing DLLM speculation strategies.' The abstract names no specific models, publishes no absolute tokens-per-second figures, and does not quantify what 'comparable task performance' means in accuracy terms.
Originally reported by paper
Read the original article →Original headline: SpecFold Delivers Up to 2× Throughput on Diffusion LLMs by Folding Multi-Branch Redundancy