GRACE Cuts Wan2.1-14B Video Latency 11.1× at Matched VBench
TL;DR
- GRACE reduces Wan2.1-I2V-14B tokens nearly 8× and latency 11.1× at 480×832 with 81 frames, matching the pretrained pipeline's VBench-I2V total of 87.90 vs 87.92.
- The retrofit costs 38.5 H200 GPU-days total: 8.5 for the autoencoder stage and 30 for DiT adaptation, with inference running on a single A100.
- At 736×1280 with 81 frames the pipeline runs the same backbone in 218.8 seconds against 3396.8 for pretrained Wan2.1-14B, a 15.5× speedup.
GRACE, a video-diffusion compression method posted to Hugging Face Papers, reduces the token count of Wan2.1-I2V-14B by nearly 8× and cuts its latency by 11.1× at 480×832 with 81 frames while matching the pretrained pipeline's VBench score. The paper notes the work "was done while the first three authors were interns at Kakao Corp."
The method keeps the frozen pretrained encoder's base latent and learns a residual latent for information lost under stronger compression. The autoencoder, the paper writes, is aligned "with the pretrained latent in the feature space of the frozen DiT, so that the autoencoder is optimized for generation rather than reconstruction alone." A second stage adapts the DiT with lightweight fine-tuning and asymmetric denoising, where "the base latent is denoised ahead of the residual latent" in a shared forward pass.
The headline VBench parity is cleaner than the reconstruction numbers. On VBench-I2V, GRACE totals 87.90 against the pretrained Wan2.1-14B's 87.92. On Panda-70M at 256×256×81, PSNR falls from 35.15 to 32.63 and rFVD jumps from 1.13 to 13.53. The retrofit cost sits at 38.5 H200 GPU-days (8.5 for the autoencoder, 30 for the DiT), and at 736×1280×81 the authors report 218.8 seconds against 3396.8 for the pretrained baseline on a single A100, a 15.5× speedup. It lands in a dense run of inference-efficiency work we have been tracking this quarter.
Originally reported by huggingface.co
Read the original article →Original headline: GRACE Latent-Compressed Video Diffusion Runs 11.1x Faster at 480x832 With No VBench Quality Drop