DIET halves LingBot-Video expert bank, VBench score ticks up
TL;DR
- DIET prunes LingBot-Video 30B-A3B from 6,144 experts to 3,072 without fine-tuning, cutting the checkpoint from 57 GB to 30 GB.
- Official VBench Total rises from 0.7941 to 0.8115 at 50% retention; quality goes 0.8125 to 0.8325, semantic 0.7204 to 0.7278.
- The method caches expert outputs in one calibration pass and replays deletions, selecting what to keep by minimizing Overall Diversity Loss.
A training-free pruning pass on LingBot-Video 30B-A3B drops half of the model's 6,144 experts and nudges its VBench Total score up from 0.7941 to 0.8115, according to a paper on arXiv by Jiachang Zhang, Teng Hu and co-authors. The checkpoint shrinks from 57 GB to 30 GB, which the authors say "enables single-card deployment on 48 GB GPU."
The pruning method, DIET, runs a single calibration pass to cache expert outputs, then evaluates deletion candidates by replaying from those cached tensors "with zero additional model forward passes," per the paper. Experts are retained by minimizing an Overall Diversity Loss that "penalizes functional divergence via cosine distance" between deleted and retained ones.
Not every retention budget holds up. Pulling the expert count down to 1,229 (20% retention) drops the VBench Total to 0.7596, and 80% retention lands at 0.8061, slightly below the 50% result. The quality subscore moves from 0.8125 to 0.8325 and semantic from 0.7204 to 0.7278, modest gains on a single benchmark on a single model.
Shared on Bluesky by 2 AI experts
Originally reported by arxiv.org
Read the original article →Original headline: DIET: Deletion-response Expert Trimming for Video Diffusion Transformers