Survey Paper Maps Inference-Efficiency Mechanisms Across 125 Video and Audiovisual LLMs
Summary
A new arXiv survey reviews 125 efficiency papers on video and audiovisual LLMs, splitting mechanisms into four pipeline stages: input construction, encoder computation, token reduction and LLM execution. Top findings include HoliTom retaining ~100% relative accuracy at 25% token retention, HieraVid cutting prefill FLOPs to ~25% of baseline while keeping ~98% accuracy, and OmniZip delivering 3.42× speed-up with 1.4× memory reduction at 35% token retention. The authors flag heterogeneous input protocols as the biggest obstacle to cross-paper comparison.
Originally reported by huggingface.co
Read the original article →Original headline: Survey Paper Maps Inference-Efficiency Mechanisms Across 125 Video and Audiovisual LLMs