NeurIPS Paper: VideoLLMs Forget Frame Order Before the Output
TL;DR
- A NeurIPS 2026 paper finds that reversing a video's frame order often leaves a VideoLLM's final answer unchanged.
- The authors show temporal information is acquired at intermediate layers but decays before reaching the output.
- Their Temporal Activation Injection method reinforces the signal at inference time with no retraining and no non-temporal regression.
Reversing the frame order of a video should flip a model's answer to "what happened first?" In today's VideoLLMs it often does not, and that single observation anchors a NeurIPS 2026 paper, Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs, submitted to arXiv on October 1 by Youngwoo Shin, Yusung Ro, Minseo Kim and Junmo Kim.
The authors argue that "VideoLLMs acquire temporal information at intermediate layers but fail to maintain it to the output," and they introduce a per-layer measurement, the temporal divergence vector τₗ, to show the signal rising then fading as depth increases. Their fix, Temporal Activation Injection, is training-free: it "extracts τₗ at the peak of the profile for each input and reinjects it into subsequent layers following the measured decay."
The method is evaluated on three off-the-shelf VideoLLMs — Qwen2.5-VL-7B, Qwen3-VL-8B and InternVL2.5-8B — across four benchmarks including TempCompass, TVBench, AoTBench and the general-purpose MVBench. The paper reports "consistent improvements in temporal reasoning across three VideoLLMs and four benchmarks" with "negligible impact on non-temporal tasks," and the authors have posted code at github.com/Youngwoo-git/Before-It-Fades.
It lands in a dense week of video-model research on our tracker, alongside OneStreamer's streaming video model trained on a one-million-record dataset and KAIST's World Observer work on video world models, both posted the same day.
Originally reported by arxiv.org
Read the original article →Original headline: NeurIPS Paper Shows VideoLLMs Lose Temporal Signal Mid-Network, Proposes Inference-Time Activation Injection Fix