Tsinghua SGF+ Yields 24-Hour Video From 5-Second Training
TL;DR
- SGF+ assigns separate parameters to context writing and denoising while keeping causal attention shared, resolving a gradient conflict Tsinghua's team measured in Self Gradient Forcing.
- Trained on only 5-second rollouts, SGF+ supports continuous autoregressive video generation up to 24 hours without long-video fine-tuning or auxiliary losses.
- Generator parameters double from 1.4B to 2.8B, lifting training time per step by about 8.2%; inference latency barely moves at 4.969 versus 4.962 seconds.
Researchers from Tsinghua University, Joy Future Academy / JD, and the Chinese University of Hong Kong report that in autoregressive video diffusion models the gradients that write context for future frames and the gradients that denoise the current frame systematically point in opposite directions, and that giving each role its own parameters yields 24-hour continuous video from a model trained only on 5-second clips. The paper, posted on Hugging Face papers on October 8, 2026, calls the method Self Gradient Forcing Plus (SGF+).
The gradient-conflict measurement is the load-bearing claim. Analyzing a Self Gradient Forcing baseline across 128 prompts and four denoising timesteps, the authors measure a mean angular separation of 104.2° in Attention and 106.3° in FFN between context-writing and denoising gradients, with every one of the 512 paired cosine similarities negative in both module families. "Conflicting gradient contributions partially cancel on shared parameters," the paper writes, motivating the parameter split.
SGF+ keeps SGF's two-pass training (a no-gradient autoregressive rollout followed by a differentiable reconstruction pass) and preserves causal attention between the context writer and the denoiser, but routes each role's updates to its own weight set. Training uses only 5-second rollouts under the original generation objective; the authors state the method needs "no auxiliary losses, additional video training data, or long-horizon fine-tuning."
The efficiency numbers temper the free-lunch read. Generator parameters double from 1.4B to 2.8B. Training time per step rises from 11.79 to 12.76 seconds, roughly 8.2%, and inference memory grows from 24.85 to 27.96 GB. Inference latency barely moves at 4.969 seconds for SGF+ versus 4.962 for SGF on an 81-frame workload, because each token touches only its role-specific parameters per forward pass.
Evaluated on VBench-Long at 60s and MovieGen-128 at 240s, generation horizons of 12× and 48× the training window, SGF+ beats Self Forcing and SGF on subject consistency, background consistency, temporal flickering, motion smoothness, aesthetic quality and imaging quality. The one exception is Dynamic Degree, where SF or SGF score higher; the authors argue the higher scores reflect "scene jumps, object deformation, subject disappearance, or additional people appearing" that inflate motion metrics without improving quality. The paper publishes no per-hour quality curve inside its 24-hour rollouts.
The demo arrives in a dense run of autoregressive-video work AI Weekly has been tracking, with 46 stories in the last 90 days including this week's GRACE latency result on Wan2.1-14B and CtrlCache's compute cuts for interactive video world models. The authors also note the gradient-direction split is "already observable at teacher-forcing (TF) initialization," pointing toward role-aware parameterization well upstream of the forcing stage.
Originally reported by huggingface.co
Read the original article →Original headline: SGF+ Decouples Context-Writing and Denoising Gradients, Enables 24-Hour Continuous Autoregressive Video From 5s Clips