huggingface.co web signal

SGF+ Decouples Context-Writing and Denoising Gradients, Enables 24-Hour Continuous Autoregressive Video From 5s Clips

Video Generation ai-research

Summary

Tsinghua/CUHK's SGF+ finds that in autoregressive video models the gradients from denoising current frames and from writing KV context for future frames are systematically negatively aligned, and splits parameters for the two roles while keeping causal attention shared. Trained on only 5-second rollouts without any auxiliary losses or long-video fine-tuning, the model extrapolates continuous generation up to 24 hours while improving framewise and chunkwise quality over the evaluated baselines.