paper web signal

TMLR Survey Sorts Video-Gen Post-Training Into Four Buckets

TL;DR

  • A TMLR survey groups post-training work for video models into four families: supervised fine-tuning, self-training and distillation, preference and reward-based, and inference-time methods.
  • The authors split alignment into implicit vs explicit modes, based on how alignment signals are enforced during training or deployment.
  • They flag four open problems: scalable reward design, long-horizon temporal consistency, stability-expressiveness trade-offs, and safety-aware generation.

The first comprehensive review of post-training methods for video-generation models sorts the field into four method families and names four problems the authors say remain unsolved. The TMLR survey, by Chaoyu Li and twelve co-authors, frames post-training as the way labs can "adapt pretrained models without retraining them from scratch."

The paper organizes existing work into "supervised fine-tuning methods, self-training and distillation methods, preference- and reward-based methods, and inference-time methods," and introduces a split between implicit and explicit alignment based on how alignment signals are enforced.

Video alignment, the authors argue, is harder than text or image alignment because of "error accumulation over time, motion-appearance coupling, multi-objective trade-offs, and limited supervision for temporal properties." The open problems they flag — "scalable reward design, long-horizon temporal consistency, stability-expressiveness trade-offs, and safety-aware generation" — are the same issues most video models surface when they are pushed from short demo clips into longer, controlled output.

The abstract publishes no head-to-head numbers across methods or benchmarks; the contribution is the map, not a ranking.