paper web signal

FilmBench Tests AI Video Against Beijing Film Academy Standards

TL;DR

  • FilmBench draws 1,169 prompts reverse-engineered from award-winning films across 20 genres, with 1,056 requiring multi-shot generation rather than single clips.
  • The taxonomy spans 3 axes, 12 components and 38 sub-metrics, built with directors and faculty from the Beijing Film Academy and Hujing Digital.
  • Leading video models scored substantially below their prior web-benchmark results, with a notable decline moving from single-shot to multi-shot generation.

Every AI video demo you have seen this year was probably graded on prompts scraped from the web and scored by another neural net. A group of 30 authors led by Shengyi Wang, working with directors and faculty from the Beijing Film Academy and Hujing Digital Media & Entertainment Group, argue in a new arXiv paper that this is the wrong test, and they have built a replacement called FilmBench.

The benchmark reverse-engineers 1,169 prompts from award-winning films across 20 genres. Crucially, 1,056 of those prompts involve multiple shots rather than a single clip, which is closer to how film work is actually structured. The evaluation taxonomy has 3 axes, 12 components and 38 sub-metrics, and it is paired with what the authors describe as an in-house expert-grade automatic evaluation agent, plus open-source cinematic language operators called FilmOps. The paper reports model-level correlation reaching Spearman ρ = 0.95 for text-to-video and 0.96 for reference-to-video.

The finding that matters is quieter than the scanner headline. When leading video generation models are graded against these film-school standards, performance scores fall substantially below prior web-based benchmarks. Consistent gaps show up in dynamic aesthetics, and a notable performance decline appears when transitioning from single to multi-shot generation, particularly affecting weaker models. In plain terms, the models look better than they are because the previous tests were easier than the job.

The honest caveats are worth naming. The automatic evaluator, however well correlated, is still the authors' own system, and single-lab benchmarks usually shift once outside teams retest. The abstract does not name which specific commercial models were evaluated or publish per-model scores, so anyone quoting a leaderboard off this paper is filling in blanks the source has not filled. Take the specifics as reported, not settled.

The upside is direction rather than verdict. If FilmBench and the FilmOps operators are picked up by other labs, the next round of model releases will have to demonstrate multi-shot coherence and not just one gorgeous clip, and the studios evaluating these tools for real production will finally have a scoring rubric written by people who make movies for a living.