huggingface.co web signal

Alibaba's FilmBench grades AI video on 1,169 film prompts

TL;DR

  • Alibaba, Hujing's Moku Lab and Beijing Film Academy released FilmBench, a 1,169-prompt benchmark reverse-engineered from award-winning films across 20 genres.
  • The taxonomy has three L1 axes (Instruction Following, Temporal Continuity, Aesthetic Quality), 12 L2 components and 35 L3 sub-metrics, with 1,056 multi-shot prompts.
  • Auto-evaluator scores correlated with 90 Beijing Film Academy raters at Spearman ρ=0.95 on T2V and 0.96 on R2V; Seedance 2.0 topped both leaderboards.

A benchmark is a claim about what counts as good, and cinematic video just got a very opinionated one. In a paper posted to Hugging Face, researchers from Alibaba Group, Moku Lab at Hujing Digital Media & Entertainment, and Beijing Film Academy introduce FilmBench, a 1,169-prompt benchmark for text-to-video and reference-to-video models that is reverse-engineered from award-winning films across 20 cinematic genres.

The structure is where the ambition shows. FilmBench organizes evaluation into three L1 axes, Instruction Following, Temporal Continuity and Aesthetic Quality, then fans those out into 12 L2 components and 35 L3 sub-metrics covering things like shot scale, camera movement, viewing angle, composition, focus, and tone. Crucially, 1,056 of the 1,169 prompts script multiple shots following real shot lists, rather than the single-clip prompts most public benchmarks lean on. Prompts were drafted by Gemini 3.1 Pro plus a FilmOps operator suite, then rewritten by directing staff from Beijing Film Academy.

The headline validation number is a Spearman ρ of 0.95 on T2V and 0.96 on R2V against roughly 90 professional raters from Beijing Film Academy, across about 30% of the benchmark. On the leaderboard, Seedance 2.0 tops both tasks at 88.93 (T2V) and 86.66 (R2V), with the HappyHorse and Kling families close behind, and Hailuo 2.3 the clear laggard. Nobody saturates. The authors' reading is that the gap between models now lives in professional camera language and multi-shot staging, not in static frame quality, where scores already sit in the 90s.

The honest caveat is that this is a self-published benchmark from one of the labs whose models are being scored on it, and the top of both leaderboards is an Alibaba-adjacent Seedance model. The paper also does not spell out how source clips from Academy, Golden Horse and Hundred Flowers award winners were licensed for reverse-engineering, and it explicitly scopes out short-form vertical video and UGC. Whether non-Alibaba labs will adopt this taxonomy or publish counter-benchmarks is the interesting bit to watch.

For studios, post-production shops and anyone trying to write a real procurement checklist for generative video, though, a 35-sub-metric rubric written with a state film school beats another VBench-style aggregate score. That is the actual shift here.