PAWBench Tests 11 Video Generators; None Consistently Matches Reference Distributions

Found first: a primary source the press has not covered yet.

A new benchmark called PAWBench evaluates eleven video generation models across 50 physical scenarios and finds that none consistently matches the correct probability distribution over possible outcomes. The paper, submitted in August 2026, formalizes a distributional criterion for world model claims that no current system meets.

What the source says

The fourteen-author team introduces PAWEval, an evaluation protocol that converts repeated video rollouts into empirical distributions over physical behaviors and compares them against reference probabilities. The core argument is that a video generator functioning as a world model must reproduce the correct distribution of possible outcomes from a given initial observation, rather than generating a single plausible trajectory. Across 50 scenarios and all eleven systems tested, the paper reports that no model consistently matches the reference probabilities while recovering the range of valid behaviors. The researchers also tested whether language prompts, noise sampling, or model training could reshape predictive distributions toward alignment, and found these insufficient. Affiliations are not listed in the abstract.

Why it matters

Several major video generation labs have described their systems as world models, a framing that implies the system encodes something like causal physics rather than learned visual correlations. PAWBench provides a concrete, testable criterion for that claim: distributional alignment across repeated rollouts from the same initial condition. By that criterion, no current system qualifies, across a 50-scenario evaluation spanning eleven models. The paper also establishes that plausible, diverse, or controllable generations do not, by themselves, constitute probabilistically aligned world modeling, which constrains what evidence can legitimately support world model claims going forward.