Six Video Generators Score Below 0.42 on Physics Bench, ~0.80 on VBench

Found first: a primary source the press has not covered yet.

A new benchmark called Principia tests six state-of-the-art video generators on Newtonian physics and finds none scores above 0.42, while the same models score approximately 0.80 on VBench. The paper, "Principia: Relational Physics Tests for Video Models", was submitted September 3, 2026.

What the source says

The work comes from Varun Varma Thozhiyoor, Shivam Tripathi, Venkatesh Babu Radhakrishnan, and Anand Bhattad. Principia covers eight Newtonian phenomena: gravity, restitution, friction, rotational inertia, projectile motion, momentum, pendulum dynamics, and mass-spring oscillation, spanning translational, rotational, collisional, and oscillatory motion. The metric is calibration-independent, measuring relational consistency between paired objects in the same scene rather than absolute motion, requiring no frame rate, scale, or camera data. Across thousands of generations, no tested model exceeded 0.42. The best vision-language model evaluated for physics-violation detection reached 67% accuracy; most performed near chance.

Why it matters

VBench scores around 0.80 have functioned as shorthand for video generation quality. Principia shows those scores do not capture whether a model generates physically plausible motion. Any application that depends on physical correctness in video, including robot learning and physical simulation, faces models that fail basic Newtonian tests regardless of their VBench standing. Because the metric operates on relational consistency between objects rather than calibrated absolute quantities, it cannot be improved by learning scale or speed priors from training data.