Principia benchmark: SOTA video models cap at 0.42 on physics
TL;DR
- Six state-of-the-art video generators score no higher than 0.42 on Principia, a new benchmark that measures physics through relational consistency between paired objects.
- The same models score around 0.8 on VBench, suggesting mainstream video-quality scores do not reflect whether outputs obey Newtonian physics.
- Vision-language models struggle to police physics violations: the best reaches 67% accuracy and most perform near chance level.
Six state-of-the-art video generators cap at 0.42 on Principia, a new benchmark that scores physical realism through relationships between objects in the same scene rather than absolute motion measurements. The same models score roughly 0.8 on VBench, the widely cited video-quality benchmark, according to the paper posted to arXiv by Varun Varma Thozhiyoor and colleagues.
The authors argue that absolute physics evaluation is awkward for generated video, because "absolute motion measurements depend on frame rate, object scale, and camera calibration, all of which are often ambiguous or unavailable in generated video." Their workaround: when two objects share a scene and obey the same physical law, the ratios of their motions must line up, and those ratios hold independent of calibration.
Principia covers eight phenomena: gravity, restitution, friction, rotational inertia, projectile motion, momentum, pendulum, and mass-spring oscillation. All are recorded from real-world scenes under controlled protocols, spanning translational, rotational, collisional, and oscillatory dynamics. Across thousands of generations, "no model exceeds 0.42 on Principia despite all scoring around 0.8 on VBench," the abstract reports.
The paper also puts vision-language models to work as judges of physics violations. "The best model achieving only 67% accuracy and most performing near chance level," the abstract states. The Principia team publishes no per-generator score in the abstract itself.
Originally reported by paper
Read the original article →Original headline: Principia Benchmark: Six SOTA Video Generators All Score Below 0.42 on Relational Physics, vs 0.80 on VBench