Apple-π Benchmark: Top Video Model Hits 0.473 on Physics Test
TL;DR
- The Apple-π benchmark tests 11 video generation models against 400 videos covering ten canonical classical mechanics tasks, and the best system scores only 0.473.
- Its three-stage protocol grades Perception, Formulation, and Deduction separately, treating the generated video as the model's visible reasoning trace rather than a pass/fail image.
- The authors report a Perception-to-Formulation-to-Deduction bottleneck, weak multi-law state transfer, and a persistent Sim-to-Real gap, arguing current models are far from reliable world simulators.
'World model' is doing a lot of load-bearing work in the current video-generation marketing cycle, and a new benchmark from a large academic team is a useful counterweight. The paper, posted to arXiv, tests 11 video generation models against 400 videos covering ten canonical classical mechanics tasks, and the best system scores only 0.473 on the authors' normalized scale. That is far from a passing grade for anything anyone would call a physical simulator.
The interesting design choice is that Apple-π refuses to grade only on whether the output looks plausible. Its three-stage protocol splits the work into Perception, Formulation, and Deduction, uses what the authors call chain-of-frames prompting on infographic-annotated first frames, and treats the generated video as the model's visible reasoning trace. A hybrid evaluation suite combines MLLM-based subjective scoring with physics-law-grounded objective measures, so you get a stage-resolved diagnosis of not just whether a model failed but where. That distinction matters, because a video that happens to look right for the wrong reasons is exactly what breaks when you try to use it downstream.
Why this matters if you are shipping or buying these models: the 'emergent world model' claim is the pitch that unlocks robotics simulation budgets, autonomous-vehicle training pipelines, and physics-aware creative tools. The paper's stage-, pillar-, and source-resolved analyses call out a Perception-to-Formulation-to-Deduction bottleneck, weak multi-law state transfer, and a persistent Sim-to-Real gap. Translated: models can sometimes see the setup, less often frame it as physics, and rarely deduce the correct rollout. Compound those and 0.473 stops looking surprising.
The honest caveat is that this is one paper on 11 models, and the readout depends on the authors' rubric; a different weighting between subjective plausibility and law-grounded objective measures would move numbers, though probably not the overall story. What the paper's abstract does not give you is a per-model leaderboard, so anyone doing procurement should press vendors for their own Apple-π-style stage-resolved scores rather than trusting a demo reel. The upside is that the benchmark is diagnostic by design: if failures really do concentrate in formulation and deduction, that is a much clearer research target than 'make the video look better'.
Originally reported by paper
Read the original article →Original headline: Apple-π Benchmark: Best Video Generation Models Score 0.473/1 on Physical Law Reasoning — 'Far From World Simulators'