paper web signal

4DCodeBench: AI codes static scenes well, flails on dynamics

TL;DR

  • 4DCodeBench asks agents to turn video into executable graphics code that reconstructs 100 real and 100 synthetic scenes.
  • Across 18 frontier models, static reconstruction looks solved while deformation, fluid flow, and fracture remain unreliable.
  • Top model GPT-6 Astra hits 0.91 on perceptual and 3D geometry but only 0.60 on 2D dynamics.

Frontier AI models can write executable graphics code that reconstructs how a scene looks from video, but reliably reproducing how it moves — deformation, fluid flow, fracture — is still out of reach. That is the headline of 4DCodeBench, a new benchmark posted to arXiv on October 2 by a team spanning Stanford, Johns Hopkins and MIT.

The setup asks agents to turn a video into an "executable graphics program" that renders back out as a 4D scene. The authors curated 100 real-world videos and 100 synthetic scenes covering diverse physical phenomena, and ran 18 frontier models through them.

The paper's blunt summary, from the abstract: "strong static reconstruction capabilities do not yet translate into reliable reconstruction of complex dynamics." The gap shows up metric by metric. The top performer, GPT-6 Astra at its highest reasoning setting, scores 0.91 on perceptual quality and 0.91 on 3D geometry but only 0.60 on 2D dynamics, for an overall of 0.79.

How agents choose to represent motion matters: "67% of solutions describe motion analytically; stronger reasoning leads to more simulation and better dynamics reconstruction," the authors report. Pushing Astra's reasoning effort from Low to Max also lifts its 3D dynamics score from 0.63 to 0.73, suggesting the dynamics bottleneck is at least partly a reasoning-and-simulation problem rather than a perception one.