16B World Model Trained on 1.5M Object-Permanence Clips Ranks First in Continuation Category

Found first: a primary source the press has not covered yet.

A 31-author team has released WROP, a 1.5-million-sample training corpus built from 150 cognitive-science tasks designed to teach video generation models object permanence and physical solidity. Their 16-billion-parameter world model, PWM-WROP, ranks first among continuation-style video models and third overall in a blind Elo evaluation covering 14 video models.

What the source says

WROP spans 150 hand-designed tasks across six cognitive categories, with 10,000-plus Blender-generated samples per task, totaling 1.5 million training clips. The evaluation is a 300-question exam scored via blind Elo across 14 models: 3 reference-to-video, 7 edit, and 4 continuation. PWM-WROP placed first among continuation models and third overall, behind a statistical tie between two reference-to-video models at the top. The release includes all training data, the exam, model answers, scores, weights, and a native-PyTorch training stack tested on AWS Trainium2. Lead author is Haotian Zhang; co-authors include Yilun Du, Alan Yuille, and Nikolaus Kriegeskorte among the 31-person team. Institutional affiliations are not listed on the abstract page.

Why it matters

Video generation benchmarks have typically measured perceptual quality or motion smoothness, not whether a model understands that a hidden object continues to exist. WROP is the first large-scale training corpus grounded explicitly in developmental psychology principles, giving researchers a concrete way to probe and improve this specific class of physical reasoning. The open release of training data, benchmark, and weights means other teams can build on this scaffold directly rather than reconstructing it. The third-place overall finish behind reference-to-video models, which receive explicit visual anchors, sets a measurable baseline for where continuation models currently stand on cognitively grounded reasoning.