WROP releases 1.5M-sample dataset to train object permanence
TL;DR
- WROP publishes 1.5 million training samples and a 300-question exam built from 150 tasks across six developmental-psychology-inspired categories.
- Fourteen video models were scored in a blind pairwise Elo study, with the authors' 16-billion-parameter PWM-WROP finishing first among continuation models and third overall.
- The release includes data, exam, model answers, scores, weights and PWM, a native-PyTorch training stack running on AWS Trainium2.
Object permanence has a benchmark now. A paper on arXiv releases WROP, or 'World Reasoning with Object Permanence,' which the authors describe as 'a data infrastructure of 150 hand-designed cognitive science inspired tasks, divided into six cognitive categories.' The training corpus is 1.5 million samples; the exam is 300 questions.
The scenes come from Blender generators that 'randomize speed, lighting, camera angle, and other nuisance parameters while preserving each task's cognitive structure,' with more than 10,000 samples per task.
Fourteen video models were run through the exam: three reference-to-video, seven edit, and four continuation. The authors' own 16-billion-parameter model, PWM-WROP, 'ranks first among continuation models and third overall, behind only a statistical tie between two reference-to-video models,' the paper reports of a blind pairwise Elo study.
Also released: the data, exam, model answers, scores, weights, and PWM, described as a 'native-PyTorch training stack on AWS Trainium2.'
Originally reported by paper
Read the original article →Original headline: WROP: First Cognitive-Science-Inspired Dataset Trains Object Permanence Into Video World Models