NeurIPS 2026: PhysVista Finds Frontier VLMs Score 18-43% on Physical Scale Estimation

Found first: a primary source the press has not covered yet.

A team from the University of Science and Technology of China, ByteDance, and City University of Hong Kong has published PhysVista, a benchmark accepted to NeurIPS 2026 Main Track that tests 17 vision-language models across a three-stage framework covering physical perception, dynamics reasoning, and plausibility assessment. The headline result: on quantitative scale estimation, scores across all 17 models range from 18.33% to 43.33%, with GPT-6 Sol, the top overall performer, scoring 35%.

What the source says

The benchmark draws on real-world video from WISA-80K and AI-generated video from VideoPhy-2, spanning 13 tasks across three stages. Models tested include GPT-6 Sol, Claude Opus 5.5, Gemini 3.1 Pro, Grok 4, Kimi K2.5, Doubao Seed 2.0 Pro, and several Qwen3.5 variants. On physical dynamics prediction, scores range from 22.96% to 58.52% across the field; on spatial state perception, the easier end of the benchmark, scores climb to between 74.5% and 85%. The paper also documents a "score collapse" failure on plausibility assessment: models concentrate predictions at the highest rating level for both real and AI-generated video, regardless of content. Six annotators verified 40% of the dataset across linguistic neutrality, precision, and validity.

Why it matters

Scale estimation is a prerequisite for most physical reasoning, not a specialist skill. Scores between 18% and 43% on that sub-task, from models that score 75% and above on spatial perception in the same benchmark, show that visual recognition and physical understanding decouple sharply in current VLMs. The score collapse on plausibility assessment is a distinct problem: models cannot rank physical realism within a video set, defaulting to uniform high scores rather than calibrated judgments. The benchmark covers both real and AI-generated video precisely to stress this loop, and results are consistent across both content types. Code and data are at github.com/Helen1p/PhysVista.