PhysVista: top VLM scores just 35% on physical scale tasks
TL;DR
- GPT-6 Sol, the top-ranked model, scored 81% on spatial state perception but only 35% on PhysVista's quantitative scale estimation task.
- On harder stages, GPT-6 Sol managed 46.67% on physical dynamics prediction and 44.54% on counterfactual reasoning.
- Plausibility correlations ranged from 0.1945 to 0.6141 Spearman, with the paper describing a 'score collapse' toward the highest ratings.
PhysVista, a NeurIPS 2026 benchmark that threads perception, reasoning, and plausibility assessment into a single evaluation loop, finds that frontier vision-language models fall off sharply when asked to make quantitative physical judgments. In the paper, the top performer, GPT-6 Sol, scored 81% on spatial state perception but only 35% on quantitative scale estimation.
The authors evaluated 17 VLMs including GPT-6 Sol, GPT-5.2, Claude Opus 5.5, Claude Opus 4.6, Gemini 3.1 Pro, Doubao Seed 2.0 Pro, Grok 4, multiple Qwen variants, and Kimi K2.5. The framework distinguishes event-level and scale-level reasoning across real-world and AI-generated videos, and reformulates scale estimation to require explicit relative ratio responses as fractional values.
The reasoning stage looked worse. GPT-6 Sol scored 46.67% on physical dynamics prediction and 44.54% on counterfactual reasoning. The paper reports that "strong performance in observable state perception does not necessarily translate into equally strong causal reasoning," with models ranking third in perception often falling much lower on reasoning.
Plausibility judgments drifted further. Spearman rank correlations with ground-truth plausibility ratings ran from 0.1945 to 0.6141 across models, and the authors describe a "score collapse" phenomenon where models concentrated predictions on the highest plausibility ratings regardless of actual physical validity.
The authors call the overall result "a persistent gap between visual recognition and genuine physical understanding," pointing toward "more principled designs for physically grounded multimodal intelligence."
Originally reported by paper
Read the original article →Original headline: NeurIPS 2026 Benchmark: Frontier VLMs Score as Low as 18% on Physical Scale Tasks