Zhejiang SpaceCast-Bench: Top VLM Hits 58% vs 87.2% Humans
TL;DR
- Zhejiang University's SpaceCast-Bench covers 3,862 questions across 182 real-world scenes and 16 task types at three levels: static perception, local prediction, and global prediction.
- Of 21 models evaluated, the strongest reaches 58.0% versus a 87.2% human baseline, while models branded spatially specialized remain near random chance.
- Fine-tuning on the authors' programmatically generated data lifts Qwen3-VL-4B from 34.0% to 65.7%, with macro-average gains across six out-of-domain benchmarks.
The strongest vision-language model tested on SpaceCast-Bench, a new benchmark from Zhejiang University, scored 58.0% against a human baseline of 87.2%, with models the authors describe as "spatially specialized" landing "near random chance."
The benchmark targets predictive spatial reasoning rather than the perceptual tasks most spatial evaluations lean on. Its 3,862 questions span 182 real-world scenes and 16 task types at three levels (static perception, local prediction, and global prediction), organized around what the paper calls an "observe-transform-infer framework" that asks a model to build a scene from observations, anticipate how an intervention changes it, and reason about an unseen outcome.
Two findings sit beyond the headline gap. The paper reports that "bridge views are critical for integrating distributed observations," and that "explicit 3D evidence benefits models more reliably than generated outcome images or videos." Fine-tuning on the authors' programmatically generated data lifted Qwen3-VL-4B from 34.0% to 65.7%, with macro-average gains across six out-of-domain benchmarks.
Twenty-one models were evaluated in total. The abstract does not name which one holds the 58.0% top score. Code and the dataset are released on GitHub and Hugging Face; the paper joins a run of multimodal-benchmark releases on our tracker over the past quarter.
Originally reported by huggingface.co
Read the original article →Original headline: Zhejiang's SpaceCast-Bench Finds Top VLM Scores 58% on Predictive Spatial Reasoning vs 87% Human Baseline