huggingface.co web signal

'Show, Don't Tell' Benchmark Puts Image Generators Against VLMs on Spatial Reasoning; GPT Image 2 Solves 37% of Items GPT-5.4 Misses

Summary

The ProVisE framework and SpatialGen-Bench (470 samples, 14 spatial subtasks, 4 capability levels) evaluate 31 model-interface systems by demanding pixel-level answers rather than text or coordinates. Best text model GPT-5.4 hits 61.0% overall vs 54.5% for GPT Image 2, but the image model 'visually rescues' 37% of samples GPT-5.4 fails — evidence that generative pixels and text VLMs are complementary rather than substitutable. Human reference: 87.8%.