arxiv.org web signal

VLM spatial binding runs through vision encoder, paper finds

TL;DR

  • Two mechanisms represent spatial binding in VLMs, but the vision encoder is the dominant source, not the language model backbone.
  • Spatial information in the vision encoder is distributed globally across visual tokens, extending beyond objects into background regions.
  • Globally amplifying vision-derived spatial representations corrects binding failures on COCO across models of various sizes.

Vision-language models bind objects to spatial positions through two parallel pathways, and one of them does almost all of the work. That is the claim of a new arXiv paper from Kelly Cui, Nikhil Prakash, Shoval Messica, Ayush Raina, David Bau, Antonio Torralba and Tamar Rott Shaham, which argues that intermediate layers of the language model backbone compute content-independent spatial relations over visual tokens but that 'this mechanism plays only a secondary role in shaping model predictions.'

The dominant source sits upstream. 'The dominant source of spatial information originates in the vision encoder, whose representations encode the layout of objects and are directly exploited by the language model backbone,' the authors write. The signal is also not where you would expect it: it is 'distributed globally across visual tokens, extending beyond object regions into surrounding background areas.'

On COCO images, the authors report that 'globally amplifying the vision-derived spatial representations across all image tokens corrects spatial variable binding failures across models of various sizes.' The abstract names no specific VLMs and publishes no per-model accuracy figures. Two of the researchers we follow surfaced the preprint in the days after its posting.

Shared on Bluesky by 2 AI experts