arxiv.org web signal

Koepke, Efros paper: cross-modal convergence breaks at scale

TL;DR

  • Scaling from 1,024 text-image pairs to 15 million drops measured cross-modal alignment from 13.5% to 0.81%, the paper reports.
  • Testing 55 language models, the authors found that newer and stronger LLMs do not seem more aligned with vision models.
  • One-to-one image-text pairings inflate alignment scores; real many-to-many data with multiple captions per image reduces it further.

Scaling from 1,024 text-image pairs to 15 million drops the measured alignment between vision and language models from 13.5% to 0.81%. That is the result at the center of a new arXiv paper by A. Sophia Koepke, Daniil Zverev, Shiry Ginosar and Alexei A. Efros, which argues the much-cited Platonic Representation Hypothesis does not hold up at realistic data scales.

'The Platonic Representation Hypothesis posits that neural networks trained on different modalities (e.g., text and images) converge toward a shared representation of reality,' the abstract says. The paper's response is blunt: 'the experimental evidence for this hypothesis is fragile and depends critically on the evaluation regime.'

Two design choices did most of the work in the original story. The first was dataset size. With roughly 1,024 samples, the mutual-nearest-neighbor metric looks like it is detecting a shared geometry, but the signal degrades as the pair count climbs into the millions. The second was 'one-to-one image-text pairings,' when real data is 'many-to-many,' and adding multiple captions per image further reduces alignment. The same behavior shows up, the paper reports, 'for text-audio and text-video alignment.'

Testing 55 language models, the researchers found that newer and stronger LLMs 'do not seem more aligned' with vision models, cutting against the prediction that scale drives representational convergence. What remains, they write, is 'coarse semantic overlap rather than consistent fine-grained structure.'

Shared on Bluesky by 2 AI experts