arxiv.org web signal

Language models encode color structure without visual input

TL;DR

  • The paper reports significant structural correspondence between text-derived color-term representations and CIELAB, a perceptually meaningful color space.
  • Warmer colors align better than cooler ones with the perceptual color space, echoing recent work on efficient communication in color naming.
  • Differences in alignment are partly mediated by collocationality and syntactic usage patterns in text.

Pretrained language models learn color relationships that correspond with the perceptual structure of color space, even though they have never seen a color. That is the central finding of "Can Language Models Encode Perceptual Structure Without Grounding? A Case Study in Color," a CoNLL 2021 paper by Mostafa Abdou, Artur Kulmizev, Daniel Hershcovich, Stella Frank, Ellie Pavlick, and Anders Søgaard.

The team compared text-derived representations of monolexemic color terms to color chips in CIELAB, "a color space with a perceptually meaningful distance metric." Using two methods of evaluating structural alignment, they report "significant correspondence" between the two.

The alignment is uneven. "Warmer colors are, on average, better aligned to the perceptual color space than cooler ones," the authors write, flagging an "intriguing connection to findings from recent work on efficient communication in color naming." Further analysis attributes part of the gap to collocationality and differences in syntactic usage, "posing questions as to the relationship between color perception and usage and context."

Shared on Bluesky by 2 AI experts