paper web signal

MIST benchmark: VLM judges flip on image presence, not content

TL;DR

  • In MIST, an aligned image shifted 20.5% of VLM labels and a misleading image shifted 19.4%, nearly the same across thirteen judges.
  • Only 37% of labels that differed between the two image conditions moved toward the sense the image actually depicted.
  • Human-annotator agreement was unchanged whether the image was absent, aligned, or misleading; only the VLM judges destabilised.

Across thirteen vision-language models asked to judge sentences, an aligned image changed 20.5% of labels and a misleading image changed 19.4%. The gap is almost nothing.

That is the headline result of "It's Not What the Image Shows", a paper by Nagham Omar, Mahmoud Jabarin, Kinan Ibraheem, and Lotem Peled-Cohen, accepted at the Trust-AI-Eval workshop at NeurIPS 2026. The authors built MIST, the Misleading-Image Stress Test, from 200 English sentences, each readable literally or figuratively, each shown with an aligned image depicting one reading, a misleading image depicting the opposite, or no image at all.

The authors' own summary: "what moves a judge is that an image is there, not which of the two it is."

Deleting the ignore-the-image instruction while leaving the image in place produced a smaller shift, 11.6% of labels, than either an aligned or misleading image. Only 37% of the labels that differ between the two image conditions moved toward the sense actually shown. Agreement with human annotators was unchanged whether the image was absent, aligned, or misleading, so the destabilisation sits inside the judges, not the task.

The effect was smaller in judges that pass the alt-test and larger in those that never do, but the paper reports it was present in all thirteen.