paper web signal

VDiff-Bench: Grok 4.3 trails Kimi on low-level image diffs

TL;DR

  • VDiff-Bench pairs 1,756 four-way multiple-choice questions across 10 change categories to test whether MLLMs can spot the specific difference between two similar images.
  • Grok 4.3 hits 82.0% on semantic differences but 40.7% on low-level ones, collapsing to 5.3% on noise and 15.3% on texture.
  • Overall scores across the 11 tested MLLMs range from 35.8% to 89.6%, with Kimi K2.5 leading low-level accuracy at 88.8%.

Grok 4.3 answers 82.0% of VDiff-Bench's semantic image-comparison questions correctly, then falls to 40.7% on low-level changes, and to 5.3% on noise and 15.3% on texture specifically. Open-source Kimi K2.5, by contrast, posts 88.8% on the same low-level split, with Kimi K3 at 82.8%.

The benchmark, described in an arXiv paper by Yixin Wan, Tianle Zheng and Kai-Wei Chang, is 1,756 four-way multiple-choice questions over paired images across ten change categories including motion, illumination, OCR/text and appearance. Each item pairs the true difference with hard-negative descriptions and a 'no difference' distractor, curated to force models to 'distinguish the actual change from nearby semantic alternatives,' the authors write.

Across 11 open- and closed-source MLLMs tested, overall scores span 35.8% to 89.6%. The authors conclude that 'fine-grained visual comparison remains brittle,' with 'persistent failures on subtle low-level changes like noises and textures.'