paper web signal

Moonshot's PerceptionBench: No Frontier MLLM Clears 60%

TL;DR

  • PerceptionBench isolates ten atomic visual skills across 3,000 verified questions, and no model in a field of sixteen frontier MLLMs cleared 60% accuracy.
  • Perception-related hallucination is the weakest capability on average, and many correct answers fail to reproduce when the same question is asked again.
  • Models with similar overall scores show sharply divergent profiles across the ten categories, so top-line accuracy hides skill-specific weaknesses.

A new paper from Moonshot argues something deflating about the current wave of vision-language models: on 3,000 questions carefully stripped of reasoning and world knowledge, none of sixteen frontier systems cleared 60% accuracy.

The benchmark, called PerceptionBench, is built bottom up. The authors went back through 42 existing multimodal benchmarks, catalogued where frontier models fail, and distilled those failures into ten atomic perceptual capabilities: visual relation, counting, attribute recognition, depth and 3D, localization, comparison, fine-grained recognition, context integration, OCR, and hallucination. Each of the 3,000 questions targets a single one of those skills and has a short, unambiguous answer. According to the paper on arxiv, the point is to isolate seeing from thinking.

Two findings do more work than the headline number. First, perception-related hallucination is the weakest capability on average across the sixteen models, meaning the systems most often fail not by misreading what is in front of them but by asserting things that are not. Second, models with almost identical overall scores can show sharply divergent profiles across the ten categories, so the aggregate leaderboard figure a vendor quotes can hide a real weakness in the specific skill you happen to be building on.

The honest caveat is that this is a benchmark paper from a group that also ships a frontier model in the mix, and the authors' own write-up on the Kimi blog has Kimi K3 out front. Take the specifics as reported, not settled. What the retrieved reporting does not give you is a per-model score table across the ten skills, so you cannot yet look up whether the model you use is the one that flunks OCR or the one that flunks counting.

The interesting move for anyone shipping product on top of these systems is not to argue with the benchmark but to copy its posture. If a large share of correct answers fail to reproduce when the same question is asked again, confidence in any single response is a poor signal, and evaluating your own workflow on per-capability slices, does it read the label reliably, does it count the boxes reliably, is probably a better use of an afternoon than tracking the next leaderboard shuffle.

Shared on Bluesky by 1 AI expert