arxiv.org web signal

InSituMeasure: best of 24 MLLMs scores 25.7% on gauges

TL;DR

  • The best of 24 MLLMs tested reached 25.7% joint value-unit accuracy and 51.8% confidence-diagnosis F1 on InSituMeasure.
  • The benchmark contains 2,922 real industrial monitoring scenes across eight functional categories of professional engineering instruments.
  • Authors trace failures to text-induced shortcuts, overconfident responses, and industrial noise including occlusion and viewpoint deviation.

The best of 24 state-of-the-art multimodal models tested on InSituMeasure scored 25.7% joint value-unit accuracy at reading industrial gauges. Confidence-diagnosis F1 topped out at 51.8%.

The benchmark, released this month by Chao Shen and co-authors, contains 2,922 real industrial monitoring scenes across eight functional categories of professional engineering instruments, with "dense gauge-attribute annotations and noise tags for failure diagnosis."

The authors are direct about the gap. They describe "a substantial gap between general multimodal competence and reliable situated measurement," and pin the failures on "text-induced shortcuts, overconfident responses, and authentic industrial noise, including mixed disturbances, viewpoint deviation, occlusion, and environmental interference."

Two experts in our Who's Who directory had already shared the preprint by the time it came across our feed.

Shared on Bluesky by 2 AI experts