paper web signal

Pumpire tests 29 3D models on real point-to-point distances

TL;DR

  • Pumpire evaluates 29 baseline configurations of image- and video-level 3D foundation models on point-to-point metric distance rather than depth or intrinsics alone.
  • The dataset holds 100 real-world scenes and 6,400 frames (64 per scene), indoor and outdoor, annotated with physically measured point-pair distances.
  • The authors argue existing protocols 'cannot directly reflect' models' point-to-point distance estimation, a scale question they say prior work 'largely overlooked.'

A new arxiv benchmark called Pumpire argues that the standard way of grading 3D foundation models is measuring the wrong thing. Instead of scoring depth maps and camera intrinsics as two separate numbers, it asks a plain-spoken question: pick two points in the reconstructed scene, measure the distance between them, and see how close the model's answer gets to a tape-measure reading.

The dataset carries 100 real-world scenes and 6,400 frames at 64 per scene, indoor and outdoor, each annotated with "physically measured point-pair distances." The protocol back-projects a model's predicted depth together with its predicted camera intrinsics into geometry, then compares the resulting point-to-point distance with the physical measurement.

Against that stick, the authors run 29 baseline configurations of "image- and video-level 3D foundation models, with or without depth priors."

The methodological complaint is blunt. Previous approaches that evaluate depth and intrinsics separately, or compare point clouds with geometric similarity metrics, "cannot directly reflect models' point-to-point distance estimation capability," and physical scale perception has, the paper says, been "largely overlooked."

The summary surfaced here publishes the dataset design, the 29-configuration coverage and the protocol; it does not spell out per-model error figures for which systems held up and which did not.