huggingface.co web signal

Scal3R Cuts KITTI Trajectory Error Over 60% With Pose Queries

TL;DR

  • Scal3R reports over 60% lower average absolute trajectory error on KITTI than its online baseline, with training converging in 8 hours on one GPU.
  • The authors argue online 3D reconstruction fails on long video because regressing poses to a fixed first-frame anchor pushes inputs outside the training distribution.
  • Learnable tokens making up about 1% of parameters are injected into a frozen backbone to query pose relative to multiple past keyframes.

A new paper called Scal3R reports it cuts average absolute trajectory error by more than 60% on KITTI against an online baseline, with training converging in eight hours on a single GPU.

The diagnosis is the interesting part. The authors argue that online 3D reconstruction breaks on long video for a specific structural reason: "regressing poses relative to a fixed first-frame anchor forces extrapolation far beyond the training distribution," causing small drifts to accumulate into "significant geometric collapse." But they observe that per-frame depth stays reliable through this failure. In their words, the "backbone's local geometry remains intact; only the global pose head breaks down."

Rather than retrain the backbone, Scal3R injects lightweight learnable tokens, "about ~1% of the parameters," into a completely frozen model via asymmetric attention, and queries pose relative to multiple past keyframes instead of a single fixed anchor. An online pose-graph optimization with loop closure suppresses long-range drift.

Beyond the KITTI headline, the authors claim state-of-the-art performance across Virtual KITTI, Sintel, TUM-Dynamic, ScanNet, and 7-Scenes. The abstract publishes no per-dataset numbers, no ablations, and no named comparisons to prior systems. It lands in a run of online 3D geometry results our tracker has been logging this quarter, alongside LatentStream's evolving latent video memory from the same day.