arxiv.org web signal

InfiniHand Merges Hand Pose and SLAM in One Streaming Model

Computer Vision Multimodal ai-research computer-vision

TL;DR

  • InfiniHand jointly estimates MANO hand parameters, camera trajectory, and hand location from uncalibrated egocentric video in a single feed-forward model.
  • The paper reports a 21.4% reduction in ARCTIC PA-p error versus ViDiHand and 11.19 FPS throughput, over twice HaWoR.
  • Training runs in two progressive stages on roughly 5,000 hours of egocentric video aggregated from multiple public datasets.

A single feed-forward model can now jointly recover 3D hand pose, camera trajectory, and world-space hand location from egocentric video, according to InfiniHand, a new paper from Kerui Ren and collaborators. The team reports a 21.4% reduction in ARCTIC PA-p error against ViDiHand and 11.19 FPS throughput, which the abstract calls "more than twice the throughput of HaWoR."

The pitch cuts against the usual cascade of a hand pose estimator on top of a SLAM system, which the authors argue produces "error accumulation, complex pipelines, and severe computational overhead." InfiniHand instead "integrates persistent spatiotemporal memory with hand-centered visual features, explicitly coupling camera motion with local hand geometry within a unified architecture." Training runs in two progressive stages, learning camera-space hand priors first and then extending to streaming world-space reconstruction, on roughly 5,000 hours of egocentric video aggregated from multiple public datasets.

The abstract publishes only the headline figures. It does not specify the hardware behind the 11.19 FPS measurement, name the datasets that make up the 5,000-hour corpus, or report how world-space drift behaves over long sequences.