InfiniHand Merges Hand Pose and SLAM in One Streaming Model
TL;DR
- InfiniHand jointly estimates MANO hand parameters, camera trajectory, and hand location from uncalibrated egocentric video in a single feed-forward model.
- The paper reports a 21.4% reduction in ARCTIC PA-p error versus ViDiHand and 11.19 FPS throughput, over twice HaWoR.
- Training runs in two progressive stages on roughly 5,000 hours of egocentric video aggregated from multiple public datasets.
A single feed-forward model can now jointly recover 3D hand pose, camera trajectory, and world-space hand location from egocentric video, according to InfiniHand, a new paper from Kerui Ren and collaborators. The team reports a 21.4% reduction in ARCTIC PA-p error against ViDiHand and 11.19 FPS throughput, which the abstract calls "more than twice the throughput of HaWoR."
The pitch cuts against the usual cascade of a hand pose estimator on top of a SLAM system, which the authors argue produces "error accumulation, complex pipelines, and severe computational overhead." InfiniHand instead "integrates persistent spatiotemporal memory with hand-centered visual features, explicitly coupling camera motion with local hand geometry within a unified architecture." Training runs in two progressive stages, learning camera-space hand priors first and then extending to streaming world-space reconstruction, on roughly 5,000 hours of egocentric video aggregated from multiple public datasets.
The abstract publishes only the headline figures. It does not specify the hardware behind the 11.19 FPS measurement, name the datasets that make up the 5,000-hour corpus, or report how world-space drift behaves over long sequences.
Originally reported by arxiv.org
Read the original article →Original headline: InfiniHand Paper Streams World-Space Hand Motion From Egocentric Video at 11 FPS