huggingface.co web signal

RLHND Turns Cosmos 3 Video Model Into Hand-and-Force Tracker

TL;DR

  • RLHND adapts the Cosmos 3 Nano video diffusion backbone into a deterministic clip-level feature extractor via clean-latent conditioning.
  • Anatomical constraints cut MANO's 45 rotational degrees of freedom to 29, and a per-video cached shape parameter eliminates ~8mm size drift.
  • The paper reports state-of-the-art hand pose results on HOT3D, ARCTIC and zero-shot EgoDex, plus best contact and force scores on OpenTouch and PressureVisionDB.

A new paper from Seoul National University and RLWRLD takes Nvidia's Cosmos 3 Nano video diffusion model and repurposes it as a hand tracker that also reads contact and force. Posted to Hugging Face papers, the work is called RLHND and is authored by Subin Jeon, Sangwoo Kim, Hanbyul Joo and Jinwoo Shin.

The pitch is that monocular egocentric video is now the cheap input for robot policy training, but off-the-shelf hand trackers wobble: they regress pose per cropped frame, drift on hand size by about 8mm within a single clip, and emit no physical cues at all. RLHND's answer is to route video through Cosmos 3 Nano's conditioning-frame interface rather than its denoising path, carrying the model's priors on hand motion and hand-object interaction into tracking. The authors describe RLHND as "a video foundation model-based hand tracking model that jointly estimates hand pose and realistic tactile information from monocular egocentric videos."

Two design choices do most of the work. The MANO hand model's 45 rotational degrees of freedom are cut to 29 by freezing anatomically infeasible axes at the PIP and DIP finger joints, with the thumb left unconstrained because its saddle joint has non-orthogonal axes. A single shape parameter is then cached per video so the estimated hand does not change size frame to frame. On top of the pose stream, a separate tactile expert, trained with the pose stream frozen, "predicts dense contact and force over the hand surface," spread from 16 bone-level tokens to all 778 MANO vertices via linear-blend-skinning weights.

The paper claims state-of-the-art pose results on HOT3D and ARCTIC, zero-shot gains on the held-out EgoDex set, and best-in-class contact and force scores on OpenTouch and PressureVisionDB. The authors also report that "RLHND achieves state-of-the-art performance across various benchmark datasets for pose estimation, while also achieving state-of-the-art performance in contact and force estimation," and demonstrate improved success rates when retargeted trajectories feed a Dexterous Point Policy on real hardware. The abstract does not attach numbers to those robot-task gains.

It lands into an active patch of robotics coverage on our tracker, including yesterday's Mecka Series B for robot motion data. Code is promised at seungjun-moon.github.io/rlhnd.