paper web signal

RoboTok Mines Web Video to Train Dexterous Robot Policies

TL;DR

  • RoboTok retrieves human hand demonstrations from web video to train dexterous robot policies, treating the internet as a continually growing supervision source.
  • It learns a latent motion space from 3D hand trajectories in actor-centered reference frames, invariant to camera viewpoint, scene appearance, and occlusions.
  • Nine authors, eight from Rice University and one from NVIDIA; the preprint was posted to arxiv on September 2, 2026.

A new preprint from Rice University and NVIDIA proposes treating the open web as a continuously indexable training corpus for dexterous robot hands. The system, described on arxiv and posted September 2, 2026, is called RoboTok: 'an internet-scale data engine that, given a query human manipulation video, retrieves manipulation-relevant human demonstrations from web videos for training dexterous robot policies.'

The technical move is to represent hand motion in a way that survives the mess of internet video. The authors 'learn a latent motion space from 3D hand trajectories expressed in estimated actor-centered reference frames.' That embedding, they write, 'enables manipulation behaviors to be compared across variations in camera viewpoint, scene appearance, and actor occlusions, while remaining compact enough for efficient search and continual indexing over internet-scale video collections.'

The pitch is a supply-side one. 'Robot learning increasingly depends on broad and diverse demonstrations, yet collecting robot data remains expensive and poorly suited to covering the long tail of real-world tasks,' the paper opens, and RoboTok is offered as a way around that bottleneck: instead of paying to teleoperate every new skill, retrieve the human hand doing it, somewhere on the web.

Against unnamed baselines, RoboTok 'retrieves more relevant manipulation demonstrations and improves downstream task success,' the authors report. The abstract does not publish the magnitude of the improvement, list the baselines, or describe the robot hardware used for the downstream evaluation. Nine authors are listed; eight are at Rice, and one, Bowen Wen, is at NVIDIA.