Found first: a primary source the press has not covered yet.
A new paper introduces PointWAM, a 3D world-action model that pre-trains on human hand-demonstration videos and improves average success on the DexJoCo dexterous manipulation benchmark by 56.9 percentage points over training from scratch. The method jointly forecasts scene and hand trajectories as 3D point clouds, treating environment and actor in a shared space-time frame rather than as separate prediction targets.
What the source says
PointWAM decomposes the manipulation world into scene and hands, forecasting how both co-evolve as 3D point trajectories given a colored point cloud and a language instruction. Pre-training on large-scale human demonstration videos yields the 56.9 percentage-point DexJoCo gain; adding scene-trajectory supervision on top of hand-only forecasting contributes a further 10.9 points. The model surpasses the prior multi-task state of the art across all ten DexJoCo tasks by 11.7 points and outperforms vision-language action model baselines on real-robot evaluation. Authors are Chunghyun Park, Beomjun Kim, Seungcheol Park, Heeseung Kwon, Yashu Shukla, Seunghoon Sim, Jinwoo Shin, and Minsu Cho; institutional affiliations are not listed on the arXiv abstract page.
Why it matters
The 56.9-point gain arrives without task-specific keypoint annotation or object selection, which has been the standard friction in transferring human demonstrations to robot hands. Representing manipulation as 3D point trajectories preserves contact geometry and spatial structure that 2D pixel or latent-space approaches discard. The scene-trajectory ablation result, contributing 10.9 additional points, suggests that the environment's movement carries signal that current baselines leave on the table. For labs trying to scale dexterous manipulation cheaply, this positions large libraries of human hand video as a viable pre-training corpus without bespoke annotation pipelines.