HuRo's human-video pretraining lifts robot completion to 80.3%
TL;DR
- Pretraining on 142M robotized human-video frames raised a VLA policy's task completion from 51.5% to 80.3% across four real-world manipulation tasks.
- Out-of-distribution completion under spatial and visual shifts rose from 34.9% to 72.2% as the human-video pretraining data scaled up.
- Ablations found end-to-end pretraining with retargeted actions outperformed visual-only transfer for closing the human-to-robot embodiment gap.
A pretraining pipeline that turns human videos into robot-compatible supervision lifted a vision-language-action policy's completion on four real-world manipulation tasks from 51.5% to 80.3%, according to a paper accepted at CoRL 2026.
Under spatial and visual shifts, the same scaling took out-of-distribution completion from 34.9% to 72.2%.
The authors assemble HuRo, 'about 630K robotized episodes and 142M processed frames from five human-video sources,' the abstract states. Their pipeline 'converts heterogeneous human videos into robot-aligned observations and action trajectories while inferring missing intermediate signals across annotation levels.'
Two ablation findings carry the argument. The paper reports that 'visual robotization improves OOD robustness,' and that 'end-to-end pretraining with retargeted actions outperforms visual-only transfer.' The abstract does not name the base VLA policy, the four tasks, or the five source datasets, and reports only aggregate completion rather than per-task figures.
Originally reported by paper
Read the original article →Original headline: HuRo (CoRL 2026): 142M Human Video Frames Lift Real-World Robot Completion From 51% to 80%