Found first: a primary source the press has not covered yet.
A paper from Joy Future Academy introduces ZimaBlue, a World Action Model that reaches 77.8% zero-shot success on real-robot manipulation tasks, up from 36.1% on target-robot data alone. The gain comes from scaling to over 120,000 hours of egocentric video, most of it human footage carrying no robot action labels. The preprint is led by Xionghao Wu, with Nan Duan and Haoyang Huang among its 20 co-authors.
What the source says
ZimaBlue uses a three-stage curriculum: causal embodied video pre-training on human and robot egocentric footage, a video-action mid-training phase that grounds the learned visual dynamics in heterogeneous robot trajectories through a unified action representation, then specialization to a target robot. The model pairs a 5B-parameter Slow world model with a 0.5B-parameter Fast action branch, an architecture that achieves 30 Hz closed-loop action prediction on a single NVIDIA RTX 4090. The 36.1%-to-77.8% improvement is measured by scaling from target-robot data alone to over 120,000 hours of embodied video. Gains are particularly pronounced on unseen tasks.
Why it matters
Robot demonstration data is expensive to collect at scale. Unlabeled egocentric video, whether from humans or other robots, is not. ZimaBlue's result suggests the gap between them is smaller than assumed: over 120,000 hours of video requiring no action labels produces more than twice the success rate of target-robot trajectories alone. The sharpest gains on unseen tasks are the more significant finding, since performance on training-distribution tasks can always improve with more target-domain data. A single consumer GPU handles inference at 30 Hz, reducing the hardware barrier for deployment outside the lab.