Meta FAIR Ships 8B RoboJEPA With Multi-Embodiment Scaling Laws
TL;DR
- Meta FAIR's RoboJEPA trains on 15,022 hours of robot video across 23 datasets and 12 embodiments, scaling from 22M to 8B parameters.
- Imagination error follows a second-order power law L(C) = E + A·C^(α - γ ln C) that extrapolates fits made on 22M–2B models to held-out 4B–8B runs.
- Capabilities emerge at compute thresholds: end-effector reaching near 10^20 FLOPs, obstacle avoidance near 10^21, fine object dynamics near 10^22.
Meta's FAIR lab trained a robot world model on 15,022 hours of video across 23 manipulation datasets and 12 embodiments, scaled a Joint Embedding Predictive Architecture up to 8B parameters, and open-sourced the checkpoints together with the training and robot-deployment code. The paper, led by Artem Zholus with joint last authors Jeannette Bohg, Nicolas Ballas and Mahmoud Assran, calls this "the first work to establish scaling laws for multi-embodiment robotic world models trained on real robot data, and RoboJEPA, at 8B parameters, is the largest JEPA predictor model trained to date."
The core claim is a fitted curve. The authors write that "RoboJEPA's imagination error, the error of its latent rollouts, follows a second-order power law in compute, allowing us to predict model quality well beyond the scale at which the law is fit." The form is L(C) = E + A·C^(α − γ ln C). On the DROID holdout the fit lands with an extrapolation error of 0.6×10⁻³, against 2.0×10⁻³ for a first-order power law — fits made on 22M-to-2B models project onto held-out 4B and 8B runs. The irreducible-error term E comes in at 0.218 on DROID and 0.1709 on RoboCasa, numbers the authors read as near-saturation with the current data budget.
Capabilities land at specific compute thresholds. End-effector 3D reaching emerges around 10^20 FLOPs, obstacle avoidance around 10^21, and fine object dynamics such as pushing around 10^22; long-horizon non-greedy planning sits above 10^22, which is the 2B-to-8B band. The paper also reports that "downstream robotic planning performance improves predictably with compute, and that imagination error is strongly correlated with it, making it a reliable proxy for real-robot evaluation."
Real-world evaluation runs on a single-arm Franka through the DROID stack across grasp, object lift and pick-and-place, with planning handled by the Cross-Entropy Method against a single goal image. The abstract reports trends with scale rather than per-task success numbers, and the authors flag that comparisons with the VLA baselines π₀-FAST and π₀.5 are contextual because the goal specifications differ. Code, checkpoints and the deployment stack are at github.com/facebookresearch/robo_jepa. The release lands into a busy week for the architecture, following a separate JEPA-latent recovery result we covered the same day.
Originally reported by huggingface.co
Read the original article →Original headline: Meta FAIR Ships RoboJEPA, Open-Sources 8B Latent World Model With Scaling Laws for Multi-Embodiment Robots