paper web signal

Xiaomi-Robotics-1 tops VLA benchmarks on 100K-hour dataset

TL;DR

  • Xiaomi-Robotics-1 was pre-trained on over 100,000 hours of real-world trajectories collected via UMI devices across more than 1,700 scenarios.
  • On RoboDojo it averages 20.07 versus a prior industry best of 13.07, with success rate rising from 8.80% to 13.93%.
  • On VLABench it reports 59.1% average success and 70.3% average progress, and on RoboCasa it beats NVIDIA GR00T N1.6 and Pi-0.5.

Xiaomi's robotics group is making the bluntest possible bet in generalist manipulation, that if you throw enough real-world hours at a vision-language-action model, generalization falls out the other end. Their new paper on arXiv reports that Xiaomi-Robotics-1 was pre-trained on more than 100,000 hours of real-world trajectories collected via UMI (Universal Manipulation Interface) rigs across more than 1,700 scenarios spanning household, commercial premises, industrial sites, and outdoor spaces.

The headline numbers are on RoboDojo, where the team reports an average score of 20.07 against a prior industry best of 13.07, with success rate climbing from 8.80% to 13.93%. On VLABench they claim 59.1% average success and 70.3% average progress. And on RoboCasa the paper says Xiaomi-Robotics-1 surpasses NVIDIA's GR00T N1.6 and Physical Intelligence's Pi-0.5, two of the most-watched generalist manipulation models, alongside Cosmos Policy, Pi-0-FAST and RLDX-1.

Why this matters if you are not in robotics: for the last two years the received wisdom in VLA work has been that real-world trajectory data is the scarce input, more so than model architecture. A leaderboard result of this size is a direct stress test of that thesis, and if the gains hold up outside the benchmark harness, the practical moat in generalist manipulation shifts toward whoever can collect trajectory hours cheapest and label them at scale.

The honest caveat is that these are the team's own numbers, and 13.93% success on RoboDojo, even if it doubles the prior best, is still low in absolute terms. Benchmark leadership on curated task suites has repeatedly failed to translate to unfamiliar rooms and factory floors. What the reporting does not give you is the cost of collecting 100,000 hours, how much of that data is reusable across embodiments, or how the model behaves on the long tail of tasks where it still fails.

The direction is what to watch. If Xiaomi's play holds, expect NVIDIA, Physical Intelligence and the Chinese robotics startups around them to answer with their own scaled real-world datasets, and expect the vendors selling UMI-compatible collection rigs to have a very good next twelve months.