paper web signal

Xiaomi-Robotics-1 tops VLA benchmarks on 100K-hour dataset

TL;DR

  • Xiaomi-Robotics-1 was pre-trained on over 100,000 hours of real-world trajectories collected via UMI devices across more than 1,700 scenarios.
  • On RoboDojo it averages 20.07 versus a prior industry best of 13.07, with success rate rising from 8.80% to 13.93%.
  • On VLABench it reports 59.1% average success and 70.3% average progress, and on RoboCasa it beats NVIDIA GR00T N1.6 and Pi-0.5.

Xiaomi's robotics group is making the bluntest possible bet in generalist manipulation, that if you throw enough real-world hours at a vision-language-action model, generalization falls out the other end. Their new paper on arXiv reports that Xiaomi-Robotics-1 was pre-trained on more than 100,000 hours of real-world trajectories collected via UMI (Universal Manipulation Interface) rigs across more than 1,700 scenarios spanning household, commercial premises, industrial sites, and outdoor spaces.

The headline numbers are on RoboDojo, where the team reports an average score of 20.07 against a prior industry best of 13.07, with success rate climbing from 8.80% to 13.93%. On VLABench they claim 59.1% average success and 70.3% average progress. And on RoboCasa the paper says Xiaomi-Robotics-1 surpasses NVIDIA's GR00T N1.6 and Physical Intelligence's Pi-0.5, two of the most-watched generalist manipulation models, alongside Cosmos Policy, Pi-0-FAST and RLDX-1.

For anyone outside robotics, the stake here is a specific one: for the last two years the received wisdom in VLA work has been that real-world trajectory data is the scarce input, more so than model architecture. A leaderboard result of this size is a direct stress test of that thesis, and if the gains hold up outside the benchmark harness, the practical moat in generalist manipulation shifts toward whoever can collect trajectory hours cheapest and label them at scale.

Take the numbers with the usual care: these are the team's own results, and 13.93% success on RoboDojo, even if it doubles the prior best, is still low in absolute terms. Benchmark leadership on curated task suites has repeatedly failed to translate to unfamiliar rooms and factory floors. The paper also leaves several things unsaid: the cost of collecting 100,000 hours, how much of that data is reusable across embodiments, and how the model behaves on the long tail of tasks where it still fails.

The direction is what to watch. If Xiaomi's play holds, expect NVIDIA, Physical Intelligence and the Chinese robotics startups around them to answer with their own scaled real-world datasets, and expect the vendors selling UMI-compatible collection rigs to have a very good next twelve months.