paper web signal

Alibaba DAMO Scales RynnBrain 1.1 to 122B, Leads Spatial Tests

TL;DR

  • RynnBrain 1.1 ships as a three-tier family in 2B, 9B, and 122B-A10B parameter configurations from Alibaba's DAMO Academy.
  • The 122B-A10B model reportedly outperforms all evaluated proprietary and open-source models on VSI-Bench, MMSI, and RefSpatial-Bench.
  • A companion RynnBrain-VLA policy is deployed on three robot platforms: Unitree G1, Astribot-S1, and Tianji-Wuji.

Alibaba's DAMO Academy has pushed its embodied AI stack up another rung, with a new paper on arXiv describing RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and a 122B-A10B tier aimed at giving robots a shared spatial and physical understanding of the world.

The pitch is that the biggest model outperforms every proprietary and open-source model the team evaluated on three spatial benchmarks: VSI-Bench, MMSI, and RefSpatial-Bench. Alongside the leaderboard claim, the release adds contact-point prediction across the model family and native 3D grounding for the 2B and 9B variants, features aimed at manipulation and grasp planning rather than pure visual question answering.

On the robotics side, DAMO has wrapped a policy layer called RynnBrain-VLA around the base models with what the paper describes as a unified cross-embodiment action space and embodiment-specific masking, and deployed it on Unitree G1, Astribot-S1, and Tianji-Wuji. That mix matters because cross-embodiment is where most current robot foundation models are still brittle, and getting one policy to drive a humanoid alongside other form factors is the harder test.

The honest caveat is that these are the paper's own numbers on the paper's own benchmark selection, and the abstract does not quantify how well the 122B tier holds up on real-robot success rates versus the smaller siblings, which are the ones that actually got the 3D grounding upgrade. What the reporting does not give you is inference cost, training data provenance, or a like-for-like comparison against the closed embodied stacks that mostly do not publish weights.

Still, the direction is what to watch. An open, large-scale spatial reasoning backbone released with three real robot integrations pulls the embodied race closer to the same open-versus-closed dynamic that reshaped chat models, and tightens the pressure on labs still keeping this class of model behind an API.