huggingface.co web signal

Xiaomi Details MiMo-V2.6 RL: 1,568-Sample Steps, 1M Context

TL;DR

  • The MiMo-V2.6 training report says RL consumes 1,568 samples per step at 2.7 to 3.7 billion tokens per step.
  • Context lengths during RL reach up to 1M, with code, general, visual, and cyber domains run through a mix of agent harnesses.
  • Xiaomi is open-sourcing the training dynamics, RL environments, and RL framework alongside the weights, with the MoE router frozen during training.

Xiaomi's MiMo team has published the training report for its MiMo-V2.6 series on Hugging Face, and the throughput figures do the work. The system consumes 1,568 samples per step at 2.7-3.7B tokens per step, with context lengths reaching up to 1M.

The paper introduces MiMo-V2.6 as 'an omni-modal family that pushes the frontier of model intelligence by scaling RL compute,' built on a pretrained hybrid-SWA base. Scaling happens along three dimensions: batch size and throughput under an asynchronous training architecture; environment diversity across code, general, visual, and cyber domains with a mixture of agent harnesses; and grader compute delivered via groupwise agentic grading, which the authors argue yields more accurate reward signals on long-horizon tasks and pushes the model toward shorter, token-efficient solutions.

The engineering choices sit on the stability side. The MoE router is frozen to maintain stability during RL. Multi-layer defense mechanisms are deployed against reward hacking. The infrastructure section describes a unified trajectory representation, high-concurrency multi-framework rollout, decoupled control and data planes, and training-inference consistency.

Xiaomi is open-sourcing the training dynamics, RL environments, and the RL framework alongside the weights. The abstract names no benchmark scores, no parameter count, and no hardware configuration. It arrives in a crowded week of Chinese open-source infrastructure releases on our tracker, including Tsinghua's TokenRouter serving system.