Xiaomi Details MiMo-V2.6 RL: 1,568-Sample Steps, 1M Context
TL;DR
- The MiMo-V2.6 training report says RL consumes 1,568 samples per step at 2.7 to 3.7 billion tokens per step.
- Context lengths during RL reach up to 1M, with code, general, visual, and cyber domains run through a mix of agent harnesses.
- Xiaomi is open-sourcing the training dynamics, RL environments, and RL framework alongside the weights, with the MoE router frozen during training.
Xiaomi's MiMo team has published the training report for its MiMo-V2.6 series on Hugging Face, and the throughput figures do the work. The system consumes 1,568 samples per step at 2.7-3.7B tokens per step, with context lengths reaching up to 1M.
The paper introduces MiMo-V2.6 as 'an omni-modal family that pushes the frontier of model intelligence by scaling RL compute,' built on a pretrained hybrid-SWA base. Scaling happens along three dimensions: batch size and throughput under an asynchronous training architecture; environment diversity across code, general, visual, and cyber domains with a mixture of agent harnesses; and grader compute delivered via groupwise agentic grading, which the authors argue yields more accurate reward signals on long-horizon tasks and pushes the model toward shorter, token-efficient solutions.
The engineering choices sit on the stability side. The MoE router is frozen to maintain stability during RL. Multi-layer defense mechanisms are deployed against reward hacking. The infrastructure section describes a unified trajectory representation, high-concurrency multi-framework rollout, decoupled control and data planes, and training-inference consistency.
Xiaomi is open-sourcing the training dynamics, RL environments, and the RL framework alongside the weights. The abstract names no benchmark scores, no parameter count, and no hardware configuration. It arrives in a crowded week of Chinese open-source infrastructure releases on our tracker, including Tsinghua's TokenRouter serving system.
Originally reported by huggingface.co
Read the original article →Original headline: Xiaomi Paper Details MiMo-V2.6 RL Stack: 1M Context, 1,568-Sample Steps on Trillion-Param MoE