律动BlockBeats|Sep 17, 2026 01:15
[Luofuli Held Back for Half a Year, Xiaomi Live Streams MiMo-V2.6 Reinforcement Learning Directly]
Beating AI News: Xiaomi is currently live streaming the reinforcement learning training of the next-generation MiMo-V2.6. Project lead Luofuli stated that the team has been silent for nearly half a year, focusing on exploring how far reinforcement learning can scale. At present, V2.6 is still in training, and the specific technical solutions will be gradually open-sourced in the coming weeks. This time, they started with a very large scale. Each step uses **1,568 prompts, generating 16 rollouts per prompt**, totaling approximately **2 billion tokens**, running asynchronously throughout. Tasks involving code, general-purpose functions, vision, cybersecurity, and chat agents are also mixed into the same training session, utilizing multiple harnesses simultaneously. External observers can view real-time internal metrics such as task composition for each step, context length, and training duration. UC Berkeley Ph.D. student Yichuan Wang observed from the dashboard that each step involves approximately 1,500 tasks, with code-related tasks accounting for about two-thirds, and more than 40,000 active sandboxes running concurrently. [Original Article Link]
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink