Xiaomi’s MiMo team today officially released and open-sourced the MiMo-V2.6 series, framing the work as a step along a recursive self-improvement path built on verifiable complex tasks and large-scale reinforcement learning.

The series includes two native multimodal models: MiMo-V2.6-Pro and the lighter MiMo-V2.6-Flash. Xiaomi says Pro scores 46 on the Artificial Analysis Intelligence Index, ahead of named open rivals in its comparison, while still trailing the strongest closed models it cites—Claude Fable 5.1 and GPT-6 Astra. On agent benchmarks, Xiaomi reports Pro performance on par with Claude Opus 5 and GPT-5.6 Sol for most of the tasks it lists, with Flash outperforming the prior MiMo-V2.5-Pro across the board.

The training disclosure is unusually concrete. Xiaomi says Flash and Pro each completed 30 RL steps in less than six days, processing about 750,000 trajectories combined, at roughly $850,000 and $2.62 million in training cost respectively (about $3.47 million total). It reports relative training-task pass-rate gains of about 25% (Flash) and 12% (Pro), and out-of-sample DeepSWE v1.1 gains of roughly 17 and 14 points. The company says it live-shared the experimental run and is open-sourcing weights, a technical report, RL task environments, and an end-to-end RL training framework.

API pricing remains unchanged from the V2.5 series. Xiaomi also launched a MiMo Desktop client with Pro UltraSpeed mode (up to 20× inference speed) and published weights on Hugging Face under its MiMo-V2.6 collection, including a Distill-Qwen-9B research checkpoint.

DigiEditorial verified these claims against Xiaomi’s primary MiMo announcement and corroborating coverage from 36Kr and Temperature2; independent reproduction of the vendor’s agent and DeepSWE numbers was not available at publish time.