Xiaomi opens two V2.6 checkpoints under MIT

Xiaomi released and open-sourced its MiMo-V2.6 series in a post titled MiMo-V2.6: Scaling Up Reinforcement Learning for Self-Improvement. The post carries an update time of 22 September 2026. The MiMo-V2.6-Pro-RL repository on Hugging Face went up at 15:39 UTC on 21 September, or 23:39 SGT.

The series has two natively multimodal models, Pro and Flash. Both model cards list the MIT licence in their metadata. The MiMo-V2.6 collection also holds MiMo-V2.6-Distill-Qwen-9B and two MOPD checkpoints.

Xiaomi says MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index. Xiaomi says that result puts it ahead of Kimi K3 and Qwen3.8 Max. The same post says a gap remains to Claude Fable 5.1 and GPT-6 Astra.

Xiaomi's MiMo team and teams that serve open weights

Xiaomi's MiMo team published the weights, the technical report and the RL resources. The MIT licence permits commercial use and fine-tuning, provided copies keep the copyright and permission notice.

The release matters most to teams that self-host agent models. It also matters to researchers who want to reproduce large-scale RL on agent tasks.

API customers see the same prices as the V2.5 series. The API model names are mimo-v2.6-pro, mimo-v2.6-flash and mimo-v2.6-pro-ultraspeed.

A 1.02T MoE and more than 7,000 RL environments

The model card describes Pro as a sparse mixture-of-experts model. It has 1.02T total parameters and 42B active parameters. It takes text, image, video and audio input, with a 1M-token context.

MiMo-V2.6 architecture diagram with visual and audio encoders feeding a hybrid sliding-window attention backbone and multi-token prediction blocks Caption: MiMo-V2.6 architecture, Figure 1 in the model card · Source: Xiaomi MiMo, MiMo-V2.6-Pro-RL model card, Sep 2026 · link

Xiaomi also published training resources:

  • More than 7,000 RL task environments across software engineering, vulnerability reproduction, knowledge work and web development.
  • An end-to-end RL framework built on verl, uni-agent and mini-swe-agent.
  • Minimal harnesses that separate system prompts, tools and context management.

The post gives the cost of the live RL phase. Flash and Pro each ran 30 steps in under 6 days. Together they produced about 750,000 trajectories. Xiaomi puts the training cost at around $850,000 for Flash and $2.62 million for Pro.

The model card table reports these Pro scores against Claude Opus 5:

  • DeepSWE v1.1: 71.9 for Pro and 74.0 for Opus 5.
  • Terminal Bench 4.0: 34.9 for Pro and 49.0 for Opus 5.
  • OSWorld-Verified: 82.0 for Pro and 83.4 for Opus 5.

Vendor scores and a two-node serving recipe

Every benchmark number above comes from Xiaomi. No independent lab has published a reproduction yet.

The post and the model card give different DeepSWE v1.1 figures for Pro. The post says RL raised Pro from 58.4 to 72.6. The model card table lists 71.9. Xiaomi gives no reason for the difference.

The Pro repository holds about 573 GB of files. Its config file lists FP8 quantisation. The SGLang recipe in the model card uses two nodes, with tensor parallelism of 16. The vLLM recipe uses tensor parallelism of 8. Xiaomi gives no minimum GPU type.

The Pro repository has no standalone LICENSE file. The MIT grant appears only in the model card metadata.

Test Pro on your own agent tasks first

  1. Run Flash or Pro on your own agent tasks before you trust the vendor table. Check Terminal Bench style tasks, where the gap to Opus 5 is widest.
  2. Budget for multi-node serving if you plan to self-host Pro. Start from the SGLang recipe in the model card.
  3. Keep a copy of the model card metadata with your licence records. The repository has no separate LICENSE file.
  4. Research teams can start RL experiments from MiMo-V2.6-Distill-Qwen-9B and the released environments.