Categories AI and Tools

Xiaomi Made MiMo-V2.6 Open Source Under the MIT License

Xiaomi Just Dropped a Top-Tier Open Model

On September 21–22, 2026, Xiaomi released MiMo-V2.6, a family of open-weight AI models under the MIT license. The headline numbers: MiMo-V2.6-Pro is a 1.02-trillion-parameter mixture-of-experts model with 42 billion active parameters per token, a 1-million-token context window, and native multimodal input across text, image, video, and audio. It scores 46 on the Artificial Analysis Intelligence Index — the highest mark of any open-weight model to date.

Xiaomi didn’t just release weights. The company published more than 7,000 reinforcement-learning environments plus its end-to-end RL training framework, and it live-streamed the reinforcement-learning training runs in public. That’s an unusual degree of openness even by open-source standards, and it’s the part researchers are most excited about.

The MiMo-V2.6 Family: Four Checkpoints

The release ships in four checkpoints, which has caused some confusion online. Here’s the lineup:

Model Total params Active params Context Best for
MiMo-V2.6-Pro 1.02T 42B 1M tokens Demanding long-horizon agentic work
MiMo-V2.6-Flash 309B 15B 1M tokens Best balance of intelligence, speed, cost
MiMo-V2.6-Pro UltraSpeed ~1T 42B 1M tokens Pro quality at faster output
MiMo-V2.6-Distill-Qwen-9B 9B 9B — Single-GPU use, RL research starting point

A few things worth understanding. “Natively omnimodal” means the same model takes text, images, video, and audio as input, rather than bolting a vision adapter onto a text model. The 1-million-token context window is roughly several thousand pages of text at once. And the sparse mixture-of-experts architecture is why a “1 trillion parameter” model stays cheap to run: you pay for the 42 billion active parameters per token, not the total.

The 9B distill is a supervised fine-tuned checkpoint based on Qwen3.5-9B, explicitly positioned as a starting point for reinforcement-learning research rather than a production model.

Benchmarks: What Xiaomi Claims

Xiaomi’s model card claims Pro performs on par with Claude Opus 5 and GPT-5.6-class models on most agent benchmarks, with near-parity scores on DeepSWE v1.1, AutomationBench, and OSWorld-Verified. Reported figures include 78.6% on SWE-Bench Verified (thinking mode), 71.9 on DeepSWE v1.1, 89.9 on Terminal-Bench 2.1, and 82.0 on OSWorld-Verified.

The standard caveat applies: these are vendor-reported numbers on the model card, and independent verification takes time. The Artificial Analysis Intelligence Index score of 46, however, comes from an independent benchmark aggregator, which gives the “best open-weight” claim more weight.

The MIT License Is the Real Story

Plenty of “open” models come with restrictive licenses that limit commercial use or require sharing modifications. MiMo-V2.6 ships under the MIT license — one of the most permissive licenses in existence. You can use it commercially, modify it, and build products on it with minimal obligations.

For context on why this matters: several of Xiaomi’s own earlier releases in the V2 line were proprietary. The V2.6 generation going fully MIT-licensed, including Pro, is a meaningful escalation in the open-model race. Combined with the published RL environments and training code, a small team can now study and build on a training pipeline that cost Xiaomi a reported $2.62 million in RL compute for Pro and about $850,000 for Flash — excluding pretraining.

The Anthropic Dispute

Shortly after launch, a dispute erupted: Anthropic accused Xiaomi of misusing Claude on a large scale for data collection during MiMo’s development. As of this writing, this is an allegation, not a ruling. It’s worth knowing about because it touches the central tension in open-model development: training data provenance. If you build a product on MiMo-V2.6, keep an eye on how this resolves, since licensing disputes can create downstream uncertainty. For now, the MIT license stands as published.

Who Should Care About MiMo-V2.6

Developers and Indie Hackers

The MIT license plus Flash’s efficiency profile is the combination to look at. If you’re building an AI feature and the API bills from frontier labs hurt, an open-weight model you can self-host or run through a cheap inference provider changes the unit economics. Flash at 309B total / 15B active is designed exactly for high-volume request processing.

Researchers

The 7,000+ published RL environments covering software engineering, vulnerability reproduction, knowledge-intensive tasks, and web development, plus the complete training framework, make this one of the most reproducible frontier-training efforts ever published. If you do RL research, the environments alone are worth the download.

Businesses Evaluating AI Vendors

The practical takeaway: the gap between open-weight and proprietary models keeps narrowing on agentic benchmarks. If your AI strategy assumes you must pay frontier-lab prices for capable agents, MiMo-V2.6 is evidence worth pricing into your planning. Pilot it against your current provider on your actual tasks before making any decisions.

How to Try It

The weights are available on Hugging Face under the MIT license. For most people, the fastest path is an inference provider or Xiaomi’s own AI Studio rather than self-hosting a trillion-parameter model. The 9B distill is the only checkpoint realistic for single-GPU experimentation.

One caution: “open weights” doesn’t mean “cheap to run yourself.” Pro needs serious GPU infrastructure. Start with hosted inference, benchmark against your current model on real tasks, and only consider self-hosting if the math works at your volume.

Sources: Open Source For You’s coverage of the release, Wikipedia’s Xiaomi MiMo overview. Benchmark figures are Xiaomi-reported unless noted; verify current scores on Artificial Analysis.

Frequently Asked Questions

What is Xiaomi MiMo-V2.6?

MiMo-V2.6 is a family of open-weight, natively multimodal AI models released by Xiaomi on September 21–22, 2026, under the MIT license. It includes Pro (1.02T parameters, 42B active), Flash (309B, 15B active), a Pro UltraSpeed variant, and a 9B distill, all with 1M-token context windows (except the distill).

Is MiMo-V2.6 really free to use commercially?

The weights are released under the MIT license, which permits commercial use, modification, and redistribution with minimal obligations. That’s among the most permissive licenses available for a model of this class.

How good is MiMo-V2.6 compared to GPT or Claude?

Xiaomi claims near-parity with Claude Opus 5 and GPT-5.6-class models on agentic benchmarks. The independent Artificial Analysis Intelligence Index score of 46 is the highest of any open-weight model. Treat vendor benchmark claims as directional until independently verified.

What does “natively omnimodal” mean?

The model accepts text, image, video, and audio inputs in a single architecture, rather than attaching separate vision or audio adapters to a text model.

What’s the controversy with Anthropic?

Anthropic has accused Xiaomi of misusing Claude at scale for data collection during MiMo’s development. This is an allegation, not a ruling, and the MIT license stands as published. Worth monitoring if you build on the model.

Can I run MiMo-V2.6 on my own hardware?

Realistically, only the 9B distill fits single-GPU setups. Pro and Flash need data-center-class GPU infrastructure; most users should start with hosted inference providers.

Leave a Reply

Your email address will not be published. Required fields are marked *