Most AI labs announce a new model with benchmarks and a launch post. Xiaomi announced MiMo-V2.6-Pro-RL this week with something rarer: an itemized bill.
The model now ranks first among open-weights models on Artificial Analysis’ Intelligence Index with a score of 46. That puts it level with Grok 4.7 and performing similarly to GPT-6 Sol at lower cost per task, ahead of GLM-5.3 (45) and Kimi K3 (44). Its own predecessor, MiMo-V2.5-Pro, scored 26. A one-generation jump from 26 to 46 is not a normal iteration. It’s the kind of leap that usually comes with a new architecture or a budget nobody talks about.
Here, the budget got talked about. Xiaomi says training the Pro model cost $2.6 million. The smaller Flash variant cost $0.9 million. Both figures ship alongside the full reinforcement learning recipe: more than 7,000 task environments and the code needed to train on them, released for anyone to inspect.
That combination of a competitive model, an auditable recipe, and the actual bill is worth an hour of your attention even if you never download the weights.
What you actually get
The specs are serious for a free download. MiMo-V2.6-Pro-RL is a mixture-of-experts model with 1.02 trillion total parameters, 42 billion of them active per token. It accepts up to 1 million tokens of input across text, images, video, and audio. The license is MIT, which means commercial use with no strings.
There’s also a smaller sibling. MiMo-V2.6-Flash, at 309 billion parameters with 15 billion active, currently ranks first among open-weights models on the Vals Index. If you’re deploying locally or trimming inference costs, the Flash model is likely the one to test first.
The training mix tells you where the capability lives: 68 percent coding, 13 percent visual artifact design, 12 percent tool use, 4 percent cybersecurity. This is a model built for work rather than for trivia benchmarks.
Why the receipt matters more than the leaderboard
Open weights are common now. Open budgets are not.
When a lab publishes a model without training details, you can use it but you can’t learn from it. You can’t compare your own fine-tuning costs, you can’t sanity-check your agent eval setup against theirs, and you can’t tell whether their results came from a clever method or a clever method plus a thousand GPUs nobody mentions.
Xiaomi’s release changes that for one data point, and one data point is still valuable. $2.6 million for frontier-adjacent capability is a concrete number you can put in a planning doc. If you’re a startup deciding whether to fine-tune an open model or rent a proprietary API, or a research group budgeting an RL run, you now have a real, recent, disclosed cost to argue with. Per The Batch’s coverage, the open environments matter as much as the money: 7,000-plus reinforcement learning task environments are the kind of infrastructure most teams could never afford to build alone, and they’re now downloadable.
The detail that shouldn’t get buried
Buried under the headline numbers is a training insight that anyone building coding agents should steal.
Xiaomi’s key reinforcement learning innovation was rewarding code quality rather than raw test passage. A separate AI grader examined solutions that already passed the tests and ranked them by whether a human reviewer would accept them. Passing but ugly or bloated code scored lower. Passing code that silently ignored edge cases scored lower too.
This matters because test-passing alone teaches models bad engineering habits. A model rewarded only for green tests learns to add unneeded code, paper over failures, and let errors slide as long as the assertions hold. Xiaomi found that grading the way a senior engineer grades fixed this. The model stopped shipping the extra junk.
If you’re training or even just evaluating coding agents, this is directly applicable. Add a review-quality score next to your pass rate. It’s cheap, and per Xiaomi’s own results, it’s the difference between a model that passes tests and one whose output you’d actually merge.
The allegation hanging over the release
None of this lands in a vacuum. Anthropic has previously alleged that Xiaomi routed MiMo users’ conversations to Claude and harvested more than 400,000 exchanges as training data. Xiaomi has not publicly answered that accusation, and as of this release it remains unresolved.
I’m not in a position to judge who’s right. But if you’re choosing a foundation model for anything sensitive, an unanswered data-provenance dispute is a legitimate line item in your risk assessment. The MIT license and the published recipe make the technical side transparent. The training-data question is the part still sitting in the dark.
What to do with this now
Three concrete moves, depending on who you are:
- If you run agents in production: benchmark MiMo-V2.6-Flash against your current model this week. The Vals Index result and the 15 billion active parameters make it a strong candidate for cost-sensitive workloads, and the MIT license means no negotiation required.
- If you train or fine-tune models: download the RL environments and study the reward design before anything else. The code-quality grader is the single most transferable idea in the whole release.
- If you write AI strategy or budgets: quote the $2.6 million figure with its date. Frontier-adjacent capability at that price is a data point that shifts build-versus-rent math, and it’s one of the few numbers in this space you can actually cite.
The open-weights leaderboard changes weekly. Disclosed training costs with published recipes change almost never. That’s the part of this release worth remembering after the benchmark scores go stale.
Xiaomi’s MiMo-V2.6 release was covered in The Batch (DeepLearning.AI), issue 373.


