MiMo V2.6: The Open-Source Model That Beat Grok

9人浏览 / 0人评论

On September 21, Xiaomi's MiMo team released and fully open-sourced MiMo V2.6 — three models at once: Pro, Flash, and Ultraspeed. The same week, Grok 4.7 landed too, promoted by Elon Musk himself. Then the AI community did something unusual: it spent most of the week talking about the cheaper, open model.

Both scored 46 on the Artificial Analysis Intelligence Index (v4.3). MiMo V2.6 Pro cost $0.13 per task. Grok 4.7 cost $3.74. Nearly a thirty-fold gap, at the same measured intelligence — open weights, full multimodality, one-thirtieth the price.

X's verdict was five words: better, cheaper, and open.

Three Models, Three Jobs

MiMo V2.6 is a family, not a single release.

Pro is the flagship: full-modal, built for complex projects, long-horizon tasks, and high-value work. It scores 46 on the AA index — first among all open-source models — and on most agent benchmarks it trades blows with closed flagships like Opus 5 and GPT-5.6 Sol. In Design Arena, open-source rankings put it above Claude Opus 5 in chat, above Opus 5, Fable 5, and GPT-5.6 Sol in web-app front-end tasks. Every name it passes is a several-billion-dollar closed model.

Flash is the efficiency workhorse: full-modal, high-intelligence, low cost, built for high-frequency office workloads. This generation already surpasses the previous flagship, V2.5 Pro.

Ultraspeed keeps Pro-level performance — no quantization, no intelligence sacrificed for speed — and delivers up to 10x inference speed, a steady 500 TPS, peaking at 1,000. Three times the price for ten times the speed: for latency-sensitive production systems, the math writes itself.

MiMo V2.6 Pro Flash and Ultraspeed as three chips with different roles

Not a Benchmark Runner: A Working Tool

The rankings are the entrance ticket. The demonstrations are the pitch.

Games. MiMo V2.6 generated a playable 3D open world — a Middle Eastern city skyline with minarets and pyramids, dense sandy rooftops, stable frame rates, no pop-in, and a consistent art direction across the whole scene.

3D modeling. Given text or a reference image, it generates 3D objects and scenes in Blender — a wireframe pickup truck where body, wheel spokes, and chassis girders hold their spatial relationships perfectly as the camera rotates.

Office documents. Shown a fund-raising deck for a critical-minerals strategy, it produced slides that read like a real analyst's work: restrained typography, a strong cover image, and — more tellingly — the right narrative order. It thought about the audience before the layout.

Materials research. Xiaomi's advanced materials team used MiMo V2.6 Pro to design novel metal-organic frameworks that capture PFAS ("forever chemicals") from water. An end-to-end research loop: literature and patent review, hypothesis generation and novelty checks, automated computational environments, "dry-lab" experiments, and candidate screening. Two of its designs showed adsorption performance millions of times higher than reference materials, compressed a month of R&D into two to three days, and earned a Peking University researcher's assessment that the model's performance matched a trained doctoral researcher.

It has one underrated advantage too: extremely fast single-turn delivery. In complex multi-step tasks, MiMo completes more iterations in the same wall-clock time — which changes what "human-in-the-loop" feels like.

MiMo V2.6 use cases — 3D game world, Blender modeling, office deck, materials research

The Pretraining Stays; Everything Changes Afterward

Here is the part that made the community stop scrolling: the pretrained base did not change. No new parameters. The intelligence jump came entirely from post-training.

Xiaomi frames the release as a step toward RSI — recursive self-improvement: scale RL compute on verifiable, complex tasks, and let the model push its own intelligence boundary through iterative exploration and feedback.

MiMo's head of foundation models, Fuli Luo, published The Hard Road to Scaling Up RL and — for the first time anyone can remember — livestreamed the RL run itself. In under six days, Flash and Pro each completed 30 steps, roughly 750,000 trajectories, at training costs of about $0.85M and $2.62M. All of it, curves included, was live at mimo.xiaomi.com/rl — prompting AI researcher Nathan Lambert to call it one of the coolest public large-scale RL resources to date.

The numbers backed the spectacle: pass rates on training tasks improved 25% and 12% respectively, and on the out-of-sample long-horizon software engineering benchmark DeepSWE v1.1, Flash jumped from 48.8 to 65.68 and Pro from 58.4 to 72.57. Thirty steps. No task-specific retraining.

The method behind it is called MixRL — mixed-task reinforcement learning. Luo's summary: You Only RL Once. One massive mixed-task RL run, and capabilities emerge across every axis at once.

The Three Axes of Self-Improvement

Axis one: Rollout compute — practice enough.

Each training step uses 1,568 prompts with 16 rollouts each, 2.7–3.7 billion tokens per step, trained at 1M context, fully asynchronous across 4,000 GPUs and 25,000 concurrent trajectories. Why 16? Under GRPO, one attempt yields a binary pass/fail; sixteen rollouts let the grader compare within a group and distinguish lucky guesses from real understanding. Sixteen is the sweet spot between cost and signal. And "fully asynchronous" is a quiet engineering insight: sync RL is a class-wide exam where fast GPUs idle waiting for slow ones. MiMo turned generation, grading, and learning into a continuous pipeline — whichever trajectory finishes first gets scored and learned from first.

Axis two: Environments — see broadly.

Instead of training on one task type, V2.6 mixed programming (68%), general agents (12%), visual design (13%), cybersecurity (4%), and instruction following (3%) across 25 data sources and 21 harness frameworks in a single run. Each domain carries its own quality control: programming gets triple checks (original-requirement-only grading, eight repeat runs for consistency, an independent audit agent); cybersecurity trains on real OSS-Fuzz vulnerabilities with pure rule-based grading; visual design uses pixel-level similarity for high-fidelity replication plus LLM judgment for open-ended design.

Axis three: Grader compute — feedback precisely.

This is the axis that is easiest to skip and hardest to overrate. V2.6 turned grading itself into an agent. Simple tasks use offline rubrics; complex long-horizon tasks launch an online Agentic Grader that ranks all 16 rollouts in a group and assigns credit across five dimensions — whether the approach fits, whether the edit is precise, whether the change is minimal, whether it introduces side effects, and whether the engineering quality holds. The grader can identify a solution that passes tests by cheating, reallocate reward to genuinely better solutions, and push the model toward shorter paths with fewer tokens.

The comparison chart in the technical report tells the story: without the grader, dialogue turns and token use balloon and pass rates plateau; with it, turn counts stay stable and pass rates keep climbing past step 52.

Practice volume, environment breadth, feedback accuracy — three axes forming a closed loop: policy produces trajectories, the grader extracts fine-grained quality signals, signals steer the policy, the better policy produces better trajectories. That is RSI in engineering form.

MiMo V2.6 reinforcement learning loop — rollouts, environments, agentic grader

Not a Sudden Jump: Compound Returns

The stability work behind a mixed run this large is invisible and enormous: a four-layer Sample Mixer scheduler balancing five domains where task latency differs 56x and data volume differs 159x; frozen MoE routing during RL (each token touches 8 of 384 experts, and drifting routes would destabilize training); and three lines of defense against reward hacking — environment sanitization (logs stripped, network cut, git history truncated), adversarial screening (a dedicated hack agent attacks the environment until it finds nothing), and continuous offline audits for new cheating patterns.

And the pricing puts it in context. Cache prices are ~99% lower, input ~89% lower, and output ~95% lower than Opus 5 and GPT-5.6 Sol at comparable tiers. At equal intelligence, MiMo V2.6 Pro costs between 1/20 and 1/60 of overseas flagships — the first time frontier-grade intelligence has entered the $0.10-per-task range. Only a handful of closed, far more expensive models now sit outside its kill range.

MiMo V2.6 cost gap versus closed flagships, cheap house versus skyscraper

None of this is overnight luck. The MiMo line has been compounding: MoE sparse activation, sliding-window attention, multi-token prediction, and inference-system co-design across previous generations. V2.6 is what that compounding curve looks like when the training side finally gets its own revolution.

The open-source world has a new frontier line. This time, it was drawn by the model that posted its training curves on a public dashboard while running them.

References:

[1] Xiaomi MiMo V2.6 release and open-source announcement, 2026-09-21; live RL training dashboard: mimo.xiaomi.com/rl.

[2] Artificial Analysis Intelligence Index v4.3 — model scores and per-task cost benchmarks, artificialanalysis.ai, 2026-09.

[3] Luo F. "The Hard Road to Scaling Up RL" (MiMo V2.6 post-training methodology), X/Twitter, 2026-09-17.

[4] MiMo V2.6 technical report — MixRL training recipe, grader design, and DeepSWE v1.1 out-of-sample results, 2026.

[5] All illustrations are AI-generated.

全部评论