Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Xiaomi trained a model live and spent 3.5 million dollars in six days

MiMo-V2.6-Pro has 1.02 trillion parameters, 42 billion of them active. In an open live training run it scored 46 points on the Artificial Analysis index and took first place among open models.

AI & modelsNewsGrace OkonkwoPublished: 23 September 20265 min readSources 2
Xiaomi trained a model live and spent 3.5 million dollars in six days

Xiaomi turned model training into a public spectacle. For six days it streamed a reinforcement learning run, with a spending counter ticking up in real time. As the Chinese outlet 量子位 reported, both models, MiMo-V2.6-Pro and MiMo-V2.6-Flash, finished 30 training steps each. The final bill came to 3.5 million dollars, or about 23.44 million yuan.

The cost breakdown shows how expensive large-scale reinforcement learning has become. The Pro model alone needed 5 days and 3 hours, at about 2.6 million dollars. Flash took 3 days and 10 hours for roughly 900,000 dollars. Each step held 1,568 prompts, and for every task the model tried 16 different paths. That came to more than 25,000 rollouts per step. At 2.7 to 3.7 billion training tokens per step, the Pro model alone processed about 81 to 111 billion tokens over the whole run. The outlet calculated that this is the volume of more than 80,000 context windows of one million tokens each.

The result and the comparison with the frontier

MiMo-V2.6-Pro has 1.02 trillion parameters, 42 billion of them active, and scored 46 points on the Artificial Analysis intelligence index, taking first place among open models. On most agent benchmarks the Pro result comes close to Claude Opus 5 and GPT-5.6 Sol. The head of the MiMo project, 罗福莉, who previously worked on DeepSeek-R1, called it one of the largest single reinforcement learning runs any team has carried out on an open model. She added that the engineering challenge surpassed the one she remembers from DeepSeek-R1.

Openness as a strategy, not a gesture

Xiaomi is also publishing architecture details. MiMo-V3 builds on HySparse2, which is optimised for long, multi-step agent tasks: two-level KV sharing, sparse token selection and an earlier exit from the prefill phase. Prefill compute drops to one fifth of what the HySparse solution with a hybrid attention window needs, and KV cache falls from 12 GB to 2.7 GB at a one million token context. According to Xiaomi, the ability to search long texts does not fall but rises.

That is the same direction visible in the DeepSeek family: Chinese models now compete not on the peak of parameters but on the cost of holding a long context. For the end user one more thing matters here. The weights are opened, so the results do not have to be taken on faith.

Comments 0

Sources

2
  1. 016 天烧光 2000 多万,拿下开源第一!小米史无前例「炼丹直播」收官 (量子位)ZH
  2. 02小米公开 MiMo-V3 核心架构 HySparse2 (IT之家, tag AI)ZH

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.