Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Open models for 3.5 million dollars: Xiaomi and DeepSeek rewrite the cost bill

MiMo-V2.6-Pro ran 30 reinforcement learning steps, cost about 2.6 million dollars and beat closed models on some agentic tasks.

TechnologyOpinionGrace OkonkwoPublished: 25 September 20267 min readSources 3
Open models for 3.5 million dollars: Xiaomi and DeepSeek rewrite the cost bill

For years the rule was simple. The best model is closed, and it costs whatever the operator decides to charge. The Chinese outlet qbitai (量子位) described an experiment by Xiaomi that undermines at least the second half of that rule.

During a live broadcast, the MiMo team led by Luo Fuli showed the training bill. Luo Fuli had previously worked on DeepSeek-R1. MiMo-V2.6-Pro completed 30 reinforcement learning steps in 5 days and 3 hours, using about 2.6 million dollars. The lighter MiMo-V2.6-Flash did the same in 3 days and 10 hours for about 900,000 dollars. The total came to around 3.5 million dollars.

The scale of a single step explains why those numbers are not extravagant. Each step covers 1,568 queries, and the model tries 16 different paths for each one. A single round produces more than 25,000 attempts. At 2.7 to 3.7 billion training tokens per step, the Pro model processed about 81 to 111 billion tokens in total. That is the equivalent of more than 80,000 context windows of one million tokens each.

Results that can be verified

Xiaomi points to the effects in two places. The average pass rate rose after 30 steps by 25 percent for Flash and 12 percent for Pro. The test outside the training set matters more. On DeepSWE v1.1, which checks how an agent works in real code repositories, Pro improved from 58.4 to 72.57 points and Flash from 48.7 to 65.68. The model did not just memorize tasks. It carried the skills over to entirely new engineering problems.

In the Artificial Analysis Intelligence Index ranking, MiMo-V2.6-Pro scored 46 points and took first place among open models. On AutomationBench it scored 53.1 points, ahead of Claude Opus 5 (50.3) and GPT-5.6 Sol (45.8). Prices stayed unchanged from the previous generation: Pro about 3 yuan per million input tokens and 6 for output, Flash 1 and 2 respectively. Artificial Analysis estimates the average cost of completing one task from the index at 0.13 dollars for Pro. Grok 4.7 scored an identical 46 points the same day. In its xHigh variant it is to cost about 3.74 dollars, close to 29 times more.

Not only code

The scientific use is more interesting. Xiaomi set the model to designing MOF materials meant to capture PFAS, persistent pollutants that do not break down. The model reviewed the literature and patents, checked the novelty of the idea on its own, then built a simulation environment and calculated the binding strength. It proposed two structures, A50 and B50, based on zirconium UiO-67 frameworks. Their PFAS adsorption is to be a million to ten million times better than the reference material. Professor Dou Jinhu of Peking University, quoted in the report, called the result work close to that of an experienced doctoral student. The research cycle at the company shrank from a month to two or three days.

Openness is no longer a marketing slogan, but a way to bring the cost of a single task down from several dollars to a dozen or so cents.

The same mechanism is at work behind DeepSeek-V4.1-Flash, which with 552 billion parameters and an MIT licence lowered inference memory requirements. Xiaomi published the weights and a technical report, but also more than 7,000 task environments for reinforcement learning, the full training framework and a small model distilled from Qwen3.5-9B. That is a qualitative change in the debate about openness. It is no longer about releasing a file with weights, but about handing over the entire production process.

Comments 0

Sources

3
  1. 01量子位 (qbitai): 小米 MiMo-V2.6 强化学习直播 — 350万美元的训练ZH
  2. 02DeepSeek-V4.1-Flash ReleaseEN
  3. 03DeepSeek-V4.1-Flash – model cardEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.