Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Six months of open-source LLMs: Xiaomi MiMo, DeepSeek and Chinese models on Hugging Face

Xiaomi streamed an entire reinforcement learning run live, and MiMo-V2.6-Pro took the top open-source spot with 1.02 trillion parameters and 42 billion active. DeepSeek also released a vision experiment under the MIT license.

AI & modelsAnalysisRachel NwosuPublished: 23 September 20266 min readSources 3
Six months of open-source LLMs: Xiaomi MiMo, DeepSeek and Chinese models on Hugging Face

Competition among Chinese open-source large models has sped up over the past six months. On September 22, Xiaomi wrapped up a six-day "alchemy livestream" that ran reinforcement learning training for a large model on camera. The final bill came to 3.5 million US dollars, about 23.44 million yuan.

The results came out at the same time. MiMo-V2.6-Pro has 1.02 trillion parameters and 42 billion active, and scored 46 on the Artificial Analysis Intelligence Index, first among open-source models. MiMo lead Luo Fuli said this was one of the largest single reinforcement learning runs any open-source model team has carried out. She added that the research innovation and engineering challenges behind it exceeded those of DeepSeek-R1, which she had worked on.

Out-of-sample generalization matters more. On DeepSWE v1.1, which tests coding agents, Pro went from 58.4 to 72.57 points and Flash from 48.7 to 65.68. After only 30 steps, the models kept getting stronger and carried their abilities over to software engineering tasks they had never seen.

Prices did not rise with the capabilities. MiMo-V2.6 keeps the previous generation's API pricing: Pro costs about 3 yuan per million tokens for input and 6 yuan for output, Flash about 1 yuan and 2 yuan. According to Artificial Analysis estimates, MiMo-V2.6-Pro completes a task for an average of 0.13 US dollars. The closed-source Grok 4.7, released in the same period, also scores 46 on the same index, but its high-end version needs an average of 3.74 US dollars per task.

The open-source list is getting longer too. Xiaomi released the model weights, the technical report, more than 7,000 RL task environments, an end-to-end training framework and a freely combinable mini-harness. It also put out a distilled model built on a Qwen3.5-9B base and supervised-fine-tuned on 77.4B tokens of data generated by MiMo. A former Hugging Face researcher focused on speed: the final RL training started less than a week before the model and training framework were handed over.

DeepSeek put DeepSeek-V4-Flash-Vision-Exp, the first multimodal model in the V4 series, on Hugging Face on August 31 under the MIT license. It published model files, the tokenizer, a reference prompt encoding implementation and a minimal PyTorch inference implementation. With both environment and ecosystem pushing, Chinese models are becoming more visible in the open-source community.

Comments 0

Sources

3
  1. 016 天烧光 2000 多万,拿下开源第一!小米史无前例炼丹直播收官ZH
  2. 02DeepSeek-V4-Flash-Vision-Exp 模型已开源ZH
  3. 03DeepSeek-V4.1-Flash 模型卡EN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.