Mozilla report puts China's best open-weight models about 4.4 months behind US frontier
The best Chinese open-weight models now trail US frontier offerings by roughly 4.4 months, according to version 1.1 of Mozilla's State of Open Source AI report, published on Sept. 15 with data current to Sept. 1.

Mozilla fitted METR task-horizon data and came up with an open-closed gap of about 4.4 months, in line with Epoch AI's four-month estimate. Tom's Hardware reported the figure on Sept. 16. METR is a research nonprofit that scores models by the length of task, in human working time, that they complete half the time.
By Mozilla's fitted estimate, closed models handle tasks that take human experts 8 to 12 hours. Open models reach that level about four months later. Mozilla calculates that open capability doubles every 3.9 months, against 5.5 months for closed models.
The report is a recurring assessment, first published on July 14 on the Mozilla blog. It draws on a Mozilla/SlashData survey of roughly 1,400 developers, OpenRouter traffic data and third-party benchmark indices. Mozilla advocates for open models, and TIME reported on July 14 that Raffi Krikorian, Mozilla's chief technology officer, described the report as partly advocacy.
"Open weights" here means downloadable weights, not training data or code. The report counts 16 notable open releases, but none delivers the data recipe required by the Open Source Initiative's definition.
Benchmarks and prices tell different stories
The best open model trailed the closed leader on the Artificial Analysis Intelligence Index by three points at 60% of the price, and sat two points behind Claude Fable 5 at 30%, according to Tom's Hardware's reading of the report.
Mozilla also charted vals.ai's Terminal-Bench 2.1 results, which run every model through the same harness, the software layer that offers a model its tools. On that board, Z.ai's GLM-5.2 scored within a point of Claude Opus 4.7 and about four points behind Opus 4.8, at less than one-fifth the cost per test.
On OpenRouter, a marketplace that routes developer traffic to hundreds of models, Mozilla counted eight of the top ten models by August token volume as open weights, seven of them Chinese-built. Closed providers nevertheless took 96% of model-layer revenue on OpenRouter from May to September 2025, the Linux Foundation reported.
"We see the decision to pay for closed [models] as workload-specific rather than organization-specific," Krikorian told Ars Technica in an email.
Mozilla's own chart caption is blunt about the durability of its headline number: "the gap resets every release cycle."
The four-month figure has caveats
One caveat is that the four-month gap and the 30% token price figure are measured API to API on hosted endpoints and at list price. The report's own hardware chart puts the best open model that fits one server at 52.6 and the best on one GPU at 40. The drop from the top is 10 and 23 points, respectively, a larger gap than the reported four months.
Serving requirements are the other catch. Kimi K3's native MXFP4 checkpoint runs about 1.56TB across 96 shards, and Mozilla's serving configuration lists 64 or more accelerators. vLLM calls for at least eight GB300 GPUs, with multiple nodes for production traffic.
The report describes this as open but not runnable by most who hold it, and Tom's Hardware put the memory need near 1.5TB in July.
One exception is Thinking Machines' Inkling-Small model, under the Apache 2.0 license, whose NVFP4 version fits one B300 at a 180GB floor.
There is also a provenance dispute attached to one of the models. K3 carries an allegation detailed in the Sept. 8 NSA/CISA/FBI joint advisory (AA26-251A). The claim, which Mozilla's report states as "asserted, and unshown," is that Moonshot extracted Claude Fable 5 data to train K3 through distillation, the practice of training one model on another model's outputs.
Mozilla's data stops at Sept. 1. Since then, Artificial Analysis has moved its index to v4.3 with a different evaluation set. The live board has Claude Fable 5.1 at 53 on its highest effort setting, with Kimi K3 at 44, not comparable to the v4.1.1 numbers Mozilla plotted. vals.ai's Terminal-Bench 2.1 board, updated Sept. 11, is now led by GPT-6 Astra at 87.27%, with Fable 5.1 at 85.02%.
For context on how quickly those rankings move: on July 17, Artificial Analysis had K3 at 57 versus Fable 5's 60, while on Sept. 1, Mozilla had it two points back.
The definition fight underneath the numbers
While Mozilla measures a capability gap, the Open Source Initiative argues the industry is measuring the wrong thing. In a Sept. 19 blog post, OSI's president (uncredited in the post) wrote that open weights let users exercise some of the four software freedoms, the freedom to use, study, modify and share without asking permission from the rights holder, but not all of them, and only to a certain degree.
AI models consist of three main components, according to OSI: training data, weights and parameters, and code for data preparation and training. A user's ability to access those components determines whether a model is closed, open-weight or Open Source. Open-weight releases share the weights, which lets users run a model locally, pay a third party to host it, or fine-tune it on their own data. They do not share the code or the training data.
OSI concedes that open weights give more choice than fully closed systems such as ChatGPT, Claude and many autonomous driving systems. Its objection is about verification. Without the code and the training data, OSI argues, you cannot fully inspect a model to explain an output or confirm you can trust it.
The post cites Ai2's Olmo as an example of the stricter category. Researchers used the full training dataset and model checkpoints to inject new information into the training data and then observe how the model memorized or forgot it, OSI wrote. That kind of study, the organization argues, is only possible under an Open Source AI release.
The distinction is not academic. Mozilla's report counts 16 notable open releases and finds none that delivers the data recipe the Open Source Initiative requires.
Small models, narrow jobs
Away from the frontier argument, some open-weight releases are deliberately small and single-purpose. Superwhisper published S1-mini, a 0.6B-parameter text normalizer for speech-to-text output, according to its model card on Hugging Face.
It takes a raw ASR transcript and rewrites it as clean written text: fillers removed, false starts resolved to the value the speaker landed on, punctuation and capitalization applied, and spoken numbers, dates, times, currency and email addresses rendered in written form.
On a held-out set of 7,519 English cases it reaches 94.8% token accuracy, and the quantized build is a 462 MiB file that runs on a laptop CPU. It is fine-tuned from Qwen/Qwen3-0.6B and carries an Apache 2.0 license plus a naming clause. The model card is explicit that this is not a chat model and will not follow general instructions.
A similar pattern shows up in a Sept. 22 deployment write-up by James Cruce on ASTGL. He replaced a deterministic router with Laya, an open-weight typed-decision model with about 421 million parameters built on a ModernBERT-large encoder and a 1,024-token limit, served over a loopback-only API on a Mac Studio with an M3 Ultra and 256 GB of unified memory.
He chose it over TypeSafe's Jev, an early-access hosted service that lists $0.042 per million input tokens with no output-token charge and supports choices with up to 255 options, because he wanted routine decisions to stay on his machine. Cruce is careful about what the result shows: the model beat his deterministic router on a frozen replay, but he did not test whether Laya beats Jev in general.
He also flags a limit that applies to every number in the Mozilla report. The model card warns Laya is overconfident and needs domain-specific calibration. A typed answer, he writes, can be structurally perfect and factually wrong.
Sources
4- 01China's open-weight AI models are now just 4 months behind frontier US offerings, Mozilla report claimsEN
- 02Open Weights Are Good. Open Source Is Better.EN
- 03S1-mini, Superwhisper's first open-weights language modelEN
- 04Local Laya vs Hosted Jev: Why My Agents Make Typed Decisions on My MacEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.