Open weights are not open source: what the AI debate keeps getting wrong
Mozilla's latest State of Open Source AI report claims China's best open-weight models now trail US frontier offerings by about four months, but the Open Source Initiative argues open weights deliver only part of what open source actually means.

Open weights have become the most contested phrase in AI policy. Mozilla published version 1.1 of its State of Open Source AI report on Sept. 15, using data current to Sept. 1. Its central claim: the best open-weight models now sit roughly four months behind frontier US offerings. Tom's Hardware covered the report on Sept. 16. Mozilla fits METR task-horizon data to put the open-closed gap at around 4.4 months, in line with an Epoch AI estimate of four months.
The Open Source Initiative does not dispute that open weights are useful. It disputes the label. In a blog post published Sept. 19, OSI draws a line between open-weight models, which release weights and nothing else, and Open Source AI, which also releases the code used to train the model and either the training data or a detailed account of how that data was built. Only the second category, OSI argues, lets users exercise all four software freedoms: to use, study, modify and share without asking the rights holder.
The four-month gap, and what it rests on
Mozilla's number is narrower than it sounds. The report counts 16 notable open releases, and none delivers the data recipe required by OSI's Open Source AI definition. The four-month figure is measured API to API on hosted endpoints and at list price, per Tom's Hardware. It describes what a developer pays a provider, not what a model costs to run on your own hardware.
On the Artificial Analysis Intelligence Index, Mozilla found the best open model trailed the closed leader by three points at 60% of the price, and sat two points behind Claude Fable 5 at 30% of the price. On vals.ai's Terminal-Bench 2.1 board, which runs every model through the same software harness, Z.ai's GLM-5.2 scored within a point of Claude Opus 4.7 and about four points behind Opus 4.8, at less than one-fifth the cost per test. On OpenRouter, Mozilla counted eight of the top ten models by August token volume as open weights, seven of them Chinese-built. Closed providers still took 96% of model-layer revenue on OpenRouter from May to September 2025, according to the Linux Foundation.
Then there is the hardware caveat, the part of the report that undercuts the headline. Mozilla's own chart puts the best open model that fits one server at 52.6 and the best that fits one GPU at 40, drops of 10 and 23 points from the top. That is a wider gap than four months. Kimi K3's native MXFP4 checkpoint runs about 1.56TB across 96 shards, and Mozilla's serving configuration lists 64 or more accelerators. vLLM calls for at least eight GB300 GPUs, with multiple nodes for production traffic. The report itself describes this as open but not runnable by most who hold it. One exception noted: Thinking Machines' Inkling-Small, under Apache 2.0, whose NVFP4 version fits one B300 at a 180GB floor.
"We see the decision to pay for closed [models] as workload-specific rather than organization-specific," Mozilla chief technology officer Raffi Krikorian told Ars Technica in an email.
Mozilla is an advocacy organization, and it says so. Tom's Hardware notes that TIME reported on July 14 that Krikorian described the report as partly advocacy. The report is built on a Mozilla and SlashData survey of roughly 1,400 developers, plus OpenRouter traffic data and third-party benchmark indices. Mozilla's chart caption is candid about the moving target underneath all of it: "the gap resets every release cycle."
Why the definition fight matters more than the benchmark
OSI's argument is not that open weights are bad. It is that the term has been stretched to cover models that let you fine-tune and self-host while leaving the training data and training code closed. That distinction has teeth in policy debates, where open weights are often treated as a proxy for openness, competition and auditability.
"Open-weight models limit your ability to modify the model, and you cannot fully study it," the OSI post states. The organization's example of the alternative is Olmo, a large language model from Ai2 released with its full training dataset and model checkpoints. Researchers can inject new information into the training data and watch how the model memorized or forgot it.
That kind of study is not academic housekeeping. A separate paper on arXiv, "Extracting memorized pieces of (copyrighted) books from open-weight language models," submitted in May 2025 and revised through July 2026, applies an extraction technique to 200 books and 14 open-weight LLMs across more than 3,000 experiments. The authors, a group that includes A. Feder Cooper, Mark A. Lemley and Percy Liang, found that most LLMs do not memorize most books, either in whole or in part, but that there are notable exceptions. Llama 3.1 70B entirely memorizes some books, the paper says, including Harry Potter and the Sorcerer's Stone, to the point that the whole book can be extracted almost verbatim from the first few words as a prompt.
The paper's conclusion is deliberately unsatisfying for both sides of the copyright fight: memorization varies by model and by book, and the results do not unambiguously favor plaintiffs or defendants. Open weights are what made that measurement possible at all. Closed models cannot be probed this way from the outside.
So the practical picture is messier than either camp's talking points. Open-weight releases are competitive on price and, on some benchmarks, within a point or two of the frontier. They are also expensive to run at the top end, unevenly documented, and, by OSI's definition, not open source. The gap that Mozilla measures is real, and so is the gap OSI describes between releasing weights and releasing a model you can actually inspect.
Sources
3- 01China's open-weight AI models are now just 4 months behind frontier US offerings, Mozilla report claimsEN
- 02Open Weights Are Good. Open Source Is Better.EN
- 03Extracting memorized pieces of (copyrighted) books from open-weight language modelsEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.