What open weights actually means, and what the models still hide
Open-weight releases let anyone download a model's parameters, but a Mozilla report published on 15 September claims the best open model still trails the closed leader by about four months, and the Open Source Initiative argues the label covers far less than the four software freedoms.

"Open weights" is now a marketing phrase as much as a technical one. The plain version: the trained parameters of a model are published as files you can download, run and fine-tune. That is it. It says nothing about the data the model learned from, the code that trained it, or the licence attached to it.
The distinction matters because the term is being used in policy fights, procurement decisions and, increasingly, national security arguments. So it is worth separating what an open-weight release gives you from what it withholds.
The four freedoms test
The Open Source Initiative, which maintains the definition most of the software industry uses, laid out its position in a blog post on 19 September. AI models, it notes, have three main components: training data, weights or parameters, and code for data preparation and training. Which of those you can access decides whether a model is closed, open-weight, or Open Source AI.
Open weights get you one of the three. You can run the model locally and pay only for electricity and hardware, or hand it to a third-party host. You can fine-tune it on your own data. That is more freedom than a fully closed system offers, where you use the vendor's infrastructure and pay the vendor's fees.
Open-weight models limit your ability to modify the model, and you cannot fully study it.
The OSI's argument is that this is not enough to satisfy the four software freedoms: to use, study, modify and share without asking permission from the rights holder. Without the training code and either the dataset or a detailed account of how it was built, you cannot inspect why a model produces a given output. The organisation points to Ai2's Olmo as a counterexample, a large language model released with its full training dataset and checkpoints, which let researchers inject new information into the training data and watch how the model memorised or forgot it.
That kind of experiment is the practical difference between the two labels. It is also, on the OSI's reading, the difference between being able to verify a model and having to take someone's word for it.
The gap, and how it is measured
Mozilla published version 1.1 of its State of Open Source AI report on 15 September, using data current to 1 September, according to Tom's Hardware. The headline finding: the best open model trailed the closed leader on the Artificial Analysis Intelligence Index by three points, at 60% of the price, and sat two points behind Claude Fable 5 at 30% of the cost.
Mozilla's fit on METR task-horizon data puts the open-closed gap at roughly 4.4 months, in line with Epoch AI's four-month estimate. METR, a research nonprofit, scores models by the length of task, measured in human working time, that they complete half the time. By Mozilla's estimate, closed models handle tasks that take human experts 8 to 12 hours; open models reach that about four months later. Mozilla computes open capability as doubling every 3.9 months against 5.5 months for closed.
The report is not a neutral document, and Tom's Hardware says so. Mozilla is the nonprofit behind Firefox and an advocate for open models. TIME reported on 14 July that Mozilla chief technology officer Raffi Krikorian described the report as partly advocacy. It is built on a Mozilla/SlashData survey of roughly 1,400 developers, plus OpenRouter traffic data and third-party benchmark indices, and counts 16 notable open releases, none of which delivers the data recipe the OSI's definition requires.
There are other caveats worth keeping in view. The four-month figure and the 30% token price comparison are measured API to API on hosted endpoints, at list price. Mozilla's own hardware chart tells a starker story: the best open model that fits on one server scores 52.6, and the best that fits on one GPU scores 40, drops of 10 and 23 points from the top. Kimi K3's native MXFP4 checkpoint runs about 1.56TB across 96 shards, and Mozilla's serving configuration lists 64 or more accelerators, while vLLM calls for at least eight GB300 GPUs with multiple nodes for production traffic. The report describes this as open but not runnable by most who hold it.
What the market is actually doing
Adoption and revenue point in different directions. On OpenRouter, a marketplace that routes developer traffic to hundreds of models, Mozilla counted eight of the top ten models by August token volume as open weights, seven of them Chinese-built. Yet closed providers took 96% of model-layer revenue on OpenRouter between May and September 2025, according to the Linux Foundation.
"We see the decision to pay for closed [models] as workload-specific rather than organization-specific," Krikorian told Ars Technica in an email.
The benchmark picture moves fast enough to complicate any static comparison. Mozilla's data stops at 1 September. Since then, Artificial Analysis has moved its index to v4.3 with a different evaluation set; the live board shows Claude Fable 5.1 at 53 on its highest effort setting with Kimi K3 at 44, numbers Tom's Hardware notes are not comparable to the v4.1.1 figures Mozilla plotted. The vals.ai Terminal-Bench 2.1 board, updated 11 September, is led by GPT-6 Astra at 87.27% with Fable 5.1 at 85.02%. Mozilla's own chart caption reads: "the gap resets every release cycle."
One more thread runs through the Kimi K3 discussion. The model carries an allegation detailed in the 8 September NSA/CISA/FBI joint advisory AA26-251A, which Mozilla's report states as "asserted, and unshown": that Moonshot extracted Claude Fable 5 data to train K3 through distillation, the practice of training one model on another's outputs. On 17 July, Artificial Analysis had K3 at 57 against Fable 5's 60; by 1 September, Mozilla had it two points back.
Why the label is being fought over
Both the OSI and Mozilla are arguing for more openness, from different angles. The OSI wants the term "open source" reserved for releases that carry all three components, on the grounds that trust and verification questions get harder as AI spreads into daily life. Mozilla wants the measured capabilities of open models taken seriously, while conceding that the leading ones are not runnable by most of the people who can download them.
For anyone reading a model card, the practical upshot is narrow and testable. Check whether the weights are downloadable, what licence governs them, whether training code or data documentation is included, and what hardware the thing actually needs. The phrase on the front of the page will not answer any of those questions.
Sources
2- 01China's open-weight AI models are now just 4 months behind frontier US offerings, Mozilla report claimsEN
- 02Open Weights Are Good. Open Source Is Better.EN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.