What 'open weights' actually means, and why it is not the same as open source
Downloadable model files are now routinely described as open source. The Open Source Initiative and Stanford researchers say that label covers only a fraction of what would be needed to inspect, reproduce or rebuild a model.

Open weights are the learned numerical parameters produced by training a neural network. Publish them on Hugging Face and anyone with a GPU can run the model, fine-tune it or self-host it. That is genuinely useful. It is also, according to the Open Source Initiative and several prominent free-software figures, not the same thing as open source.
The distinction matters more than a labelling argument. It decides whether you can merely deploy a finished neural network or whether you can inspect, reproduce, alter and redistribute the system that produced it. As Stanford HAI director James Landay told The Register: "Open weights answer 'Can I run this?' Open source answers 'Can I trust this, improve it, and build the next thing on top of it?'"
What you actually get in the download
A typical open-weight release contains the weights, the model architecture and the inference code. Together with a runtime such as llama.cpp, that is enough to generate text. Superwhisper's S1-mini, for example, is a 0.6B-parameter text normalizer for speech-to-text output, fine-tuned from Qwen/Qwen3-0.6B and licensed Apache 2.0 plus a naming clause. Its model card lists 596M unique parameters, 28 layers, 16 query heads and 8 key/value heads under grouped-query attention, and a quantized build that is a 462 MiB file.
What is missing from that card is the training data. Superwhisper does not say what transcripts S1-mini was trained on, how they were selected, or what was removed. That is normal for open-weight releases and it is exactly the gap critics point at.
The OSI's own glossary entry draws the line bluntly. "Open Weights refer to the final weights and biases of a trained neural network," it says, adding that weights alone expose only "a fraction of the information required for full accountability." Without training data or detailed documentation, outsiders cannot determine which sources were used, what copyrighted or private material may have been included, how data was filtered, which languages were underrepresented, whether benchmark data leaked into training, or what alignment and safety methods were applied after pretraining.
"You still can't see how the thing was built, what it was trained on, or why it behaves the way it does. That's not an open model. That's open distribution." James Landay, Stanford HAI
The OSI tried to define it, and the fight is not over
The Open Source Initiative released the Open Source AI Definition, OSAID 1.0, in October 2024. It requires model parameters, including weights, to be made available under OSI-approved terms but does not prescribe a specific legal mechanism. The OSI acknowledged at release that the definition would keep evolving. Critics say the central shortcomings have not been resolved.
Luca Antiga, CTO of Lightning AI and a PyTorch contributor, argued that OSAID's treatment of weights leaves "a gaping hole that will make licenses less effective in determining whether OSI-licensed AI systems can be adopted in real-world contexts." Bruce Perens, author of the original Open Source Definition, denounced OSAID in 2024 and later said: "It's not Open Source! … It's unfortunate that the Open Source Initiative itself is now involved in Openwashing."
Bradley Kuhn of the Software Freedom Conservancy and Red Hat senior commercial counsel Richard Fontana have gone further and called for OSAID to be repealed. Their argument, as quoted by The Register: "The OSI acted too quickly to impose an overly ambitious policy compromise on the community. OSAID undeniably created a rift in the FOSS community; that rift seriously damaged the OSI's reputation, authority, and influence."
Meanwhile the Linux Foundation's Mike Dolan submitted the Open Model, Data, and Weights license, OpenMDW, to the OSI. It has existed since 2025 and lists contributors from Amazon, Meta, IBM, Microsoft and Nvidia. Its approach is to define separate terms for a model's architecture, training data and weights under one agreement. The submission has drawn objections on OSI's license review mailing list. Stefano Maffulli, OSI's former executive director, said he keeps getting the impression the review "is tainted by an ideological bias: Because we don't like big tech and AI now is big tech, then we don't like AI; therefore, we'll do anything to block it."
The capability story is real, and so is the caveat
None of this has slowed the release cadence. Mozilla published version 1.1 of its State of Open Source AI report on 15 September using data current to 1 September, according to Tom's Hardware. It puts the best open model three points behind the closed leader on the Artificial Analysis Intelligence Index at 60% of the price, and two points behind Claude Fable 5 at 30%. Mozilla's fit on METR task-horizon data puts the open-closed gap at around 4.4 months, in line with Epoch AI's four-month estimate.
Mozilla's numbers come from a Mozilla/SlashData survey of roughly 1,400 developers, OpenRouter traffic data and third-party benchmark indices. Mozilla advocates for open models; TIME reported on 14 July that Mozilla CTO Raffi Krikorian described the report as partly advocacy. The report counts 16 notable open releases, and notes that none delivers the data recipe required by the OSI definition.
- On vals.ai's Terminal-Bench 2.1, Z.ai's GLM-5.2 scored within a point of Claude Opus 4.7 and about four points behind Opus 4.8, at less than one-fifth the cost per test.
- On OpenRouter, Mozilla counted eight of the top ten models by August token volume as open weights, seven of them Chinese-built.
- Closed providers still took 96% of model-layer revenue on OpenRouter from May to September 2025, per the Linux Foundation.
- The best open model that fits one server scored 52.6 on Mozilla's hardware chart; the best on one GPU scored 40. The drop from the top is 10 and 23 points.
That last point is the practical catch. Kimi K3's native MXFP4 checkpoint runs about 1.56TB across 96 shards, and Mozilla's serving configuration lists 64 or more accelerators, while vLLM calls for at least eight GB300 GPUs with multiple nodes for production traffic. Mozilla describes this as open but not runnable by most who hold it. One exception the report flags is Thinking Machines' Inkling-Small, under Apache 2.0, whose NVFP4 version fits one B300 at a 180GB floor.
Krikorian told Ars Technica in an email: "We see the decision to pay for closed [models] as workload-specific rather than organization-specific."
Copyright cases are running into the same wall
The definition fight has a parallel in court. A paper by A. Feder Cooper, Mark A. Lemley and seven co-authors, posted to arXiv and revised in July 2026, measured memorization of 200 books across 14 open-weight LLMs through more than 3,000 experiments. Their conclusion is that plaintiffs and defendants in generative AI copyright suits both oversimplify.
Most LLMs do not memorize most books, in whole or in part, under the authors' extraction method. But there are exceptions: Llama 3.1 70B entirely memorizes some books, including Harry Potter and the Sorcerer's Stone, to the point where the whole book can be extracted almost verbatim from its first few words. The authors write that their results have significant implications for copyright cases, though not ones that unambiguously favour either side.
That is roughly where the open-weights debate sits. The files are public. The recipe is not. And the two sides cannot even agree on what to call the result.
Sources
4- 01Open weights are not open source: Why AI's favorite label is under disputeEN
- 02China's open-weight AI models are now just 4 months behind frontier US offerings, Mozilla report claimsEN
- 03S1-mini, Superwhisper's first open-weights language modelEN
- 04Extracting memorized pieces of (copyrighted) books from open-weight language modelsEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.