Open weights are not open source, and the gap is starting to show
Almost every release labelled an open model ships weights you can download and nothing else. The Register argued on 15 September that the label is being stretched. A day earlier, a Mozilla report put hard numbers on what the two categories actually buy you.

Weights are the learned numerical parameters that training produces. On their own they let you run a model, self-host it, fine-tune it on internal documents and skip a vendor's API. That is a real thing to own. It is also, per the Open Source Initiative, only "a fraction of the information required for full accountability."
The Register's Steven J. Vaughan-Nichols made the case on 15 September that the industry has quietly swapped one word for another. A company publishes model files to Hugging Face. Developers run them on their own GPUs. The release is promptly described as an open source model. Not necessarily, he argues. It may only be open-weight.
The distinction sounds academic. It is not.
Without training data or detailed documentation, outsiders cannot establish which sources went into a model, what copyrighted or private material may have been included, how data was filtered, which languages were underrepresented, whether benchmark data leaked into training, or what alignment work was done after pretraining. James Landay, director of the Stanford Institute for Human-Centered AI, told The Register that open weights are progress, because you can keep the model out of someone else's data pipeline. But he added: "You still can't see how the thing was built, what it was trained on, or why it behaves the way it does. That's not an open model. That's open distribution."
The Open Source Initiative tried to settle this with its Open Source AI Definition, OSAID 1.0, released in October 2024. It requires model parameters, weights included, to be available under OSI-approved terms, but does not prescribe a specific legal mechanism. That gap has not closed. Luca Antiga, CTO of Lightning AI, called the treatment of weights "a gaping hole that will make licenses less effective." Bruce Perens, author of the original Open Source Definition, denounced OSAID in 2024 and later said the OSI itself is now involved in openwashing. Bradley Kuhn of the Software Freedom Conservancy and Red Hat's Richard Fontana have called for the definition to be repealed.
What the weights actually get you
Mozilla published version 1.1 of its State of Open Source AI report on 15 September, using data current to 1 September. It is a useful reality check on both sides of the argument. The best open model trailed the closed leader on the Artificial Analysis Intelligence Index by three points at 60% of the price, and sat two points behind Claude Fable 5 at 30%. Fitting METR task-horizon data, Mozilla put the open-closed gap at around 4.4 months, close to Epoch AI's four-month estimate. Open capability doubles every 3.9 months against 5.5 for closed, by Mozilla's computation.
Mozilla is an advocate for open models, and the report is not neutral. TIME reported on 14 July that Mozilla CTO Raffi Krikorian described it as partly advocacy. The report's own numbers also cut against the headline. It counts 16 notable open releases, and none supplies the data recipe the OSI definition requires. The four-month gap and the 30% price figure are measured API to API on hosted endpoints at list price.
Then there is the hardware. Mozilla's chart puts the best open model that fits one server at 52.6 and the best that fits a single GPU at 40, drops of 10 and 23 points. Kimi K3's native MXFP4 checkpoint runs about 1.56TB across 96 shards. The serving configuration lists 64 or more accelerators, and vLLM calls for at least eight GB300 GPUs with multiple nodes for production traffic. Mozilla's own phrase for this is open but not runnable by most who hold it. One exception: Thinking Machines' Inkling-Small, under Apache 2.0, fits a single B300 at a 180GB floor.
Krikorian told Ars Technica by email that the choice to pay for closed models looks workload-specific rather than organization-specific. The money agrees. Closed providers took 96% of model-layer revenue on OpenRouter from May to September 2025, according to the Linux Foundation. Mozilla, meanwhile, counted eight of the top ten models by August token volume as open weights, seven of them Chinese-built.
That last detail matters for the licensing fight. Mozilla's data stops at 1 September, and Artificial Analysis has since moved its index to v4.3 with a different evaluation set, so the plotted numbers are no longer directly comparable to the live board. Mozilla's own chart caption reads: "the gap resets every release cycle." A 4.4-month lag is a snapshot, not a trend line.
The copyright angle is where the missing documentation gets expensive. A paper by A. Feder Cooper, Mark A. Lemley and co-authors, posted to arXiv and revised on 20 July 2026, ran over 3,000 experiments across 200 books and 14 open-weight LLMs. Most models did not memorize most books, in whole or in part. But Llama 3.1 70B entirely memorized some titles, including Harry Potter and the Sorcerer's Stone, enough that the whole book can be extracted almost verbatim from its first few words. The authors say the result favours neither side in court.
So the label dispute is not semantics. If you cannot see the data, you cannot audit what a model memorized, and you cannot reproduce the system that produced it. Landay's framing is the cleanest one on offer. Open weights answer whether you can run this. Open source answers whether you can trust it, improve it and build the next thing on top.
Sources
3- 01Open weights are not open source: Why AI's favorite label is under disputeEN
- 02China's open-weight AI models are now just 4 months behind frontier US offerings, Mozilla report claimsEN
- 03Extracting memorized pieces of (copyrighted) books from open-weight language modelsEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.