Open weights are not open source, and the gap is getting harder to measure
A Mozilla report says China's best open-weight models now trail US frontier offerings by about four months, while critics argue the term "open weights" is being stretched well past what the Open Source Initiative's definition allows.

Mozilla published State of Open Source AI version 1.1 on 15 September, drawing on data current to 1 September. The report counts 16 notable open releases. According to the report as described by Tom's Hardware, none of them ships the training data recipe that the Open Source Initiative's Open Source AI Definition requires.
That is the tension running under this month's open-weights news cycle. Downloading a model is easy. Knowing what went into it, or rebuilding it, is not.
The four-month number, and its caveats
Mozilla's central claim is a narrowing gap. The best open model trailed the closed leader on the Artificial Analysis Intelligence Index by three points at 60% of the price, and sat two points behind Claude Fable 5 at 30%. Mozilla's fit on METR task-horizon data puts the open-closed gap at around 4.4 months, in line with Epoch AI's four-month estimate. By Mozilla's computation, open capability doubles every 3.9 months against 5.5 for closed models.
The report leans on a Mozilla/SlashData survey of roughly 1,400 developers, OpenRouter traffic data and third-party benchmark indices. Mozilla is an advocate for open models, and TIME reported on 14 July that Mozilla chief technology officer Raffi Krikorian described the report as partly advocacy. Mozilla says the four-month figure and the 30% token price number are measured API to API on hosted endpoints and at list price.
Hardware is where the story gets less flattering. The report's own chart puts the best open model that fits on one server at 52.6 and the best that fits on one GPU at 40. The drop from the top is 10 and 23 points respectively, a larger gap than four months suggests. Kimi K3's native MXFP4 checkpoint runs about 1.56TB across 96 shards, and Mozilla's serving configuration lists 64 or more accelerators. The report describes this as open but not runnable by most who hold it.
There is a counterexample. Thinking Machines' Inkling-Small, under Apache 2.0, has an NVFP4 version that fits one B300 at a 180GB floor.
What "open" is being asked to mean
The Register's Steven J. Vaughan-Nichols argued on 15 September that the industry abuses the word. Publishing weights to Hugging Face and letting developers run them on their own GPUs, he wrote, is open-weight distribution, not open source. The distinction, he said, decides whether you can merely deploy a finished network or inspect, reproduce, alter and redistribute the system that produced it.
The OSI makes the distinction directly: "Open Weights refer to the final weights and biases of a trained neural network." Weights alone, the OSI adds, expose only "a fraction of the information required for full accountability."
James Landay, director of the Stanford Institute for Human-Centered AI, told The Register: "Open weights are progress. You can download the model, run it on your own machine, keep it out of someone else's data pipeline. But you still can't see how the thing was built, what it was trained on, or why it behaves the way it does. That's not an open model. That's open distribution."
Landay's concern is specific. Without training data or detailed documentation, outsiders cannot tell which sources were used, what copyrighted or private material may have been included, how data was selected or removed, which languages were underrepresented, whether benchmark data leaked into training, or what alignment methods shaped the model after pretraining.
The Open Source AI Definition 1.0, released in October 2024, requires model parameters including weights to be made available under OSI-approved terms but does not prescribe a legal mechanism. Luca Antiga, CTO of Lightning AI, has said that treatment leaves "a gaping hole" that will weaken licences as a signal of real-world adoptability. Bruce Perens, author of the original Open Source Definition, denounced OSAID in 2024 and later said the OSI itself is now involved in openwashing. Bradley Kuhn of the Software Freedom Conservancy and Red Hat's Richard Fontana have called for OSAID to be repealed.
Practical memorisation questions sit underneath all of it. A paper by A. Feder Cooper, Mark A. Lemley and co-authors, submitted to arXiv in May 2025 and revised through July 2026, ran over 3,000 experiments across 200 books and 14 open-weight LLMs. Most LLMs did not memorise most books, the authors found. But Llama 3.1 70B entirely memorised some titles, including Harry Potter and the Sorcerer's Stone, enough that the whole book could be extracted almost verbatim from its first few words.
Meanwhile the open-weight ecosystem keeps shipping. Superwhisper published S1-mini, a 0.6B-parameter text normaliser for speech-to-text output, fine-tuned from Qwen/Qwen3-0.6B under Apache 2.0 plus a naming clause. It reports 94.8% token accuracy on 7,519 held-out English cases and a 462 MiB quantised build that runs on a laptop CPU.
Small, downloadable, and documented enough to run. The accountability question is a separate file.
Sources
4- 01China's open-weight AI models are now just 4 months behind frontier US offerings, Mozilla report claimsEN
- 02Open weights are not open source: Why AI's favorite label is under disputeEN
- 03Extracting memorized pieces of (copyrighted) books from open-weight language modelsEN
- 04S1-mini, Superwhisper's first open-weights language modelEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.