Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Open weights are not open source: the fight over a label

Open-weight language models are shipping faster than the rules that describe them. The Open Source Initiative says weights alone expose only "a fraction of the information required for full accountability," and critics say the label is being stretched.

AI & modelsAnalysisGrace OkonkwoPublished: 27 September 20267 min readSources 3
Open weights are not open source: the fight over a label

Downloading a model file from Hugging Face takes seconds. Finding out what went into it takes a subpoena. That gap is the whole argument.

The Register's Steven J. Vaughan-Nichols drew the distinction bluntly on 15 September. A release can be open-weight without being open source. The difference decides whether you can merely run a finished neural network or actually inspect, reproduce, alter and redistribute the system that produced it. Weights are the learned numerical parameters that training produces. Together with the architecture and the inference code, they make a large language model work. You can self-host an open-weight model, fine-tune it on internal documents, and keep your prompts off a proprietary API. What you cannot do, in most releases, is see the data recipe behind it.

That is not a technicality to the people who write licences.

OSI draws a line, then gets criticised for it

The Open Source Initiative, which stewards the Open Source Definition, states plainly that "Open Weights refer to the final weights and biases of a trained neural network," and that releasing them exposes only "a fraction of the information required for full accountability." Its own Open Source AI Definition, OSAID 1.0, was meant to close that gap. It requires model parameters, including weights, to be available under OSI-approved terms, but it does not prescribe a specific legal mechanism. Released in October 2024, it was acknowledged at the time as a document that would keep evolving.

Critics say the central shortcoming never got fixed. Luca Antiga, CTO of Lightning AI and a prominent PyTorch contributor, has argued that OSAID's treatment of weights leaves "a gaping hole that will make licenses less effective in determining whether OSI-licensed AI systems can be adopted in real-world contexts." Bruce Perens, author of the original Open Source Definition, denounced OSAID in 2024 and later declared: "It's not Open Source! … It's unfortunate that the Open Source Initiative itself is now involved in Openwashing." Bradley Kuhn of the Software Freedom Conservancy and Red Hat Senior Commercial Counsel Richard Fontana have called for OSAID to be repealed. The OSI, they argue, "acted too quickly to impose an overly ambitious policy compromise on the community," and the rift "seriously damaged the OSI's reputation, authority, and influence."

James Landay, director of the Stanford Institute for Human-Centered AI, gave The Register the practical version of the complaint: "Open weights are progress. You can download the model, run it on your own machine, keep it out of someone else's data pipeline. But you still can't see how the thing was built, what it was trained on, or why it behaves the way it does. That's not an open model. That's open distribution."

"Open weights answer 'Can I run this?' Open source answers 'Can I trust this, improve it, and build the next thing on top of it?'" said Landay.

Without training data or detailed documentation, outsiders cannot determine which sources were used, what copyrighted or private material may have been included, how data was selected or removed, which languages and communities were underrepresented, whether benchmark data leaked into training, or what alignment and safety methods shaped the model after pretraining.

A licence trying to split the difference

Into that vacuum came the Linux Foundation's Open Model, Data, and Weights licence, submitted to the OSI by Mike Dolan. OpenMDW has been around since 2025 and lists contributors from Amazon, Meta, IBM, Microsoft and Nvidia, which gives it industry weight. Its approach is to define separate terms for a model's architecture, training data and weights, bringing everything a licensor supplies under one agreement. Conventional open source revolves around source code. LLMs combine code, architecture and numerical weights derived from datasets that may be proprietary, copyrighted or simply undisclosed.

The submission has run into objections on OSI's licence review mailing list. Stefano Maffulli, OSI's former executive director, who led the organisation while OSAID was being formulated, said: "I continue getting the impression that the OpenMDW review is tainted by an ideological bias: Because we don't like big tech and AI now is big tech, then we don't like AI; therefore, we'll do anything to block it."

Meanwhile the models keep shipping. Mozilla's State of Open Source AI report, version 1.1, published 15 September with data current to 1 September, counted 16 notable open releases. None delivers the data recipe OSAID requires. Mozilla is an advocate for open models, and its CTO, Raffi Krikorian, told TIME on 14 July that the report is partly advocacy, according to Tom's Hardware.

The capability gap is closing, the cost gap is not

Mozilla's numbers are the most concrete measure of where open weights now stand. The best open model trailed the closed leader on the Artificial Analysis Intelligence Index by three points at 60% of the price, and sat two points behind Claude Fable 5 at 30% of the price. Mozilla's fit on METR task-horizon data, which scores models by the length of task in human working time they complete half the time, puts the open-closed gap at around 4.4 months, in line with Epoch AI's four-month estimate. Open capability doubles every 3.9 months, by Mozilla's computation, against 5.5 months for closed models.

On vals.ai's Terminal-Bench 2.1 board, which runs every model through the same harness, Z.ai's GLM-5.2 scored within a point of Claude Opus 4.7 and about four points behind Opus 4.8, at less than one-fifth the cost per test. On OpenRouter, Mozilla counted eight of the top ten models by August token volume as open weights, seven of them Chinese-built. Closed providers still took 96% of model-layer revenue on OpenRouter from May to September 2025, per the Linux Foundation. Krikorian told Ars Technica in an email: "We see the decision to pay for closed [models] as workload-specific rather than organization-specific."

Those figures come with caveats the report itself states. The four-month gap and the 30% token price are measured API to API on hosted endpoints and at list price. Mozilla's own hardware chart puts the best open model that fits one server at 52.6 and the best that fits one GPU at 40, drops of 10 and 23 points from the top, a larger gap than four months. Kimi K3's native MXFP4 checkpoint runs about 1.56TB across 96 shards. Mozilla's serving configuration lists 64 or more accelerators, while vLLM calls for at least eight GB300 GPUs, with multiple nodes for production traffic. The report describes this as open but not runnable by most who hold it. Thinking Machines' Inkling-Small, under Apache 2.0, is one exception: its NVFP4 version fits a single B300 at a 180GB floor.

The data stops at 1 September, and the leaderboards have moved since. Artificial Analysis now runs index v4.3 with a different evaluation set; its live board has Claude Fable 5.1 at 53 on the highest effort setting and Kimi K3 at 44, numbers not comparable to the v4.1.1 figures Mozilla plotted. The caption on Mozilla's own chart reads: "the gap resets every release cycle."

There is also an unresolved allegation attached to K3. A 8 September joint advisory from the NSA, CISA and FBI (AA26-251A) asserts that Moonshot extracted Claude Fable 5 data to train K3 through distillation, the practice of training one model on another model's outputs. Mozilla's report states the claim as "asserted, and unshown." On 17 July, Artificial Analysis had K3 at 57 against Fable 5's 60; by 1 September, Mozilla had it two points back.

What the label is actually for

None of this settles the copyright fights. A paper by A. Feder Cooper, Mark A. Lemley and co-authors, first submitted to arXiv in May 2025 and revised through July 2026, tested 200 books against 14 open-weight models in more than 3000 experiments. Most LLMs do not memorise most books, in whole or in part, the authors found. But Llama 3.1 70B entirely memorises some titles, including Harry Potter and the Sorcerer's Stone, to the point where the whole book can be extracted almost verbatim from its first few words. The paper's conclusion is that memorisation varies by model and by book, and that the results carry significant implications for copyright cases, "though not ones that unambiguously favor either side."

That is the awkward middle the label sits in. Open weights tell you a model can be run and fine-tuned. They do not tell you what it ate. Landay's framing, as quoted by The Register, is the cleanest version: open weights answer whether you can run the thing, open source answers whether you can trust it, improve it, and build on it. Right now, for almost every release on the Hub, the first answer is yes and the second is still a question.

Comments 0

Sources

3
  1. 01Open weights are not open source: Why AI's favorite label is under disputeEN
  2. 02China's open-weight AI models are now just 4 months behind frontier US offerings, Mozilla report claimsEN
  3. 03Extracting memorized pieces of books from open-weight language modelsEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.