Open weights, open source and one model that memorised Harry Potter
A new arXiv paper reports that Llama 3.1 70B can reproduce at least one book almost verbatim, reviving a debate about what "open weights" actually discloses.

Two things are true about open weights language models at the same time, and the industry has spent three years pretending otherwise. They are genuinely useful. And they are not open source.
The distinction is not semantic hair-splitting. It decides whether anyone outside the releasing lab can audit what a model learned, reproduce how it learned it, or verify claims about bias and safety. The current wave of "open" releases, from Meta's Llama line to Mistral's Mixtral, publishes weights and little else. Training code, datasets and methodology stay behind the wall. Prompt Engineering's December 2023 write-up framed the gap plainly: open weights allows model use but not full transparency, while open source enables model understanding and customization but requires substantially more work to release.
What a weight actually is
Weights are the numerical output of a training run. They are not human-readable, not debuggable, and not instructions. Source code is all three. Prompt Engineering calls the conflation of the two a common error with real consequences. It leads people to describe a proprietary training pipeline as open source simply because a file of parameters is downloadable.
The practical cost shows up in auditability. If only weights are available, developers can run state-of-the-art models but cannot meaningfully evaluate biases, limitations and societal impacts, the same piece argues. Misalignment between a model and real-world needs becomes hard to identify. With full source access, researchers across organisations can conduct rigorous audits of fairness, safety and robustness. That is the trade: convenience now, inspectability never.
Releasing only a model's weights broadly enables application development but concentrates control among a small group of organisations. Enabling open source access distributes control but requires greater commitment to transparency and decentralization.
That framing has aged well, partly because the evidence for the inspectability problem keeps arriving.
The memorisation problem
On 18 May 2025, a group of researchers including A. Feder Cooper, Mark A. Lemley and Percy Liang submitted a paper to arXiv. They wanted to measure how much copyrighted book text open-weight language models can be made to reproduce. They applied their extraction technique to 200 books and 14 open-weight models, running more than 3000 experiments. Version six of the paper, revised on 20 July 2026, is the current one.
The headline finding cuts against both sides of the copyright litigation. Most LLMs do not memorise most books, either in whole or in part, according to the paper. But there are notable exceptions. Llama 3.1 70B entirely memorises some books, including Harry Potter and the Sorcerer's Stone. The memorisation is extensive enough that the whole book can be extracted almost verbatim, deterministically, using its first few words as an initial prompt.
Plaintiffs and defendants in generative AI copyright suits, the authors write, make sweeping and opposing claims about memorisation that dramatically oversimplify the relationship between memorisation and copyright. Their results, they add, have significant implications for copyright cases, though not ones that unambiguously favour either side.
Note what the experiment required. It needed weights that could be downloaded and interrogated. A restricted model would have made the same test impossible for outside researchers. Open weights, in other words, enabled the finding even though open weights is not open source.
A small counterexample
Not every open-weights release is a frontier model. Superwhisper published s1-mini, a 0.6B-parameter text normaliser for speech-to-text output, under Apache 2.0 plus a naming clause. It takes a raw ASR transcript and rewrites it as clean written text: fillers removed, false starts resolved, punctuation and capitalisation applied, spoken numbers and dates rendered in written form. On a held-out set of 7,519 English cases it reaches 94.8% token accuracy, and the quantised build is a 462 MiB file that runs on a laptop CPU.
The release notes are unusually candid about a documentation quirk. The Hugging Face sidebar reports 0.8B parameters, but the model card says the real figure is 596.0M unique parameters. The 155.6M-parameter embedding is stored twice: 751.6M tensor elements against 596.0M unique ones. The layout is inherited from Qwen/Qwen3-0.6B, which reports 0.8B on the Hub for the same reason.
That kind of disclosure is the whole argument for releasing weights with documentation. It is also, precisely, not source code. The model card tells you the base model, the precision, the expected input shape and the licence. It does not give you the training data or the full recipe.
Why the labels matter less than the access
Prompt Engineering proposed splitting the vocabulary: use "Open Weights" for licences covering neural network weights, and "Ethical Weights" for a separate category of licences designed specifically for weights. The suggestion has not been widely adopted, and the marketing has not slowed.
- Open weights: parameters released, training code and data withheld, use and fine-tuning permitted.
- Open source: architecture code, training methodology, hyperparameters, original dataset and documentation released.
- Restricted weights: parameters available only under significant legal or ethical limitations.
The distinction has a second-order effect that rarely makes the announcement posts. Open sourcing empowers decentralised progress, in Prompt Engineering's phrasing, because researchers worldwide can propose improvements and alternative development directions. Progress stops being bottlenecked by the small number of companies able to train massive proprietary models. Open weights concentrates that capability instead, while distributing the artefacts.
None of this means open-weights releases are worthless. The memorisation study exists because they are not. But a downloadable checkpoint is a licence to use a model, not a licence to understand it, and the gap between those two things is where most of the current argument about AI openness is actually happening.
Sources
3- 01Openness in Language Models: Open Source vs Open Weights vs Restricted WeightsEN
- 02Extracting memorized pieces of (copyrighted) books from open-weight language modelsEN
- 03S1-mini, Superwhisper's first open-weights language modelEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.