Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Open weights, closed recipes: what the memorization paper and the licensing fight actually show

A COLM 2026 paper measured how much 14 open-weight language models memorized 200 books. Most models do not reproduce most books, but Llama 3.1 70B can be prompted from a book's first few words into near-verbatim output of some titles. The finding sits in the middle of a dispute over what "open" is supposed to mean.

AI & modelsAnalysisGrace OkonkwoPublished: 27 September 20265 min readSources 3
Open weights, closed recipes: what the memorization paper and the licensing fight actually show

The paper is Extracting memorized pieces of (copyrighted) books from open-weight language models, arXiv:2505.12546. It was first submitted on 18 May 2025 and revised through 20 July 2026. The authors are A. Feder Cooper, Mark A. Lemley, Allison Casasola, Ahmed Ahmed, Aaron Gokaslan, Amy B. Cyphert, Christopher De Sa, Daniel E. Ho and Percy Liang. COLM 2026 accepted it.

The authors took aim at how the lawsuits argue. Plaintiffs and defendants in generative AI copyright cases, the abstract says, "often make sweeping, opposing claims about the extent to which large language models (LLMs) memorize protected expression from books in their training data." Both sides, the authors write, "dramatically oversimplify the relationship between memorization and copyright."

So they built a measurement technique and ran it at scale: 200 books, 14 open-weight models, more than 3000 experiments. The headline result undercuts the strongest claims. "With respect to our specific extraction methodology, we find that most LLMs do not memorize most books -- either in whole or in part."

Then come the exceptions. Llama 3.1 70B "entirely memorizes some books, like Harry Potter and the Sorcerer's Stone." Memorization there runs deep enough that, per the abstract, "one can deterministically extract the whole book almost verbatim using the book's first few words as an initial prompt." The finding is model-specific and book-specific, and that is the point: memorization is not a property of LLMs in general.

Whether that helps anyone in court is left open. The results, the abstract says, "have significant implications for copyright cases, though not ones that unambiguously favor either side."

The label problem underneath

The models in that study are open-weight. You can download them. That is not the same as open source, and The Register made the case on 15 September 2026 in a column by Steven J. Vaughan-Nichols. The Open Source Initiative, steward of the Open Source Definition, defines open weights as "the final weights and biases of a trained neural network," and adds that weights alone expose only "a fraction of the information required for full accountability."

James Landay, director of the Stanford Institute for Human-Centered AI, gave The Register the practical version: "Open weights are progress. You can download the model, run it on your own machine, keep it out of someone else's data pipeline. But you still can't see how the thing was built, what it was trained on, or why it behaves the way it does. That's not an open model. That's open distribution."

"Open weights answer 'Can I run this?' Open source answers 'Can I trust this, improve it, and build the next thing on top of it?'"

That second quote is also Landay's, from the same Register column. The column reports criticism of the OSI's Open Source AI Definition 1.0 from Bruce Perens, Bradley Kuhn and Richard Fontana. When the OSI released OSAID 1.0 in October 2024, it acknowledged the definition would keep evolving. The Linux Foundation's Mike Dolan has submitted the Open Model, Data, and Weights license. It has been in circulation since 2025, with contributors listed from Amazon, Meta, IBM, Microsoft and Nvidia, and the submission has drawn objections on the OSI's license review mailing list.

This matters for the memorization result. If you cannot see the training data, you cannot independently check which books went in. The extracted text is your only evidence about what came out.

Cheap, close, and hard to run

Mozilla published version 1.1 of its State of Open Source AI report on 15 September 2026, with data current to 1 September, according to Tom's Hardware. Mozilla is the nonprofit behind Firefox and an advocate for open models. The report draws on a Mozilla/SlashData survey of roughly 1,400 developers, OpenRouter traffic data and third-party benchmark indices. It counts 16 notable open releases, and Tom's Hardware notes that none delivers the data recipe the OSI definition requires.

The capability numbers: the best open model trailed the closed leader on the Artificial Analysis Intelligence Index by three points at 60% of the price, and sat two points behind Claude Fable 5 at 30%. Mozilla's fit on METR task-horizon data puts the open-closed gap at about 4.4 months, close to Epoch AI's four-month estimate. By Mozilla's computation, open capability doubles every 3.9 months, against 5.5 for closed.

The caveats are worth as much as the headline. The four-month gap and the 30% figure are measured API to API on hosted endpoints at list price. Mozilla's own hardware chart puts the best open model that fits one server at 52.6 and the best on one GPU at 40, drops of 10 and 23 points from the top. That spread is wider than four months suggests. Kimi K3's native MXFP4 checkpoint runs about 1.56TB across 96 shards; Mozilla's serving configuration lists 64 or more accelerators, and vLLM calls for at least eight GB300 GPUs. The report calls this open but not runnable by most who hold it.

On the demand side, Mozilla counted eight of the top ten models by August token volume on OpenRouter as open weights, seven of them Chinese-built. Closed providers still took 96% of model-layer revenue on OpenRouter from May to September 2025, per the Linux Foundation. Mozilla CTO Raffi Krikorian told Ars Technica by email that the choice to pay for closed models looks workload-specific rather than organization-specific.

One more thread. The report states, as "asserted, and unshown," an allegation from the 8 September NSA/CISA/FBI joint advisory AA26-251A that Moonshot extracted Claude Fable 5 data to train Kimi K3 through distillation. Unproven, and unresolved.

Put together, the two documents describe the same object from different ends. A model can be downloaded, fine-tuned and, in at least one documented case, prompted into reproducing a novel almost verbatim. The training data behind it stays undocumented, unlicensed as open source and, for the largest checkpoints, out of reach of the hardware most people have.

Comments 0

Sources

3
  1. 01Extracting memorized pieces of (copyrighted) books from open-weight language modelsEN
  2. 02Open weights are not open source: Why AI's favorite label is under disputeEN
  3. 03China's open-weight AI models are now just 4 months behind frontier US offerings, Mozilla report claimsEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.