Memorisation Study Finds Most LLMs Do Not Reproduce Most Books
A COLM 2026 paper reports that across 200 books and 14 open-weight language models, most models do not memorise most books, though Llama 3.1 70B can reproduce at least one in full.

Most large language models do not memorise most of the books they were trained on, according to a paper accepted at COLM 2026. The claim comes from a study that ran more than 3,000 extraction experiments across 200 books and 14 open-weight models. It cuts against the sweeping assertions both sides have made in copyright litigation over generative AI.
The work is titled "Extracting memorized pieces of (copyrighted) books from open-weight language models." It was first posted to arXiv in May 2025 and revised through July 2026. Its authors include A. Feder Cooper, Mark A. Lemley, Allison Casasola, Ahmed Ahmed, Aaron Gokaslan, Amy B. Cyphert, Christopher De Sa, Daniel E. Ho and Percy Liang. The abstract states plainly that the polarized positions taken by plaintiffs and defendants "dramatically oversimplify the relationship between memorization and copyright."
What the extraction method actually found
The headline result is a negative one, and it is narrower than it sounds. With respect to the authors' specific extraction methodology, most LLMs do not memorise most books, either in whole or in part. That qualifier matters. The finding is a statement about what this particular technique could pull out. It is not a general proof that training data leaves no trace.
The exceptions are where the paper gets uncomfortable for model developers. Llama 3.1 70B entirely memorises some books, the authors write, naming Harry Potter and the Sorcerer's Stone. The memorisation is extensive enough that the whole book can be extracted almost verbatim, deterministically, using the book's first few words as an initial prompt. No sampling tricks, no jailbreak, just the opening phrase.
Memorisation varied both by model and by book across the 14 open-weight systems tested. The abstract does not name the other models, and it does not give a per-book breakdown. The published summary leaves one practical question open: which titles are most exposed outside the Harry Potter example.
The authors are careful about what their results mean for court. They conclude that the findings have significant implications for copyright cases, "though not ones that unambiguously favor either side." That sentence is the paper's real contribution to the debate, and both sets of lawyers are likely to quote it out of context.
The licensing argument running alongside the research
While researchers measure what models retain, another fight is underway over what the word "open" means when attached to them. The Register published a column on 15 September by Steven J. Vaughan-Nichols arguing that the industry routinely describes releases as "open source models" when they are only open-weight.
The distinction, as the column frames it, determines whether you can merely deploy a finished neural network or whether you can inspect, reproduce, alter and redistribute the system that produced it. Weights are the learned numerical parameters created by training. Together with the architecture and inference code, they make a model run. They do not tell you what went into it.
The Open Source Initiative draws the line directly. Its position, quoted in the column, is that open weights refer to the final weights and biases of a trained neural network, and that releasing them exposes only "a fraction of the information required for full accountability."
"Open weights are progress. You can download the model, run it on your own machine, keep it out of someone else's data pipeline. But you still can't see how the thing was built, what it was trained on, or why it behaves the way it does. That's not an open model. That's open distribution."
That is James Landay, director of the Stanford Institute for Human-Centered AI, speaking to The Register. He added that there is "a wide gap between open-weight AI and open source AI." Without disclosed training data, or a documented and auditable account of it, he argued, outsiders cannot test or reproduce the work in the fullest sense.
Without the training data, the column notes, outsiders cannot determine which sources were used, what copyrighted or private material may have been included, how data was selected or removed, which languages and communities were underrepresented, whether benchmark data leaked into training, or what alignment methods shaped the model after pretraining. That list is precisely the set of questions the memorisation paper tries to answer empirically, from the outside, with extraction rather than disclosure.
The OSI's own Open Source AI Definition, OSAID 1.0, released in October 2024, requires model parameters including weights to be available under OSI-approved terms but does not prescribe a specific legal mechanism. It has drawn criticism from figures including Bruce Perens, author of the original Open Source Definition, who declared: "It's not Open Source! ... It's unfortunate that the Open Source Initiative itself is now involved in Openwashing." Bradley Kuhn of the Software Freedom Conservancy and Red Hat's Richard Fontana have called for OSAID to be repealed. The OSI, they argue, "acted too quickly to impose an overly ambitious policy compromise on the community."
A competing effort, the Linux Foundation's Open Model, Data, and Weights licence, or OpenMDW, has been submitted to the OSI. It has been around since 2025, lists contributors from Amazon, Meta, IBM, Microsoft and Nvidia, and defines separate terms for a model's architecture, training data and weights. The Register reports that the submission has met objections on OSI's licence review mailing list. Stefano Maffulli, the OSI's former executive director, said he keeps getting the impression the review is "tainted by an ideological bias."
Cheaper, closer, and still not reproducible
Mozilla published version 1.1 of its State of Open Source AI report on 15 September, using data current to 1 September. Tom's Hardware covered it the following day. The report puts the best open model three points behind the closed leader on the Artificial Analysis Intelligence Index at 60% of the price, and two points behind Claude Fable 5 at 30%. Mozilla's fit on METR task-horizon data puts the open-closed gap at around 4.4 months, in line with Epoch AI's four-month estimate.
The methodology deserves scrutiny. METR scores models by the length of task, in human working time, they complete half the time. By Mozilla's fitted estimate, closed models handle tasks taking human experts 8 to 12 hours. Open models reach that about four months later, with open capability doubling every 3.9 months against 5.5 for closed. The report rests on a Mozilla/SlashData survey of roughly 1,400 developers, OpenRouter traffic data and third-party benchmark indices. Mozilla advocates for open models, and TIME reported on 14 July that Mozilla chief technology officer Raffi Krikorian described the report as partly advocacy.
On vals.ai's Terminal-Bench 2.1, which runs every model through the same harness, Z.ai's GLM-5.2 scored within a point of Claude Opus 4.7 and about four points behind Opus 4.8, at less than one-fifth the cost per test. On OpenRouter, Mozilla counted eight of the top ten models by August token volume as open weights, seven of them Chinese-built. Closed providers nevertheless took 96% of model-layer revenue on OpenRouter from May to September 2025, according to the Linux Foundation. "We see the decision to pay for closed [models] as workload-specific rather than organization-specific," Krikorian told Ars Technica in an email.
The report counts 16 notable open releases, and Tom's Hardware notes that none delivers the data recipe required by the OSI definition. The four-month gap and the 30% token price figure are measured API to API on hosted endpoints and at list price. The report's own hardware chart puts the best open model that fits one server at 52.6 and the best on one GPU at 40, drops of 10 and 23 points from the top. That is a wider gap than four months suggests. Kimi K3's native MXFP4 checkpoint runs about 1.56TB across 96 shards, with Mozilla's serving configuration listing 64 or more accelerators, and vLLM calling for at least eight GB300 GPUs with multiple nodes for production traffic. The report describes this as open but not runnable by most who hold it. Thinking Machines' Inkling-Small, under Apache 2.0, is an exception, with an NVFP4 version that fits one B300 at a 180GB floor.
The data stops at 1 September, and the benchmarks have moved since. Artificial Analysis is now on index v4.3 with a different evaluation set, and its live board has Claude Fable 5.1 at 53 on its highest effort setting with Kimi K3 at 44. Those numbers are not comparable to the v4.1.1 figures Mozilla plotted. Mozilla's own chart caption reads: "the gap resets every release cycle."
One further item in the report is stated as an allegation rather than a finding. A 8 September joint advisory from the NSA, CISA and FBI (AA26-251A) asserts that Moonshot extracted Claude Fable 5 data to train K3 through distillation, the practice of training one model on another's outputs. Mozilla's report describes the claim as "asserted, and unshown."
Put the two threads together and the picture is awkward for anyone who wants a clean answer. A model can be downloadable, cheap, within months of the frontier, and still opaque about its training data. And when researchers go looking for what it retained, as the COLM paper did with 200 books and 3,000-plus experiments, they find that the answer depends on which model and which book you ask about. Landay's summary of the gap, as quoted by The Register: "Open weights answer 'Can I run this?' Open source answers 'Can I trust this, improve it, and build the next thing on top of it?'"
Sources
3- 01Extracting memorized pieces of books from open-weight language modelsEN
- 02Open weights are not open source: Why AI's favorite label is under disputeEN
- 03China's open-weight AI models are now just 4 months behind frontier US offerings, Mozilla report claimsEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.