Every team that trains or fine-tunes a model on material it did not write is making copyright decisions, whether it records them or not. Courts in the United States, the United Kingdom and Germany have now ruled on parts of the question, the EU AI Act puts copyright duties on general-purpose model providers, and the answers so far turn less on abstract principle than on facts an engineering team controls: where the data came from, what the crawler respected, what the model memorised and what it can be made to repeat.
This article is the engineering view. It summarises what has been decided as of 3 October 2026, separates the four acts a court actually looks at, and turns each into a control you can build: a provenance ledger, opt-out handling at crawl time, memorisation measurement and an output filter, and a takedown process. It is not legal advice; use it to produce the evidence your counsel will ask for.
What has been decided, as of October 2026
| Matter | Where | What was decided |
|---|---|---|
| Bartz v. Anthropic (2025) | US, N.D. Cal. | Training on lawfully acquired books was fair use; downloading and keeping pirated library copies was not. The class action settled for US$1.5 billion. |
| Kadrey v. Meta (2025) | US, N.D. Cal. | Summary judgment for Meta on training, on that record; the judge stressed that the plaintiffs had not developed market-harm evidence and that such arguments could succeed elsewhere. |
| Thomson Reuters v. ROSS (2025, affirmed 29 Sep 2026) | US, Third Circuit | Westlaw headnotes are copyrightable and using them to train a competing legal search tool was not fair use. The first US federal appellate ruling on AI training; the tool was not generative. |
| New York Times v. OpenAI | US, S.D.N.Y. (consolidated) | Pending; core infringement claims survived dismissal. No ruling on fair use yet. |
| Getty Images v. Stability AI [2025] EWHC 2863 (Ch) | UK High Court | Primary claims were dropped at trial; the secondary claim failed because the model weights were held not to be an infringing copy. Limited trade-mark findings over generated watermarks. |
| GEMA v. OpenAI (11 Nov 2025, 42 O 14139/24) | Germany, Munich Regional Court I | Song lyrics memorised in model parameters and reproduced in outputs infringed; text and data mining exceptions did not cover the memorisation. |
| EU AI Act Article 53 and the GPAI Code of Practice | EU | General-purpose model providers need a copyright policy, must respect machine-readable text and data mining opt-outs, and must publish a summary of training content. |
| US Copyright Office, Part 3 report (pre-publication, May 2025) | US, advisory | Argued that some training uses, especially those that compete with the originals, are unlikely to be fair use. Not binding on courts. |
Two patterns run through these decisions. US courts have so far been more receptive to training itself than to how the data was obtained or whether the product substitutes for the source. European courts and regulators focus on opt-outs and on what the model retains and reproduces. The duties and the timetable under the EU AI Act are covered in the EU AI Act engineering guide and the roles and obligations article.
Four acts, four sets of controls
Separate the problem into four acts, because each has different facts, different law and different controls.
- Acquisition. Copying works into your storage. Bartz turned on this act: the training use was excused, the pirated library was not. Controls: source allowlists, licence records, crawler behaviour, no use of known infringing sources.
- Training copies. The intermediate copies made during tokenisation, shuffling and training. This is where fair use (US) or the text and data mining exceptions (EU, subject to opt-out) are argued. Controls: purpose and opt-out records per document.
- The model. Whether the weights themselves contain a reproduction. Getty and GEMA reached opposite results on their facts; the difference was evidence of memorised, reproducible works. Controls: deduplication and memorisation measurement.
- Outputs. Whether the deployed system reproduces protected text or images on request. Controls: regurgitation testing, output filters, terms of use and a complaints channel.
The pipeline in one picture
The provenance ledger
The provenance ledger is the control everything else depends on. Without it you cannot answer the first question a rightsholder, regulator or court will ask, which is whether a given work was in a given model. Record, per document: source URL or supplier, acquisition time, licence or legal basis, the opt-out state observed at fetch time, a content hash, and the dataset versions it entered. Training runs then pin a manifest of dataset versions, so the mapping from work to model is a join, not an archaeology project. The wider governance patterns are covered in AI data governance and provenance and signing.
CREATE TABLE doc_provenance (
doc_id bytea PRIMARY KEY, -- sha256 of normalised content
source text NOT NULL, -- URL, supplier name or upload id
fetched_at timestamptz NOT NULL,
legal_basis text NOT NULL, -- 'licence:<id>', 'public-domain', 'tdm-exception', ...
robots_allowed boolean, -- NULL for non-web sources
tdm_reserved boolean, -- TDMRep or other machine-readable reservation seen
excluded_at timestamptz, -- set on takedown or later opt-out
exclusion_reason text
);
CREATE TABLE dataset_member (dataset_version text, doc_id bytea REFERENCES doc_provenance);
CREATE TABLE model_manifest (model_id text, dataset_version text);
Crawling: robots.txt, TDMRep and acquisition hygiene
Opt-outs are expressed in several machine-readable ways, and the GPAI Code of Practice commits signatories to honour robots.txt as specified in RFC 9309 and other widely used machine-readable reservations. The Code's copyright chapter also commits them not to circumvent technological protection measures such as paywalls, to exclude websites recognised as persistently and repeatedly infringing, and to tell rightsholders which crawlers they run. The W3C community group's TDM Reservation Protocol (TDMRep) is one such reservation: a tdm-reservation HTTP header or HTML meta element with value 1 or 0, an optional tdm-policy URL pointing at licensing terms, and a site-wide file at /.well-known/tdmrep.json. Per the specification, a meta element overrides the header, which overrides the well-known file; the sketch below reads only the header, so a production gate must check all three.
from urllib import robotparser
from urllib.parse import urlsplit
USER_AGENT = "examplebot-training" # publish this name and what it obeys
BLOCKED_HOSTS = load_blocklist() # placeholder: piracy and shadow-library domains
def admit(url, response, robots_cache):
"""Decide at fetch time and return the facts to store in the ledger."""
host = urlsplit(url).hostname
if host in BLOCKED_HOSTS:
return False, {"reason": "blocked-host"}
rp = robots_cache.get(host)
if rp is None:
rp = robotparser.RobotFileParser(f"https://{host}/robots.txt")
rp.read()
robots_cache[host] = rp
robots_ok = rp.can_fetch(USER_AGENT, url)
tdm = response.headers.get("tdm-reservation") # header only: also check the HTML
tdm_reserved = tdm is not None and tdm.strip() == "1" # meta element and tdmrep.json
admitted = robots_ok and not tdm_reserved and response.status_code == 200
return admitted, {"robots_allowed": robots_ok, "tdm_reserved": tdm_reserved}Three details matter. Record the decision even when you reject a page, because proving you respected an opt-out is as important as respecting it. Opt-outs change, so a page fetched in March may be reserved by September; re-check on every re-crawl and set excluded_at when a reservation appears, so future dataset builds drop it. And never fetch through login walls, paywalls or CAPTCHAs for training data, whatever the legal theory, because circumvention is exactly what the Code excludes and what makes acquisition look like Bartz's pirated library.
Memorisation, regurgitation and output filtering
GEMA turned on proof that the model reproduced lyrics, and the regurgitation risk is measurable. Published research has found that memorisation increases with model size, with the number of times a sequence is duplicated in the training data, and with the length of the prompt used to extract it, and that deduplicating training data reduces how much verbatim text models emit. Those findings give you three engineering levers: deduplicate aggressively, especially near-duplicates of popular works such as lyrics, poems and news articles that appear thousands of times on the web; measure extraction on the works you care most about; and filter outputs.
Measure by prompting the model with the first part of protected texts and comparing its continuation with the true continuation. Then deploy the same comparison as a filter: index hashed word shingles of the protected set and flag outputs that share long verbatim runs.
import hashlib, re
def shingles(text, k=12):
words = re.findall(r"\w+", text.lower())
for i in range(len(words) - k + 1):
yield hashlib.blake2b(" ".join(words[i:i + k]).encode(), digest_size=8).digest()
def build_index(docs, k=12): # docs: (doc_id, text) for protected works
index = {}
for doc_id, text in docs:
for h in shingles(text, k):
index.setdefault(h, doc_id)
return index
def verbatim_hits(output, index, k=12):
hits = {}
for h in shingles(output, k):
doc_id = index.get(h)
if doc_id is not None:
hits[doc_id] = hits.get(doc_id, 0) + 1
return hits # block or rewrite if any doc exceeds a thresholdA 12-word shingle shared with a protected work is rarely coincidence; a run of several consecutive shared shingles almost never is. Tune k and the threshold on a held-out sample so quotations of a sentence or two, common phrases and public-domain text do not trigger, and log every block so you can show the safeguard working. The same machinery doubles as a privacy control; membership-inference defence covers the related problem of proving whether a record was in training at all.
Worked example: a fine-tuned support assistant
Consider a team fine-tuning an open-weight model into a customer-support assistant for the EU market, using three sources: 40,000 licensed product manuals from the client, 2 million pages crawled from public technical forums and documentation sites, and a purchased dataset of how-to articles. The numbers below are illustrative.
- The licensed manuals enter the ledger with
legal_basis = 'licence:client-2026-14'and the licence's expiry date, so a build after expiry fails closed. - The crawl runs under a published user agent. Suppose 9 percent of hosts disallow it in robots.txt and 3 percent of pages send
tdm-reservation: 1; all are recorded and rejected, and 1.76 million pages are admitted. - The purchased dataset arrives without per-document sources. The team asks the vendor for provenance and a warranty, gets neither, and drops the dataset rather than inherit an unknown acquisition history.
- Near-duplicate removal over the admitted corpus removes 18 percent, mostly mirrored documentation. The training manifest pins the three dataset versions.
- Before release, an extraction probe over 500 sampled documents and 200 well-known copyrighted texts finds two long verbatim continuations of a popular manual mirrored across many sites. The team adds both to the shingle filter and removes the remaining duplicates for the next build.
- A month later a forum operator files a complaint. The ledger shows 1,912 pages from that host in dataset version 3, used by two models. They are excluded from version 4, added to the output filter immediately, and the next scheduled retrain drops them from the weights.
Failure modes
- No provenance. A corpus assembled from scraped dumps and third-party datasets with no per-document source cannot answer whether a work was used, cannot honour a takedown and cannot support the public training summary the AI Act requires.
- Opt-out checked once. Recording robots.txt state at first crawl and never again ignores reservations added later. Re-check on every crawl and propagate exclusions to the next dataset build.
- Laundered datasets. A third-party dataset that was itself built from a shadow library carries that history with it. Diligence on supplier acquisition is part of your acquisition.
- Weights cannot forget. Removing a document from storage does not remove it from a trained model. Plan for output filtering as the interim control and scheduled retraining as the durable one; machine unlearning methods are not yet a reliable substitute.
- Filters only in the chat interface. If the API, batch endpoints or downloadable weights bypass the output filter, the safeguard does not exist for those users.
- Treating one ruling as settled law. The decisions above conflict across jurisdictions and turn on their records. Design for the strictest market you serve.
Trade-offs
Strict acquisition rules shrink the corpus and can cost quality on niche domains; licensed data costs money and negotiation time. Aggressive deduplication improves both memorisation risk and, per published results, model quality, so it is the rare control with little downside. Output filters add latency and false positives on legitimate quotation, so tune them per product: a coding assistant, a lyrics service and a legal research tool need different thresholds. Keeping a ledger costs storage and discipline, but it is far cheaper than reconstructing provenance under a discovery order or a regulator's deadline.
What to do next
- Inventory every dataset in every model you ship and record, per dataset, its source, legal basis and acquisition method. Drop or quarantine anything with unknown provenance.
- Create the provenance ledger and model manifests above, and make dataset builds fail closed on documents without a recorded legal basis.
- Publish your crawler's user agent, honour robots.txt and TDMRep at fetch time, record every decision, re-check on re-crawl and never circumvent paywalls or logins.
- Deduplicate near-duplicates before training and run an extraction probe over high-risk works before each release.
- Deploy a verbatim-overlap output filter on every serving path and log its blocks.
- Stand up a rightsholder contact and complaint process that ends in ledger exclusions, filter updates and a retraining schedule.
- Review the table at the top of this article with counsel for each market you serve, and update it as appeals and new rulings arrive.