Until recently, documenting a training set was good practice. It is now a legal duty in two large markets. California's AB 2013 requires developers of public generative AI systems to post a summary of their training data covering twelve specific points. The EU AI Act requires every provider of a general-purpose AI model to publish a summary of training content in a template set by the Commission, and to keep more detailed technical documentation for regulators and downstream providers. The two regimes ask overlapping but different questions, at different granularity, for overlapping but different sets of models.
This article explains what each law asks for as of October 2026, then shows the engineering approach that keeps the answers true: a single machine-readable training manifest per model version, generated from the pipeline that actually built the dataset, rendered into each required document and gated at release. It includes a field crosswalk, generator code, a worked example and a checklist. It is engineering guidance, not legal advice; your counsel decides scope and wording.
California AB 2013
AB 2013, the Generative Artificial Intelligence Training Data Transparency Act, adds section 3111 to the California Civil Code. A developer, defined as a person or organisation that designs, codes, produces or substantially modifies a generative AI system or service for public use, must post documentation on its website about the data used to train the system. It applies to systems released on or after 1 January 2022 and made available to Californians, whether paid or free. The first deadline was 1 January 2026, and the documentation must be posted again before each later release or substantial modification, which includes retraining or fine-tuning that materially changes functionality or performance.
The documentation is a high-level summary of the datasets, and the statute lists twelve items: the sources or owners of the datasets; how the datasets further the intended purpose of the system; the number of data points, which may be in general ranges; the types of data points; whether the datasets include data protected by copyright, trademark or patent, or are entirely public domain; whether they were purchased or licensed; whether they include personal information; whether they include aggregate consumer information; whether there was cleaning, processing or other modification, and its intended purpose; the time period of collection, noting if collection is ongoing; the dates the datasets were first used in development; and whether the system used or continuously uses synthetic data generation.
Exempt are systems whose sole purpose is security and integrity, systems for operating aircraft in national airspace, and systems developed for federal national security, military or defence purposes and made available only to a federal entity. The law is in force but contested: xAI sued the state, a federal district court denied a preliminary injunction on 4 March 2026, and an appeal was pending as of mid-2026. Check the current status before relying on any outcome either way.
The EU training-content summary
Article 53(1)(d) of the EU AI Act requires every general-purpose AI model provider, including providers of open-source models, to publish a sufficiently detailed summary of the content used to train the model, following a template from the AI Office. The Commission published that template with an explanatory notice on 24 July 2025. It has three parts: general information such as provider and model identification, modalities, approximate data size per modality and language coverage; a list of data sources broken into public datasets, licensed private data, unlicensed private data, web-crawled data, user data, synthetic data and other sources; and data processing aspects, namely how reserved rights such as text-and-data-mining opt-outs were respected and how illegal content was removed.
Web-crawled data carries the most work. Providers list the top 10 percent of domain names by size of content scraped, and small and medium-sized enterprises the top 5 percent or 1,000 domains, together with information about crawler operation and collection periods. The explanatory notice leaves open whether size means bytes, tokens or another measure, so record the metric you used. When a model is trained further after release, the summary is updated every six months or sooner for a material change. A party that modifies a model enough to become its provider summarises only the content used for the modification.
Timing matches the EU AI Act timeline: the duty has applied to models placed on the market since 2 August 2025, the Commission's enforcement powers including fines started on 2 August 2026, and models already on the market before August 2025 must comply by 2 August 2027. Separately, Article 53(1)(a) and (b) require non-public technical documentation for the AI Office and information for downstream providers, both of which describe training data in more detail than the public summary.
One manifest, many documents
The failure pattern these laws punish is a document written by hand once, in prose, that drifts from what the training pipeline actually did. The fix is structural. Make the dataset manifest the system of record: one row per source with provenance, licence, collection window and personal-data status, written when the source is admitted, plus aggregate statistics emitted by the pipeline at build time, plus crawler statistics per domain. Bind the manifest to a model version, render every public and private document from it, and fail the release when any required field is empty. The admission side of this is covered in data governance for AI, and datasheets for datasets covers the internal documentation and how to bind it to the bytes.
The manifest schema
The manifest needs only a few dozen fields, but they must be typed so a renderer can rely on them. A minimal schema:
from dataclasses import dataclass, field
from datetime import date
@dataclass
class Source:
name: str
owner: str # AB 2013 item 1
kind: str # public | licensed | unlicensed_private | crawl | user | synthetic
purpose: str # item 2: how it furthers the intended purpose
modality: str # text | image | audio | video | code
units: int # item 3, exact here; rendered as a range
unit_name: str # "tokens", "images", "documents"
data_types: list[str] # item 4
ip_status: str # item 5: copyrighted | public_domain | mixed | unknown
acquired: str # item 6: purchased | licensed | none
personal_info: bool # item 7
aggregate_consumer_info: bool # item 8
processing: list[str] # item 9: dedup, pii_redaction, toxicity_filter ...
collected_from: date # item 10
collected_to: date | None # None means collection is ongoing
first_used: date # item 11
languages: list[str] = field(default_factory=list)
opt_out_method: str = "" # EU: robots.txt, TDM reservation protocol ...
@dataclass
class Manifest:
model: str
version: str
eu_release: date | None
synthetic_generation: bool # item 12
sources: list[Source]
top_domains: list[tuple[str, int]] # (domain, bytes) for crawled data
size_metric: str = "bytes"
Crosswalk between the regimes
The crosswalk below is the core of the design: each legal requirement maps to fields the pipeline already produces. Where a cell is empty, the law needs something your pipeline does not record, which is the gap to close first.
| Requirement | Manifest fields | Rendering |
|---|---|---|
| AB 2013 sources or owners; EU list of sources | owner, kind, name | AB: list of owners; EU: grouped by kind |
| AB 2013 purpose | purpose | One sentence per source |
| AB 2013 number of data points; EU size per modality | units, unit_name, modality | Ranges, never exact counts |
| AB 2013 types of data points | data_types | Labels or characteristics |
| AB 2013 IP; EU reserved rights | ip_status, opt_out_method | AB: yes or no per source; EU: method narrative |
| AB 2013 purchased or licensed | acquired | Per source |
| AB 2013 personal and aggregate consumer information | personal_info, aggregate_consumer_info | Per source |
| AB 2013 cleaning; EU illegal content removal | processing | Steps and their purpose |
| AB 2013 collection period and first use | collected_from, collected_to, first_used | Dates, ongoing flagged |
| AB 2013 synthetic data; EU synthetic sources | synthetic_generation, kind | Yes or no plus description |
| EU top domains | top_domains, size_metric | Top 10 percent, or SME rule |
Generating the documents and gating releases
The generator renders each document and the gate refuses to emit anything when a field is missing. Rendering counts as ranges is deliberate: AB 2013 allows general ranges, and exact counts invite needless disputes when a later dedup pass changes them by a fraction of a percent.
import math
BANDS = [10**3, 10**6, 10**9, 10**12]
def as_range(n):
lo = max([b for b in BANDS if b <= n], default=0)
hi = min([b for b in BANDS if b > n], default=None)
return f"{lo:,} to {hi:,}" if hi else f"more than {lo:,}"
def gate(m):
errors = []
for s in m.sources:
for f in ("owner", "purpose", "ip_status", "acquired"):
if not getattr(s, f):
errors.append(f"{s.name}: missing {f}")
if s.kind == "crawl" and not s.opt_out_method:
errors.append(f"{s.name}: crawl source without opt-out handling")
if any(s.kind == "crawl" for s in m.sources) and not m.top_domains:
errors.append("crawled data but no domain statistics")
if errors:
raise RuntimeError("; ".join(errors))
def top_domains(m, sme=False):
ranked = sorted(m.top_domains, key=lambda d: -d[1])
k = math.ceil(len(ranked) * (0.05 if sme else 0.10))
if sme:
k = min(k, 1000) # confirm the SME rule with counsel before relying on it
return [d for d, _ in ranked[:k]]
def ab2013_rows(m):
gate(m)
for s in m.sources:
yield {"source": s.name, "owner": s.owner, "purpose": s.purpose,
"data points": f"{as_range(s.units)} {s.unit_name}",
"types": ", ".join(s.data_types), "ip": s.ip_status,
"purchased or licensed": s.acquired,
"personal information": s.personal_info,
"aggregate consumer information": s.aggregate_consumer_info,
"processing": ", ".join(s.processing) or "none",
"collected": f"{s.collected_from} to {s.collected_to or 'ongoing'}",
"first used": str(s.first_used)}The SME branch encodes one reading of the template; the comment is there because legal interpretation belongs in review, not hidden in a constant. Item 12, synthetic data generation, describes the system rather than a source, so render it once at the top of the page from synthetic_generation. Publish the rendered pages from CI with the manifest hash in the footer, so any statement can be traced to the data build it describes.
Worked example: a fine-tuned support assistant
A company fine-tunes an open-weight model on 40,000 licensed customer-support transcripts, 20,000 synthetic dialogues generated by a larger model, and a crawl of 1,200 public documentation sites, then ships a chat product to customers in California and the EU.
In California the company is a developer, because fine-tuning that materially changes the system is a substantial modification. It posts the twelve items for its own datasets: owners (the customers under contract, itself for synthetic data, the site operators for the crawl), count ranges (the generator above puts each of the first two sources in the 1,000 to 1,000,000 band), the licensed status, personal information present in transcripts before redaction, the redaction and dedup steps, collection windows, first-use dates, and yes to synthetic generation. Counsel commonly reads the base model's own training data as the base developer's documentation to publish; if yours agrees, link to it rather than restating it.
In the EU the analysis differs. The company is the provider of its chat system, and so carries that system's Article 50 transparency duties. For the model, it becomes a provider of a general-purpose model only if its modification is significant enough under the Commission's guidelines on GPAI providers, which use the compute spent on the modification relative to the original as the main indicator. A support fine-tune usually falls well below that line, so the Article 53 summary stays with the base model provider; the company still needs that provider's documentation for its own obligations. Record the compute estimate and the reasoning in the manifest, because the conclusion must be re-checked whenever training grows.
The crawled sites raise the same opt-out questions under both regimes. The pipeline honoured robots.txt and the TDM reservation protocol at crawl time, logged the result per domain, and stored bytes per domain, so the documentation is generated rather than reconstructed. The copyright and AI training article covers the acquisition controls behind those log lines.
Failure modes
- Hand-written summaries. A prose page drafted once drifts from the data after the first retrain, and nobody notices until a complaint quotes it.
- Reconstructing provenance after training. Licences, collection dates and opt-out status that were not recorded at admission are often unrecoverable.
- Exact counts. Publishing precise numbers creates contradictions each time dedup or filtering changes; use the ranges the law permits.
- Missing domain statistics. Crawlers that do not record bytes per domain cannot produce the EU domain list without re-crawling.
- Forgetting re-publication. AB 2013 attaches to each release and substantial modification; a pipeline that publishes once misses later versions.
- Overclaiming. Stating that no personal information is present, when the claim is really that it was filtered, is a falsifiable statement you may not be able to defend.
- Disclosing secrets by accident. Internal names, customer identities or vendor prices in manifest fields leak into public pages unless the renderers whitelist fields.
Trade-offs
| Choice | Benefit | Cost |
|---|---|---|
| Generated documents from a manifest | Stays true across retrains, auditable | Schema and pipeline work up front |
| Ranges instead of exact counts | Stable, permitted by AB 2013 | Less useful to researchers |
| Per-source rows | Answers rights-holder questions directly | Longer pages, more review |
| Grouped categories | Shorter, protects commercial detail | May fall short of sufficiently detailed in the EU |
| One global page | One review cycle | Must satisfy the stricter regime in every field |
What to do next
- List every model and system you ship to California or the EU, with release and modification dates.
- Decide with counsel which are covered by AB 2013 and which make you a GPAI provider in the EU.
- Define the manifest schema and fill it for the current release from admission records and pipeline logs.
- Add per-domain byte counts and opt-out results to the crawler if they are not already recorded.
- Write renderers for the AB 2013 page, the EU summary and the technical documentation, with field whitelists.
- Gate releases on a complete manifest and publish rendered pages from CI with the manifest hash.
- Schedule re-publication on every release and, for continued EU training, at least every six months.