Data governance for AI means being able to answer, for any model you ship, three questions: what data was it trained on, were you allowed to use that data for this purpose, and what happens when that permission changes. Most organisations can answer none of them a year after training. The data was copied from a warehouse export, filtered by a notebook, deduplicated by a script, and the only lasting artefact is a checkpoint.

This page covers the training side: the data lifecycle that ends in a set of weights. The runtime side, the prompts, retrieval indexes, logs and caches an LLM application creates while serving, is covered in Data Governance for LLM Systems. Here you will build a source manifest, an admission gate, a lineage index that links records to model versions, and a policy for deletion requests against data that has already shaped a model.

Why training data needs its own governance

Training data is different from ordinary analytical data in one way that drives everything else: once used, it cannot be pulled back out. If a customer withdraws consent for a row in a warehouse, you delete the row. If that row was in a fine-tuning set, the influence of that row is spread across millions of parameters, and there is no reliable way to subtract it. Research on machine unlearning exists, but today's methods give approximate removal without strong guarantees, so a governance programme cannot rely on them as its only answer.

That makes governance a problem of control before training, plus traceability after. Control decides which records may enter a dataset, for which purposes. Traceability lets you find every dataset version and model a record touched, so that when something changes you know the blast radius. Both depend on metadata captured at the moment data enters the AI pipeline, because it is almost never recoverable later.

The four objects to govern

Govern four objects and the links between them. A source is where data comes from: a product database, a licensed corpus, a web crawl, a vendor feed, synthetic output from another model. It carries terms: licence, consent basis, allowed purposes, retention, jurisdiction. A record is one item, a document, a conversation, an image, with a stable ID and a pointer to its source. A dataset version is an immutable, content-addressed set of records with the transforms applied. A model version points to the dataset versions used to train and evaluate it.

Lineage chain: every model version can be traced back to the records and terms it was trained onSourcelicence, consent, purposeAdmission gatepolicy checks per recordDataset versioncontent-addressed manifestModel versiontrained on manifest hashingestadmittrainQuarantinerejected, with reasonrejectDeletion / objection requestsubject, source or licence changeLineage indexrecord, versions, modelsNext dataset version excludes itAffected models flaggedsuppress, retrain or documentGovernance metadata travels with the data; the model never becomes the only record of where it came from.
Records enter through an admission gate that checks source terms. Dataset versions are immutable manifests, models point at them, and a lineage index answers which models a record reached.

The important property is that metadata travels with the data. If a record's purpose restriction lives only in a wiki page about the source, the first export that drops the source column breaks governance silently. Store the source ID on every record and resolve terms from a registry at admission time.

Source manifests

Every source gets a manifest before any record is ingested. Keep it in version control and review it like code, with legal or privacy sign-off for anything carrying personal data or third-party content.

# sources/support_tickets_eu.yaml
id: src.support_tickets_eu
owner: support-platform@company
system_of_record: tickets_db.eu.tickets
contains_personal_data: true
jurisdictions: [EU]
legal_basis: legitimate_interest          # as assessed by privacy team, ref DPIA-2026-014
allowed_purposes: [support_assistant_finetune, eval]
prohibited_purposes: [marketing, foundation_pretraining]
licence: internal
retention_days: 730
deletion_feed: kafka://privacy.erasure.v1   # requests that must propagate
required_transforms: [pii_redaction_v3, signature_strip]
expires: 2027-06-30                        # re-review date

Four fields do most of the work. allowed_purposes stops data collected for support from quietly becoming pretraining data. deletion_feed connects the source to the requests that will later affect it. required_transforms forces redaction before admission rather than relying on someone to remember. expires forces a re-review, because terms change: a vendor contract lapses, a licence is updated, a consent form is reworded. For third-party and web data, record the licence and the date it was observed. Copyright and AI training covers why that matters.

The admission gate

The admission gate turns manifests into decisions per record. It runs in the pipeline that builds training sets, and it fails closed: a record with an unknown source, an expired manifest or a disallowed purpose is quarantined with a reason, never silently passed.

from dataclasses import dataclass

@dataclass
class Decision:
    admit: bool
    reason: str

def admit(record, purpose, registry, now):
    src = registry.get(record.source_id)
    if src is None:
        return Decision(False, "unknown_source")
    if now.date() > src.expires:
        return Decision(False, "manifest_expired")
    if purpose not in src.allowed_purposes or purpose in src.prohibited_purposes:
        return Decision(False, f"purpose_not_allowed:{purpose}")
    if record.subject_id and registry.is_suppressed(record.subject_id, now):
        return Decision(False, "subject_erased_or_objected")
    if record.created_at < now - src.retention:
        return Decision(False, "past_retention")
    missing = set(src.required_transforms) - set(record.transforms_applied)
    if missing:
        return Decision(False, f"missing_transforms:{sorted(missing)}")
    return Decision(True, "ok")

The suppression check is what makes deletion work going forward. When an erasure or objection request arrives, the subject goes into a suppression list, and every future dataset build excludes their records even if a stale copy survives somewhere upstream. Add the content checks your risk needs: PII scanners as a backstop for redaction, deduplication, and screening for poisoned or adversarial samples as described in Data poisoning. Report quarantine counts by reason; a sudden jump usually means an upstream schema change, not a sudden surge in bad data.

Dataset versions and lineage

A dataset version should be a manifest, not a folder. The manifest lists record IDs and content hashes, the gate version and the purpose it was built for, and its own hash is the dataset version ID. Training jobs take that ID as input and write it into the model's metadata. The lineage index is then a simple join you can query in both directions.

import hashlib, json

def build_manifest(records, purpose, gate_version):
    entries = sorted((r.id, r.content_sha256, r.source_id) for r in records)
    body = {"purpose": purpose, "gate_version": gate_version, "records": entries}
    blob = json.dumps(body, sort_keys=True).encode()
    return hashlib.sha256(blob).hexdigest(), body

# lineage tables (any warehouse works)
# dataset_records(dataset_id, record_id, source_id, subject_id)
# model_datasets(model_id, dataset_id, role)   -- role: train | eval

AFFECTED_MODELS = '''
SELECT DISTINCT md.model_id, md.role
FROM dataset_records dr
JOIN model_datasets md USING (dataset_id)
WHERE dr.subject_id = @subject_id OR dr.source_id = @source_id
'''

Publish a human-readable summary next to each dataset version, the datasheet: sources and their share of records, known gaps and biases, transforms, intended uses. Datasheets for datasets gives a template. The manifest is what machines check; the datasheet is what reviewers and auditors read.

Deletion and objection requests after training

When a request arrives, the lineage index tells you whether the subject or source reached any model. Then you have four responses, and a policy should say which applies when:

  1. Exclude going forward. Always. The subject goes on the suppression list and the next dataset version drops their records. This is cheap and should be automatic.
  2. Suppress at inference. If there is a risk the model reproduces the person's data, add output filters for their identifiers and test with targeted extraction prompts. This addresses the practical harm, though it does not change the weights.
  3. Retrain on schedule. Fine-tuned models are usually retrained regularly anyway. A policy such as a maximum age for any model trained on personal data means removed records age out within a known window. Write that window down and tell your privacy team what it is.
  4. Retrain or unlearn now. For high-risk cases, such as sensitive data that should never have been admitted, retrain from a corrected dataset version. Approximate unlearning may reduce influence, but treat it as a mitigation and record it, not as proof of erasure.

Which of these satisfies a legal obligation in your jurisdiction is a question for counsel, and it is still being argued about. What engineering can guarantee is the evidence: which models were affected, what was done, and when the data stopped influencing new versions. Licence changes follow the same path, keyed on source_id instead of subject_id.

What the regulations ask for

Several frameworks ask for exactly these artefacts. Under the GDPR, personal data used for training needs a lawful basis, is bound by purpose limitation, and is subject to rights such as erasure and objection; the source manifest and suppression list are how you show those were respected. The EU AI Act's Article 10 requires providers of high-risk AI systems to apply data governance and management practices to training, validation and testing data, covering design choices, collection processes and the origin of data, the original purpose for personal data, preparation steps such as labelling and cleaning, and examination for possible biases. For general-purpose AI models, Article 53 requires providers to put in place a copyright policy and to publish a sufficiently detailed summary of the content used for training; those obligations have applied since 2 August 2025. Application dates for other parts of the Act have been the subject of proposed changes, so check the current timeline for your system. The EU AI Act covers the risk tiers.

None of this needs a separate compliance system. A manifest per source, a gate with logged reasons, immutable dataset versions and a lineage index produce the evidence these rules ask for as a by-product of normal pipeline operation.

Worked example: a support assistant fine-tune

A company fine-tunes a support assistant on two years of EU support tickets. The source manifest above allows support_assistant_finetune, requires redaction, and connects to the erasure feed. The build for dataset version ds_9f3c sees 1.2 million tickets. The gate quarantines 41,000 past retention, 3,100 from suppressed subjects and 900 where the redaction transform was missing because a backfill job skipped it. The team fixes the backfill and rebuilds. Model support-ft-2026-09 records ds_9f3c as its training set.

Three weeks later a customer files an erasure request. The suppression list is updated within the hour. The lineage query shows the subject's 14 tickets are in ds_9f3c and therefore in the deployed model. Policy for this data class says: exclude going forward, add an inference filter for the customer's name and account number, and rely on the 90-day maximum model age. The privacy team receives an automated record: request ID, affected model, actions taken, and the date by which the retrained model replaces it. A month later a vendor contract for a translated knowledge base expires; the manifest's expires date has already blocked new builds from using it, and the same lineage query lists the two models that need a retrain before the old ones are retired.

Failure modes

Most governance failures are metadata that went missing between systems.

  • Source ID dropped in an export. A notebook flattens records to text and the connection to terms is lost. Make the gate reject any record without a source ID.
  • Purpose creep. A dataset built for evaluation is reused for training because it was convenient. Bind the purpose into the manifest hash and check it in the training job.
  • Synthetic data laundering. Outputs of a model trained on restricted data are used as "clean" synthetic data. Treat generated data as derived from its generator's sources and inherit their restrictions.
  • Evaluation sets ignored. Test data containing personal data is subject to the same rules and is often copied more widely than training data.
  • Mutable datasets. A folder that is edited in place makes lineage meaningless. Only immutable versions may be referenced by a training job.

Trade-offs

Strict admission throws away data and slows teams down; loose admission produces models you cannot defend later. A reasonable balance is to be strict on purpose, consent and licence, which are hard to fix after training, and pragmatic on quality checks, which are easy to iterate. Record-level lineage costs storage, but compared with the cost of retraining a model because you cannot prove what was in it, it is cheap. Short model lifetimes make deletion simpler at the cost of more frequent training runs.

What to do next

  • List every model in production and, for each, try to name the dataset version it was trained on; the gaps are your starting backlog.
  • Write a source manifest for each training source, with allowed purposes, legal basis, deletion feed and a re-review date.
  • Put an admission gate in the dataset build that fails closed and logs a reason for every rejection.
  • Make dataset versions immutable and content-addressed, and record their IDs in model metadata.
  • Build the lineage tables and test the affected-models query with a real deletion request.
  • Agree a written policy with privacy and legal for which deletion response applies to which data class, including a maximum model age.
  • Publish a datasheet with each dataset version and link it from the model card.
Key takeaway: Training data cannot be pulled back out of weights, so govern it before training and trace it after. Give every source a manifest with purposes and terms, admit records through a fail-closed gate, train only on immutable dataset versions, keep a record-to-model lineage index, and agree in advance how deletion and licence changes are handled for models that already exist.