Traditional data governance assumes data sits in tables with owners, schemas and access controls. An LLM application breaks that picture. A single customer email can be chunked into a vector index, pasted into a prompt, sent to a model vendor, echoed in a completion, written to a trace, cached, sampled into an evaluation set and, months later, used to fine-tune a model. Each copy has a different owner, a different retention period and a different way of leaking, and most were created by code nobody reviewed as a data flow.
This article sets out a governance design for that reality: an inventory of the stores an LLM system creates, labels that travel with data through chunking and generation, purpose binding, permission-aware retrieval, retention per store, deletion that reaches every copy, and a policy gate that enforces all of it. It is written for the engineers who build the pipelines, not only for the compliance team that audits them.
Inventory every store the system creates
You cannot govern stores you have not listed. Start with an inventory of every place your LLM system writes data, because the defaults of most frameworks are to keep everything.
| Store | What it holds | Typical governance gap |
|---|---|---|
| Source connectors | Copies of wiki pages, tickets, drives | Source permissions not copied with the content |
| Chunk store and vector index | Text chunks and their embeddings | Treated as derived and therefore harmless |
| Prompt and completion logs | Full user input and model output | Kept indefinitely for debugging |
| Traces and observability | Prompts inside spans and attributes | Shipped to a third-party APM with its own retention |
| Response and semantic caches | Answers keyed by similar questions | One user's answer served to another |
| Agent and conversation memory | Facts extracted about users | No owner, no expiry, no deletion path |
| Evaluation and fine-tuning sets | Sampled real conversations | Copied out of the governed system into notebooks |
| Model vendor | Whatever you send in requests | Retention and training use set by contract, not code |
| Model weights | Whatever was in the training data | Cannot be edited row by row |
For each store, record the owner, the data classes it may hold, its retention, who can read it and how deletion works. Gaps in the last column are the work list. The embedding row deserves attention: embeddings are not anonymised data. Research on embedding inversion has reconstructed substantial parts of the original text from embeddings, so an index inherits the classification of the text it was built from.
Labels that travel with the data
Governance needs every piece of data to carry machine-readable labels, attached at ingest and propagated through every transformation. A practical label set is small:
- Classification: public, internal, confidential, restricted.
- Data categories: personal data, special-category data such as health, credentials, payment data.
- Subject and tenant: whose data this is, for deletion and isolation.
- Purpose: the purposes it was collected for, such as support, product analytics or model improvement.
- Source and lineage: the system and record it came from, and its version.
The propagation rule for derived data is the high-water mark: a chunk inherits the labels of its document, an embedding inherits the labels of its chunk, and a completion inherits the union of labels of everything in its context window, including retrieved chunks and the conversation. That rule is conservative, and it is the only one you can enforce without understanding the content. A summary of a restricted document is restricted, even if it happens to omit the sensitive sentence, because you cannot prove it did.
from dataclasses import dataclass, field
LEVELS = ["public", "internal", "confidential", "restricted"]
@dataclass(frozen=True)
class Labels:
level: str = "public"
categories: frozenset = frozenset()
subjects: frozenset = frozenset()
purposes: frozenset = frozenset({"any"})
def combine(*ls: Labels) -> Labels:
# High-water mark for derived data such as completions.
purposes = [l.purposes for l in ls if "any" not in l.purposes]
return Labels(
level=max((l.level for l in ls), key=LEVELS.index),
categories=frozenset().union(*(l.categories for l in ls)),
subjects=frozenset().union(*(l.subjects for l in ls)),
purposes=frozenset.intersection(*purposes) if purposes else frozenset({"any"}),
)Purposes intersect rather than union: a completion built from support data and analytics data may only be used for purposes both allow. Detection of personal data at ingest will miss some cases, so labels from the source system, such as the CRM saying "this field is a customer record", are more reliable than scanners and should win when they disagree.
Purpose binding
The GDPR principles of purpose limitation and storage limitation, both in Article 5, map directly onto LLM pipelines. Data collected to answer a support question was not collected to train a model, and keeping full transcripts "in case they are useful for evaluation" is exactly the open-ended retention the principle forbids. Whatever your jurisdiction, the engineering form is the same: every flow has a declared purpose, and data may only move along a flow whose purpose is in its label.
The most common violation is quiet. An engineer exports a sample of production conversations into a notebook to build an evaluation set. The copy leaves the governed system, loses its labels and lives forever on a laptop or in a bucket. Make the governed path easier than the shortcut: provide an evaluation-set builder that samples only records whose purposes include evaluation, redacts by category, records lineage, and sets an expiry.
For systems that fall under the EU AI Act's high-risk category, Article 10, titled "Data and data governance", adds requirements for training, validation and testing data sets, covered in EU AI Act compliance. The inventory, labels and lineage described here are the evidence those obligations ask for.
Permission-aware retrieval
Retrieval is where governance most often fails in practice, because the index flattens the permissions of dozens of source systems into one similarity search. If a user cannot open a document in the source system, its chunks must not reach their prompt. A system prompt asking the model not to reveal restricted content is not a control; once the text is in the context window, the model can repeat it.
def retrieve(query_vec, user, k=8):
principals = directory.expand(user) # user, groups, roles, tenant
hits = index.search(
query_vec, k=k * 4,
filter={"tenant": user.tenant, # hard partition first
"acl": {"any_of": principals}, # copied from the source system
"level": {"lte": user.clearance}},
)
return [h for h in hits if acl_cache.still_allowed(user, h.doc_id)][:k]Three details matter. Filter inside the search, not after it, or a query can return zero permitted results while leaking timing and counts. Copy access lists at ingest and resynchronise them when source permissions change, because a revoked share that lingers in the index is a breach. And partition by tenant before anything else, as described in multi-tenant LLM isolation; caches need the same partition key, or a semantic cache will serve one customer's answer to another.
Retention per store
Each store gets a retention period chosen from its purpose, enforced by the storage layer rather than by a cleanup script that will one day stop running. Illustrative choices for a support assistant might be: raw prompt and completion logs for a short debugging window, then deletion; redacted logs for longer for quality analysis; traces with prompt bodies stripped before export to the APM vendor; cache entries with a TTL of hours; agent memory per user with an explicit expiry and a user-visible list. The numbers are a policy decision; what matters is that every store has one and it is enforced with TTLs, lifecycle rules or partitions dropped by date.
Logs deserve special care because they are needed for security investigation, which argues for keeping them. Split them: keep the metadata (who, when, which model, which documents by ID, which policy decisions) for the audit period, and keep the content for much less time or in redacted form. The design of tamper-evident audit trails is covered in LLM audit logging.
Deletion that reaches every copy
An erasure request under GDPR Article 17, or a customer offboarding, has to reach every copy. That is impossible without a lineage index that maps a subject identifier to every record derived from their data: documents, chunk IDs, vector IDs, log entries, cache keys, memory items, evaluation records and the training runs that consumed any of them.
def erase_subject(subject_id, request_id):
refs = lineage.find(subject_id) # every derived copy, by store
with audit.span("erasure", request=request_id, refs=len(refs)):
chunk_store.delete(refs.chunks)
index.delete(refs.vector_ids) # then confirm compaction ran
logs.redact(refs.log_ids)
cache.purge(refs.cache_keys)
memory.delete(refs.memory_ids)
evalsets.drop_rows(refs.eval_rows)
for run in refs.training_runs: # weights cannot be edited
retrain_queue.flag(run, reason=request_id)
return lineage.verify_absent(subject_id) # re-query every storeVector databases often mark deletions and remove data during later compaction, so confirm when the bytes are actually gone. Fine-tuned weights are the hard case: there is no reliable way to remove one person's influence from a trained model, so the practical controls are keeping personal data out of fine-tuning sets in the first place, a retraining cadence that drops erased records, and an honest record of which models were trained on what. Memorization risk is covered in LLM PII leakage.
The model vendor boundary
Every request to a hosted model is a data transfer. Whether the vendor stores prompts, for how long, whether staff can review them and whether they are used for training are set by your contract and account configuration, and they differ between vendors and between product tiers. Record them per vendor in the inventory, verify them against current contract terms rather than marketing pages, and re-check when you change tiers. Then enforce the boundary in code: route restricted data only to destinations approved for it, which may mean a self-hosted model for some classes. The broader trust boundary is covered in third-party API security.
One policy gate for every flow
The pieces come together in one enforcement point that every flow calls: ingest, retrieval into a prompt, a model call, a log write, a copy into an evaluation set. It takes the data's labels, the declared purpose and the destination, and returns allow, allow with redaction, or deny, and it logs the decision.
DESTINATIONS = {
"vendor_api": {"max_level": "confidential", "deny": {"special_category", "credentials"}},
"self_hosted": {"max_level": "restricted", "deny": {"credentials"}},
"apm_traces": {"max_level": "internal", "deny": {"personal_data"}},
"eval_set": {"max_level": "confidential", "deny": {"special_category"}},
}
def decide(labels, purpose, dest):
rule = DESTINATIONS[dest]
if "any" not in labels.purposes and purpose not in labels.purposes:
return "deny", "purpose not permitted"
if LEVELS.index(labels.level) > LEVELS.index(rule["max_level"]):
return "deny", "classification too high for destination"
blocked = labels.categories & rule["deny"]
if blocked:
return "redact", sorted(blocked)
return "allow", NoneKeep the rules in version control, review changes like code, and test them with fixtures. Count denials per flow: a sudden rise usually means a new data source was labelled wrongly, and zero denials for months usually means a flow bypasses the gate.
Worked example: an erasure request
A support assistant answers customer questions from a help-centre index and the customer's own ticket history. A customer asks for erasure. Without lineage, the team deletes the CRM record and the tickets and declares success, while the customer's messages remain in 90 days of raw logs, an evaluation set exported in spring, a semantic cache, and the APM vendor's traces. With the design above, the lineage index returns 4 tickets, 37 chunks and vectors, 212 log entries, 3 cache keys, 6 evaluation rows and 1 fine-tuning run. Deletion runs per store, the vector store confirms compaction, the fine-tuning run is flagged for the next scheduled retrain, and APM traces never held content because the gate stripped prompts before export. The verification query returns nothing, and the audit record lists every step.
Failure modes
| Failure mode | Symptom | Control |
|---|---|---|
| Labels dropped at chunking | Restricted text in a general index | Chunker copies labels; ingest rejects unlabelled chunks |
| Post-filtering retrieval | Empty results, or leaks if the filter is skipped | Filter inside the search |
| Stale ACLs in the index | Revoked users still retrieve content | Permission change events trigger resync |
| Untenanted cache | One customer sees another's answer | Tenant and ACL in the cache key |
| Shadow evaluation sets | Production data in notebooks and buckets | Governed builder with expiry and lineage |
| Prompt bodies in traces | Personal data at an APM vendor | Gate strips content before export |
| Deletion stops at the source | Erased data still in logs and indexes | Lineage index and verification query |
Trade-offs
High-water-mark labelling over-classifies: many completions will be labelled restricted when their text is harmless, which blocks useful analytics. Accept it, and add a reviewed declassification path rather than weakening the rule. Short retention makes debugging harder; metadata-only logs and on-demand capture for a specific session recover most of that value. Self-hosting restricted classes costs money and model quality. And a central policy gate is a dependency on every path, so it must be fast, highly available and fail closed for restricted data.
What to do next
- Build the store inventory, including traces, caches, memory, evaluation sets and every vendor, with an owner and retention for each.
- Define a small label set and make ingest reject unlabelled data; propagate labels through chunking, embedding and generation.
- Move retrieval permission checks inside the vector search and wire source permission changes to an index resync.
- Put a policy gate on model calls, log writes, trace export and evaluation-set creation, and log every decision.
- Set and enforce retention per store with TTLs or lifecycle rules, not cleanup scripts.
- Build the lineage index and run a test erasure end to end, including a verification query across every store.
- Record vendor retention and training terms from current contracts, and review them whenever you change vendor or tier.
- Read data lineage and contracts for AI knowledge systems to make source systems publish labels and changes in a contract rather than by convention.