Traditional data governance assumes data sits in tables with owners, schemas and access controls. An LLM application breaks that picture. A single customer email can be chunked into a vector index, pasted into a prompt, sent to a model vendor, echoed in a completion, written to a trace, cached, sampled into an evaluation set and, months later, used to fine-tune a model. Each copy has a different owner, a different retention period and a different way of leaking, and most were created by code nobody reviewed as a data flow.

This article sets out a governance design for that reality: an inventory of the stores an LLM system creates, labels that travel with data through chunking and generation, purpose binding, permission-aware retrieval, retention per store, deletion that reaches every copy, and a policy gate that enforces all of it. It is written for the engineers who build the pipelines, not only for the compliance team that audits them.

Advertisement

Inventory every store the system creates

You cannot govern stores you have not listed. Start with an inventory of every place your LLM system writes data, because the defaults of most frameworks are to keep everything.

StoreWhat it holdsTypical governance gap
Source connectorsCopies of wiki pages, tickets, drivesSource permissions not copied with the content
Chunk store and vector indexText chunks and their embeddingsTreated as derived and therefore harmless
Prompt and completion logsFull user input and model outputKept indefinitely for debugging
Traces and observabilityPrompts inside spans and attributesShipped to a third-party APM with its own retention
Response and semantic cachesAnswers keyed by similar questionsOne user's answer served to another
Agent and conversation memoryFacts extracted about usersNo owner, no expiry, no deletion path
Evaluation and fine-tuning setsSampled real conversationsCopied out of the governed system into notebooks
Model vendorWhatever you send in requestsRetention and training use set by contract, not code
Model weightsWhatever was in the training dataCannot be edited row by row

For each store, record the owner, the data classes it may hold, its retention, who can read it and how deletion works. Gaps in the last column are the work list. The embedding row deserves attention: embeddings are not anonymised data. Research on embedding inversion has reconstructed substantial parts of the original text from embeddings, so an index inherits the classification of the text it was built from.

Sourcesdocs, tickets, CRMUser promptschat, uploadsIngest gateclassify, label, lineagePolicy gatelabel x purpose x destVector indexchunks + labels + ACLModel APIvendor or self-hostedLogs and cachesTTL per storeEval / fine-tune setspurpose-bound copiesLineage indexsubject -> every copyraw datalabelled chunkspromptACL-filtered hitsallowed contextredacted recordonly if purpose allowsEvery arrow out of the policy gate is a decision that is logged and can be replayed for audit
Data governance control points in an LLM application: labels are attached at ingest, every flow passes a policy gate that checks label, purpose and destination, and a lineage index records every copy so deletion can reach all of them.

Labels that travel with the data

Governance needs every piece of data to carry machine-readable labels, attached at ingest and propagated through every transformation. A practical label set is small:

  • Classification: public, internal, confidential, restricted.
  • Data categories: personal data, special-category data such as health, credentials, payment data.
  • Subject and tenant: whose data this is, for deletion and isolation.
  • Purpose: the purposes it was collected for, such as support, product analytics or model improvement.
  • Source and lineage: the system and record it came from, and its version.

The propagation rule for derived data is the high-water mark: a chunk inherits the labels of its document, an embedding inherits the labels of its chunk, and a completion inherits the union of labels of everything in its context window, including retrieved chunks and the conversation. That rule is conservative, and it is the only one you can enforce without understanding the content. A summary of a restricted document is restricted, even if it happens to omit the sensitive sentence, because you cannot prove it did.

from dataclasses import dataclass, field

LEVELS = ["public", "internal", "confidential", "restricted"]

@dataclass(frozen=True)
class Labels:
    level: str = "public"
    categories: frozenset = frozenset()
    subjects: frozenset = frozenset()
    purposes: frozenset = frozenset({"any"})

def combine(*ls: Labels) -> Labels:
    # High-water mark for derived data such as completions.
    purposes = [l.purposes for l in ls if "any" not in l.purposes]
    return Labels(
        level=max((l.level for l in ls), key=LEVELS.index),
        categories=frozenset().union(*(l.categories for l in ls)),
        subjects=frozenset().union(*(l.subjects for l in ls)),
        purposes=frozenset.intersection(*purposes) if purposes else frozenset({"any"}),
    )

Purposes intersect rather than union: a completion built from support data and analytics data may only be used for purposes both allow. Detection of personal data at ingest will miss some cases, so labels from the source system, such as the CRM saying "this field is a customer record", are more reliable than scanners and should win when they disagree.

Advertisement

Purpose binding

The GDPR principles of purpose limitation and storage limitation, both in Article 5, map directly onto LLM pipelines. Data collected to answer a support question was not collected to train a model, and keeping full transcripts "in case they are useful for evaluation" is exactly the open-ended retention the principle forbids. Whatever your jurisdiction, the engineering form is the same: every flow has a declared purpose, and data may only move along a flow whose purpose is in its label.

The most common violation is quiet. An engineer exports a sample of production conversations into a notebook to build an evaluation set. The copy leaves the governed system, loses its labels and lives forever on a laptop or in a bucket. Make the governed path easier than the shortcut: provide an evaluation-set builder that samples only records whose purposes include evaluation, redacts by category, records lineage, and sets an expiry.

For systems that fall under the EU AI Act's high-risk category, Article 10, titled "Data and data governance", adds requirements for training, validation and testing data sets, covered in EU AI Act compliance. The inventory, labels and lineage described here are the evidence those obligations ask for.

Permission-aware retrieval

Retrieval is where governance most often fails in practice, because the index flattens the permissions of dozens of source systems into one similarity search. If a user cannot open a document in the source system, its chunks must not reach their prompt. A system prompt asking the model not to reveal restricted content is not a control; once the text is in the context window, the model can repeat it.

def retrieve(query_vec, user, k=8):
    principals = directory.expand(user)          # user, groups, roles, tenant
    hits = index.search(
        query_vec, k=k * 4,
        filter={"tenant": user.tenant,           # hard partition first
                "acl": {"any_of": principals},   # copied from the source system
                "level": {"lte": user.clearance}},
    )
    return [h for h in hits if acl_cache.still_allowed(user, h.doc_id)][:k]

Three details matter. Filter inside the search, not after it, or a query can return zero permitted results while leaking timing and counts. Copy access lists at ingest and resynchronise them when source permissions change, because a revoked share that lingers in the index is a breach. And partition by tenant before anything else, as described in multi-tenant LLM isolation; caches need the same partition key, or a semantic cache will serve one customer's answer to another.

Retention per store

Each store gets a retention period chosen from its purpose, enforced by the storage layer rather than by a cleanup script that will one day stop running. Illustrative choices for a support assistant might be: raw prompt and completion logs for a short debugging window, then deletion; redacted logs for longer for quality analysis; traces with prompt bodies stripped before export to the APM vendor; cache entries with a TTL of hours; agent memory per user with an explicit expiry and a user-visible list. The numbers are a policy decision; what matters is that every store has one and it is enforced with TTLs, lifecycle rules or partitions dropped by date.

Logs deserve special care because they are needed for security investigation, which argues for keeping them. Split them: keep the metadata (who, when, which model, which documents by ID, which policy decisions) for the audit period, and keep the content for much less time or in redacted form. The design of tamper-evident audit trails is covered in LLM audit logging.

Deletion that reaches every copy

An erasure request under GDPR Article 17, or a customer offboarding, has to reach every copy. That is impossible without a lineage index that maps a subject identifier to every record derived from their data: documents, chunk IDs, vector IDs, log entries, cache keys, memory items, evaluation records and the training runs that consumed any of them.

def erase_subject(subject_id, request_id):
    refs = lineage.find(subject_id)             # every derived copy, by store
    with audit.span("erasure", request=request_id, refs=len(refs)):
        chunk_store.delete(refs.chunks)
        index.delete(refs.vector_ids)           # then confirm compaction ran
        logs.redact(refs.log_ids)
        cache.purge(refs.cache_keys)
        memory.delete(refs.memory_ids)
        evalsets.drop_rows(refs.eval_rows)
        for run in refs.training_runs:          # weights cannot be edited
            retrain_queue.flag(run, reason=request_id)
    return lineage.verify_absent(subject_id)    # re-query every store

Vector databases often mark deletions and remove data during later compaction, so confirm when the bytes are actually gone. Fine-tuned weights are the hard case: there is no reliable way to remove one person's influence from a trained model, so the practical controls are keeping personal data out of fine-tuning sets in the first place, a retraining cadence that drops erased records, and an honest record of which models were trained on what. Memorization risk is covered in LLM PII leakage.

The model vendor boundary

Every request to a hosted model is a data transfer. Whether the vendor stores prompts, for how long, whether staff can review them and whether they are used for training are set by your contract and account configuration, and they differ between vendors and between product tiers. Record them per vendor in the inventory, verify them against current contract terms rather than marketing pages, and re-check when you change tiers. Then enforce the boundary in code: route restricted data only to destinations approved for it, which may mean a self-hosted model for some classes. The broader trust boundary is covered in third-party API security.

One policy gate for every flow

The pieces come together in one enforcement point that every flow calls: ingest, retrieval into a prompt, a model call, a log write, a copy into an evaluation set. It takes the data's labels, the declared purpose and the destination, and returns allow, allow with redaction, or deny, and it logs the decision.

DESTINATIONS = {
    "vendor_api":   {"max_level": "confidential", "deny": {"special_category", "credentials"}},
    "self_hosted":  {"max_level": "restricted",   "deny": {"credentials"}},
    "apm_traces":   {"max_level": "internal",     "deny": {"personal_data"}},
    "eval_set":     {"max_level": "confidential", "deny": {"special_category"}},
}

def decide(labels, purpose, dest):
    rule = DESTINATIONS[dest]
    if "any" not in labels.purposes and purpose not in labels.purposes:
        return "deny", "purpose not permitted"
    if LEVELS.index(labels.level) > LEVELS.index(rule["max_level"]):
        return "deny", "classification too high for destination"
    blocked = labels.categories & rule["deny"]
    if blocked:
        return "redact", sorted(blocked)
    return "allow", None

Keep the rules in version control, review changes like code, and test them with fixtures. Count denials per flow: a sudden rise usually means a new data source was labelled wrongly, and zero denials for months usually means a flow bypasses the gate.

Worked example: an erasure request

A support assistant answers customer questions from a help-centre index and the customer's own ticket history. A customer asks for erasure. Without lineage, the team deletes the CRM record and the tickets and declares success, while the customer's messages remain in 90 days of raw logs, an evaluation set exported in spring, a semantic cache, and the APM vendor's traces. With the design above, the lineage index returns 4 tickets, 37 chunks and vectors, 212 log entries, 3 cache keys, 6 evaluation rows and 1 fine-tuning run. Deletion runs per store, the vector store confirms compaction, the fine-tuning run is flagged for the next scheduled retrain, and APM traces never held content because the gate stripped prompts before export. The verification query returns nothing, and the audit record lists every step.

Failure modes

Failure modeSymptomControl
Labels dropped at chunkingRestricted text in a general indexChunker copies labels; ingest rejects unlabelled chunks
Post-filtering retrievalEmpty results, or leaks if the filter is skippedFilter inside the search
Stale ACLs in the indexRevoked users still retrieve contentPermission change events trigger resync
Untenanted cacheOne customer sees another's answerTenant and ACL in the cache key
Shadow evaluation setsProduction data in notebooks and bucketsGoverned builder with expiry and lineage
Prompt bodies in tracesPersonal data at an APM vendorGate strips content before export
Deletion stops at the sourceErased data still in logs and indexesLineage index and verification query

Trade-offs

High-water-mark labelling over-classifies: many completions will be labelled restricted when their text is harmless, which blocks useful analytics. Accept it, and add a reviewed declassification path rather than weakening the rule. Short retention makes debugging harder; metadata-only logs and on-demand capture for a specific session recover most of that value. Self-hosting restricted classes costs money and model quality. And a central policy gate is a dependency on every path, so it must be fast, highly available and fail closed for restricted data.

What to do next

  1. Build the store inventory, including traces, caches, memory, evaluation sets and every vendor, with an owner and retention for each.
  2. Define a small label set and make ingest reject unlabelled data; propagate labels through chunking, embedding and generation.
  3. Move retrieval permission checks inside the vector search and wire source permission changes to an index resync.
  4. Put a policy gate on model calls, log writes, trace export and evaluation-set creation, and log every decision.
  5. Set and enforce retention per store with TTLs or lifecycle rules, not cleanup scripts.
  6. Build the lineage index and run a test erasure end to end, including a verification query across every store.
  7. Record vendor retention and training terms from current contracts, and review them whenever you change vendor or tier.
  8. Read data lineage and contracts for AI knowledge systems to make source systems publish labels and changes in a contract rather than by convention.
Key takeaway: An LLM application copies data into indexes, prompts, logs, caches, memory, evaluation sets, vendors and weights. Inventory every store, label data at ingest and propagate labels by high-water mark, bind flows to purposes, enforce permissions inside retrieval, give each store an enforced retention period, keep a lineage index so erasure reaches every copy, and route every flow through a policy gate that logs its decisions.