Adding a large language model to a product usually creates several new processing operations at once. Users type personal data into prompts, documents about people are embedded into a vector index, prompts and outputs are logged for debugging, conversations are kept as memory, and some of it is sent to a model provider in another country. The General Data Protection Regulation (GDPR) applies to every one of those steps when the data relates to an identifiable person in its scope.

This article maps the regulation onto the architecture of an LLM application: who is responsible for what, which lawful basis you can rely on, how to minimise before the prompt, and how to honour access and erasure requests across every store the pipeline creates. It focuses on engineering obligations. It is not legal advice; decisions on lawful basis, transfers and risk acceptance belong with your data protection officer or counsel.

Advertisement

First principles: what the GDPR asks of any system

The regulation is built on a few principles in Article 5: lawfulness, fairness and transparency; purpose limitation; data minimisation; accuracy; storage limitation; integrity and confidentiality; and accountability, meaning you must be able to demonstrate compliance. Personal data is any information relating to an identified or identifiable person, which in an LLM system includes names in prompts, but also free text that identifies someone indirectly, user ids in logs, and embeddings derived from documents about people.

Everything else follows from those principles. Rights of the data subject, such as access (Article 15), rectification (Article 16), erasure (Article 17) and objection (Article 21), are the mechanisms by which people hold you to them. Requests must be answered without undue delay and in any event within one month, extendable by two further months for complex or numerous requests (Article 12(3)). Fines for infringing the principles can reach 20 million euros or 4 percent of worldwide annual turnover, whichever is higher (Article 83). The engineering consequence is simple to state and hard to do: you must know where every piece of personal data lives, why, and for how long.

Map the data before anything else

User / data subjectprompt, files, accountYour application (controller)minimise, route, logModel providerprocessor, retention, regionChat historysessions, memoryLogs and tracesprompts, outputsVector indexembedded documentsCaches, eval setsfine-tune dataData map: store, purpose, lawful basis, retention, subject keythe single source for access and erasureRights servicefan out access / erasure to every store, record proof
An LLM request fans personal data out into several stores. A data map keyed by a subject identifier is what makes access and erasure possible across all of them.

Draw the flow for one request. The user's prompt and any uploaded files reach your application. Some of that is stored as chat history or agent memory. Prompts and outputs land in logs and traces. Retrieved documents come from a vector index that may contain personal data about third parties. Responses may be cached, and interesting conversations get copied into evaluation or fine-tuning sets. The prompt is sent to a model provider, which may retain it for abuse monitoring for a period set in its terms.

For each store, record the purpose, the lawful basis, the retention period, the location and a subject key that lets you find a person's records. This is the backbone of your records of processing activities (Article 30), and it is the thing auditors and regulators ask for first. A store that has no subject key cannot honour access or erasure, and a store without a retention period violates storage limitation by default. The different memory layers an agent may keep are described in agent memory layers.

Advertisement

Roles: controller, processor, sub-processor

The controller decides the purposes and means of processing; the processor processes on the controller's behalf (Article 4). If you build a customer support assistant, you are normally the controller for your users' data, and a model provider you call through an API is normally your processor. That relationship requires a contract meeting Article 28: processing only on documented instructions, confidentiality, security, assistance with data subject requests, deletion or return at the end, and approval of sub-processors.

In practice, read the provider's data processing terms for three things: whether API inputs and outputs are used to train its models, how long they are retained and for what purpose, and in which regions processing happens. If a provider uses your customers' prompts for its own purposes, it is acting as a controller for that use, and you need a lawful basis and transparency for the disclosure. Enterprise and API terms often differ from consumer terms on exactly these points, so check the plan you actually use.

Lawful basis and purpose limitation

Article 6 lists the lawful bases. For LLM features the relevant ones are usually contract (the processing is necessary to deliver the service the user asked for), legitimate interests (which requires a documented balancing test and gives people a right to object) and consent (which must be freely given, specific and as easy to withdraw as to give). A chat reply the user requested fits contract; reusing those chats to fine-tune a model is a different purpose and needs its own basis, often legitimate interests with an opt-out, or consent.

Purpose limitation is where LLM projects most often slip. Logs kept for debugging quietly become an evaluation set, then a fine-tuning set. Each step is a new purpose that must be compatible with the original or have its own basis, and users must be told. Special categories of data in Article 9, such as health, religion or sexual orientation, are prohibited unless an explicit exception applies, and users type them into prompts all the time. Plan for their arrival rather than assuming they will not appear.

Minimise before the prompt, not after

Data minimisation (Article 5(1)(c)) and data protection by design and by default (Article 25) translate directly into code: decide which fields the purpose needs, send only those, and pseudonymise identifiers before text leaves your boundary. The allow-list belongs in your data protection impact assessment, not in a developer's head, and every call records what it sent so you can later prove it.

ALLOWED_FIELDS = {"ticket_text", "product", "order_status"}   # decided in the DPIA, not ad hoc

def build_prompt(ticket: dict, redactor, subject_id: str) -> tuple[str, dict]:
    # Data minimisation (Art. 5(1)(c)): only fields the purpose needs reach the model.
    context = {k: v for k, v in ticket.items() if k in ALLOWED_FIELDS}
    context["ticket_text"], vault = redactor.pseudonymise(context["ticket_text"])
    audit = {
        "subject_key": hash_subject(subject_id),     # lets erasure find this record later
        "purpose": "support_reply",
        "lawful_basis": "contract",
        "fields_sent": sorted(context),
        "provider_region": "eu",
        "retain_until": days_from_now(30),
    }
    return render_template("support_reply", context), audit

The detection and redaction techniques themselves, patterns, named-entity recognition, token vaults, and their honest recall limits, are covered in PII handling for LLMs. The point here is placement: minimisation happens before the model call and before logging, because anything that reaches a log or a provider must then be governed there too.

Data subject rights across every store

A request for access or erasure is only as good as the store you forgot. The table maps the common rights onto a typical pipeline.

StoreAccess (Art. 15)Erasure (Art. 17)Typical trap
Chat history, agent memoryExport conversations by subject keyDelete sessions and summariesSummaries and memories copied out of the session
Prompt and output logsSearch by subject keyDelete or shorten retentionLogs replicated to a vendor observability tool
Vector indexList chunks by source document and subjectDelete vectors and source text, re-indexStale copies in other index replicas or snapshots
Response cacheUsually not needed if short-livedExpire by key or short TTLCache keyed by prompt text, not by subject
Eval and fine-tune setsSearch by subject keyRemove and exclude from future runsWeights already trained on the data
Model providerVia processor assistanceVia provider deletion or retention expiryAbuse-monitoring retention you did not account for

Erasure has to fan out to every store, tolerate partial failure, respect legitimate exceptions such as legal holds (Article 17(3)), and produce a record that it happened without itself becoming a new copy of the personal data.

STORES = [chat_history, prompt_logs, vector_index, response_cache, eval_sets, provider_files]

def erase_subject(subject_id: str, request_id: str) -> dict:
    key = hash_subject(subject_id)
    report = {"request_id": request_id, "received": now(), "stores": {}}
    for store in STORES:
        try:
            n = store.delete_by_subject(key)            # every store must support this
            report["stores"][store.name] = {"deleted": n}
        except LegalHold as e:
            report["stores"][store.name] = {"held": e.reason}   # Art. 17(3) exception, recorded
        except Exception as e:
            report["stores"][store.name] = {"error": str(e)}
            enqueue_retry(store.name, key, request_id)
    for store in STORES:                                 # verify, do not assume
        entry = report["stores"][store.name]
        if "deleted" in entry:
            entry["verified"] = store.count_by_subject(key) == 0
    mark_training_exclusion(key)                         # keep future fine-tunes and evals clean
    return report                                        # kept as proof; contains no personal data

Rectification (Article 16) is harder than it looks: if a retrieved document says something false about a person, fix the source and re-index; do not rely on a prompt instruction to override it. Keeping the audit trail itself compliant is discussed in audit logging for LLM systems.

Model weights and training data

If you fine-tune on personal data, some of it can be memorised and later extracted. Whether a trained model itself contains personal data is not automatic in either direction: the European Data Protection Board's Opinion 28/2024 on AI models says anonymity of a model must be assessed case by case, based on how likely it is that personal data can be extracted or obtained through queries. Honouring erasure for data already in weights may require retraining, and machine unlearning is not yet a dependable substitute. The mechanics of memorisation, deduplication, differential privacy and canaries are covered in PII leakage and memorisation.

The most robust engineering answer is architectural: keep personal data in retrieval stores, where it can be deleted, rather than in weights, where it cannot easily be removed, and exclude flagged subjects from future training runs.

Automated decisions, transparency and Article 22

Article 22 gives people the right not to be subject to a decision based solely on automated processing, including profiling, that produces legal or similarly significant effects, subject to narrow exceptions and safeguards such as human intervention and the ability to contest. An LLM that drafts a reply is not making such a decision. An LLM that screens job applicants, sets credit limits or denies insurance claims without meaningful human review may be.

Meaningful review means the reviewer has the authority, time and information to disagree, not a rubber stamp on every output. Transparency obligations also apply: your privacy notice must say that an AI system processes the data, for what purposes, with which recipients, and where it is transferred.

International transfers

Sending personal data to a model provider outside the European Economic Area is a transfer under Chapter V (Articles 44 onward). It needs an adequacy decision for the destination, such as the EU-US Data Privacy Framework for certified US companies, or appropriate safeguards such as standard contractual clauses with a transfer impact assessment. Region-pinned endpoints offered by several providers and cloud platforms can keep processing in the EEA, but check whether abuse monitoring, support access or logging still leave the region. Record the chosen mechanism per provider in the data map.

DPIAs, records and breaches

A data protection impact assessment (Article 35) is required where processing is likely to result in high risk, and supervisory authorities list new technologies and large-scale processing among the indicators, so most customer-facing LLM features with personal data warrant one. A useful LLM DPIA describes the data flow above, the allow-listed fields, the provider terms, the retention per store, the prompt-injection and data-exfiltration risks, and the residual risk accepted.

LLM systems add breach scenarios that did not exist before: a prompt injection that makes the model reveal another user's data from shared memory, a retrieval index without per-user access control, or logs of prompts exposed through a misconfigured tracing tool. A personal data breach must be notified to the supervisory authority within 72 hours of becoming aware of it, unless it is unlikely to result in a risk (Article 33). Rehearse the LLM scenarios, and read data exfiltration through LLMs to design against them.

Worked example: an internal HR assistant

An HR team wants an assistant that answers employee policy questions and summarises case notes. The data map shows four stores: policies in a vector index (no personal data), case notes in a second index (special category data, including health), chat history and logs. The DPIA concludes that case-note summarisation is high risk. Decisions: case notes stay in an EEA-pinned deployment with no provider retention for training; the index enforces per-case access control so an HR partner retrieves only their own cases; logs keep metadata only, with prompts redacted and deleted after 14 days; the assistant never recommends disciplinary outcomes, keeping Article 22 out of scope; and erasure fans out to all four stores. The policy-question half launches first because its risk is low.

What to do next

  1. Draw the data flow for one request and list every store it touches, including vendor tools.
  2. For each store, record purpose, lawful basis, retention, location and subject key; add missing subject keys.
  3. Read your model provider's data processing terms for training use, retention and region, and record the transfer mechanism.
  4. Move minimisation and pseudonymisation before the prompt and before logging, driven by an allow-list.
  5. Build one erasure and access service that fans out to every store, verifies, and keeps proof without personal data.
  6. Separate purposes: do not let debugging logs become training data without a basis and notice.
  7. Run a DPIA for features touching special categories, decisions about people or large scale, and rehearse an LLM breach against the 72-hour clock.
Key takeaway: The GDPR applies to every step of an LLM pipeline that touches personal data: prompts, memory, logs, vector indexes, caches, evaluation and fine-tune sets, and the model provider. Map each store with purpose, lawful basis, retention, location and a subject key; treat the provider as a processor under an Article 28 contract and check its training, retention and transfer terms; minimise before the prompt; build one service that honours access and erasure everywhere; keep personal data in deletable retrieval stores rather than weights; and run a DPIA. This is engineering guidance, not legal advice.