Every product that wants to train or fine-tune on user content eventually ships a toggle labelled something like ‘improve the model with my data’. The toggle is the easy part. The hard part is everything behind it: what the toggle legally means in each jurisdiction, whether consent is even the right legal basis, how the decision is recorded so you can prove it later, how a dataset build two months from now knows which conversations it may use, and what happens to a model when someone changes their mind.
This article treats consent as an engineering system rather than a banner. It corrects a common misconception, that each AI purpose needs its own consent, sets out a purpose taxonomy, designs an event-sourced consent ledger and the dataset filter that reads it, works through withdrawal and objection, and walks a chat product from toggle to training manifest. It is engineering guidance, not legal advice; your counsel decides the legal basis, and the system's job is to make whatever they decide enforceable and provable.
Consent is one lawful basis, not the only one
Under the GDPR, every processing purpose needs a lawful basis, and consent is only one of six in Article 6(1). The others that matter for AI are contract (processing needed to provide the service the user asked for) and legitimate interests. So the claim that training, fine-tuning, personalisation and analytics each ‘require distinct consent’ is wrong as stated. What is true is narrower: each purpose needs its own basis, and if the basis is consent, the consent must be specific to that purpose and not bundled with others or with acceptance of the terms of service (Article 7(2) and 7(4)).
For model development specifically, the European Data Protection Board's Opinion 28/2024 (December 2024) accepts that legitimate interest can be a basis, provided the controller passes the three-step test: the interest is lawful, clearly articulated and real; the processing is necessary for it with no less intrusive way; and the interest is not overridden by the data subjects' rights, taking into account what they could reasonably expect. When a controller relies on legitimate interest, the user's lever is the right to object under Article 21 rather than consent, and the flow must offer that objection clearly and honour it.
Consent still matters a great deal. It is usually required for special category data such as health information (Article 9), it is the safer basis when users would not expect their content to train a model, and many companies choose opt-in as a product decision regardless. The engineering point is that your system must represent which basis applies to which purpose for which user, because the user-facing control and the effect of turning it off both depend on it.
A purpose taxonomy
Start with a closed list of purposes, each with a stable identifier, a plain-language description and a versioned notice. Purposes must be specific enough that a user could predict what happens; ‘improving our services’ is not a purpose, it is a category.
| Purpose id | What it covers | Typical basis (decided by counsel) | Effect of withdraw or object |
|---|---|---|---|
| serve.inference | processing a prompt to answer it | contract | not applicable; needed to provide the service |
| train.foundation | pre-training or continued training of shared models | consent or legitimate interest | excluded from all future builds |
| train.finetune.shared | fine-tunes deployed to all customers | consent or legitimate interest | excluded from future fine-tunes |
| personalise.memory | per-user memory and preferences | consent or contract | memory deleted or disabled |
| eval.human_review | humans reading conversations for quality | consent or legitimate interest | removed from review queues |
| analytics.aggregate | aggregate usage statistics | legitimate interest | excluded from new aggregates |
Separating eval.human_review from training is deliberate: many users accept automated training but not a person reading their messages, and regulators look at the two differently. Keep the taxonomy small, because every purpose multiplies UI, policy rules and test cases.
Architecture: consent as an event log
The core design decision is to store consent as an append-only event log rather than a mutable boolean column. Article 7(1) requires the controller to be able to demonstrate consent, which means knowing what the user saw, when, and through which surface, not just the current value. An event log gives you that history, and it gives dataset builds something to pin: ‘this snapshot reflects every decision up to ledger offset N’. A mutable flag can only tell you the present, and the present is not what a training run from March used.
The consent record
Each event is a consent receipt (annotated JSON below). The fields are the minimum that lets you answer an auditor's or a user's question about any past decision:
{
"event_id": "01J9ZK4Q6V3T8M2N5R7W0XYABC",
"subject_id": "u_48213", // pseudonymous; resolves via the identity service
"purpose": "train.finetune.shared",
"decision": "granted", // granted | withdrawn | objected | denied
"basis": "consent", // which lawful basis this event relates to
"notice_version": "train-notice-2026-08",
"jurisdiction": "EU",
"surface": "settings_page", // settings_page | onboarding | api | gpc_signal
"actor": "subject", // subject | guardian | admin_on_behalf
"recorded_at": "2026-09-30T14:02:11Z",
"evidence": {"ui_build": "web-5.41.0", "locale": "de-DE"}
}Three rules keep the ledger trustworthy. Events are never updated or deleted for correction; a correction is a new event. The notice_version must reference text stored in a registry, so ‘what did they agree to’ has a literal answer. And defaults are explicit: if the policy for a jurisdiction is that absence of a decision means no, the policy engine encodes that, rather than each consumer inventing its own default. When a user exercises erasure, the ledger itself becomes a question for counsel, since some proof of the withdrawal is often retained under the obligation to demonstrate compliance; see erasure across stores and models.
Flows users can trust
The consent requirements translate directly into interface rules. Consent must be freely given, specific, informed and unambiguous, given by a clear affirmative act (Article 4(11)), and withdrawing must be as easy as giving (Article 7(3)). In practice:
- No pre-ticked boxes, and no consent implied by continuing to use the product.
- One control per purpose that uses consent, not one switch for everything.
- ‘No’ is as prominent and as close as ‘yes’; a bright accept button beside a grey text link is the pattern regulators cite.
- The withdrawal control lives where users look, in settings, reachable in the same number of steps as granting.
- Service does not degrade for refusing a purpose the service does not need; making the core product conditional on training consent undermines ‘freely given’.
- Notice text says what is used (prompts, files, feedback ratings), for what, and for how long.
California adds its own layer. The CCPA as amended by the CPRA gives consumers the right to opt out of the sale or sharing of personal information and to limit the use of sensitive personal information, requires businesses to honour opt-out preference signals such as Global Privacy Control, requires opt-in for selling or sharing the data of consumers under 16 (with parental consent under 13), and its regulations state that agreement obtained through dark patterns does not count as consent. A GPC signal arriving with a request should therefore produce a ledger event with surface: "gpc_signal", scoped to the purposes it legally covers; whether model training counts as sale or sharing depends on your arrangements, which the CCPA article works through.
Enforcing consent at dataset build
Consent is only real if the training pipeline enforces it. The rule is that the dataset builder never reads the current consent state; it reads a snapshot at a pinned ledger offset and writes that offset into the dataset manifest.
from dataclasses import dataclass
@dataclass(frozen=True)
class Snapshot:
offset: int
policy_hash: str
state: dict # (subject_id, purpose) -> last decision event at or before offset
def build_snapshot(ledger, offset, policy):
state = {}
for ev in ledger.read(upto=offset): # events in ledger order
state[(ev["subject_id"], ev["purpose"])] = ev
return Snapshot(offset, policy.hash(), state)
def allowed(snapshot, policy, subject, purpose):
ev = snapshot.state.get((subject.id, purpose))
rule = policy.rule(purpose, subject.jurisdiction) # basis + default for this market
if rule.basis == "consent":
return ev is not None and ev["decision"] == "granted"
if rule.basis == "legitimate_interest":
return ev is None or ev["decision"] == "granted"
return False # unknown basis fails closed
def filter_examples(examples, snapshot, policy, purpose, subjects):
kept, dropped = [], 0
for ex in examples:
if allowed(snapshot, policy, subjects[ex.subject_id], purpose):
kept.append(ex)
else:
dropped += 1
return kept, {"ledger_offset": snapshot.offset, "policy_hash": snapshot.policy_hash,
"purpose": purpose, "kept": len(kept), "dropped": dropped}The returned record goes into the dataset manifest and from there into the model card and training metadata, so for any model you can say exactly which consent state it was trained under. Two details matter. Examples without a resolvable subject (anonymous sessions, shared accounts) need an explicit policy rather than slipping through. And content about third parties inside a user's conversation is not covered by that user's consent at all, which is a reason to run PII filtering on training data regardless of the consent result.
Withdrawal, objection and trained models
Withdrawal of consent, or a successful objection, has a precise legal effect under the GDPR: it stops processing from that point, and Article 7(3) says it does not make processing that happened before the withdrawal unlawful. For engineering this splits into what is cheap and mandatory and what is expensive and debatable.
Mandatory and cheap: stop all future use. The withdrawal event must reach every place the subject's data waits to be used: fine-tune queues, human-review queues, feature stores, evaluation sets, cached shards for the next build. Implement this as a fan-out consumer of the ledger with a service-level objective, for example 72 hours from event to exclusion everywhere, and a reconciliation job that scans each store for subjects whose latest event is a withdrawal.
Expensive and debatable: models already trained. Whether a trained model contains personal data is a case-by-case question; the EDPB's opinion says a model is not automatically anonymous. Most teams handle this by excluding the subject from every future build and letting the existing model age out on its normal retraining cycle, and reserve retraining or unlearning for cases where the model can be shown to reproduce the person's data, which is the territory of deletion requests against trained models. Small, per-customer fine-tunes are the exception: when a handful of users dominate a dataset, retraining without them is cheap and usually the right call.
Worked example: a chat assistant
A chat assistant serves users in the EU and California. Counsel decides: inference under contract; shared fine-tuning under consent in the EU and under a clearly disclosed opt-out in California; human review under consent everywhere. The settings page shows two toggles, ‘Use my chats to improve models’ and ‘Allow reviewers to read flagged chats’, both off by default in the EU.
Over September, 1.2 million EU users are active and 18% grant fine-tune consent. In California, 400,000 users are active, 6% opt out through settings and a further 9% send a GPC signal that the policy treats as an opt-out for this purpose. On 1 October the dataset builder pins ledger offset 88,412,907 and keeps conversations from about 216,000 EU users and 340,000 California users. The manifest records the offset, the policy hash and the per-jurisdiction counts.
On 3 October a user who had granted consent withdraws it. The event lands at offset 88,503,114. The fan-out removes their conversations from the review queue within the hour and from the cached shards for the November build overnight. The October fine-tune, already trained, keeps running; its manifest shows it was built at an earlier offset, at which this user's consent was valid, which is exactly what Article 7(3) contemplates. The November build excludes them automatically because it pins a later offset.
Failure modes
- Mutable consent flag: no history, so you cannot show what any past model was trained under. Use an event log.
- Builder reads live state: two shards of the same dataset see different consent. Pin an offset.
- Bundled consent: one toggle for training, review and personalisation; regulators treat it as not specific. One control per consent-based purpose.
- Lost withdrawal: the event reached the database but not the feature store. Reconcile every store against the ledger on a schedule.
- Fail-open defaults: an unknown jurisdiction or missing subject is treated as allowed. Unknown must mean excluded.
- Notice drift: the notice text changed but old grants are treated as covering the new purpose. A material change needs fresh consent; compare notice versions in policy.
What to do next
- List every purpose your AI features use personal data for, and have counsel assign a basis per purpose and jurisdiction; record the result as policy data, not prose. GDPR for LLM applications covers the surrounding obligations.
- Replace any consent boolean with an append-only ledger and a versioned notice registry.
- Change the dataset builder to pin a ledger offset and write it, with the policy hash, into each manifest, as part of your training data documentation.
- Build the withdrawal fan-out with a stated SLO and a nightly reconciliation report.
- Audit the settings UI against the rules above: no pre-ticked boxes, symmetric yes and no, withdrawal in two clicks.
- Accept and log GPC signals, and test that an unknown jurisdiction fails closed.