India's Digital Personal Data Protection Act, 2023 (the DPDP Act) and the DPDP Rules, 2025 are the first general data-protection regime that applies to every LLM product serving users in India. The Act replaced years of drafts, including the Personal Data Protection Bill that older pages still call the PDPB, and it differs from GDPR in ways that matter for engineering: there is no general legitimate-interest ground, consent must be itemised per purpose, and penalties are fixed rupee ceilings per category of failure rather than a share of turnover.
This article reads the law as a set of requirements on data flows. You will see which obligations attach to prompts, retrieval indexes, logs and training sets, how to build a consent ledger and an erasure fan-out that actually reach every copy, and where LLMs make compliance harder than for an ordinary web app. It is engineering guidance, not legal advice; your counsel decides the interpretation, and you build the plumbing that makes it provable.
What the law covers and on what grounds
The Act applies to digital personal data processed in India, and to processing outside India when it is connected with offering goods or services to people in India. A US-hosted chatbot with Indian users is in scope. The vocabulary is its own: the person is the data principal, the company deciding purposes and means is the data fiduciary, and anyone processing on its behalf, such as a hosted model API, is a data processor. The fiduciary stays responsible for what its processors do, so a model vendor's retention settings are your compliance problem.
Processing needs one of two grounds. The first is consent under section 6: free, specific, informed, unconditional and unambiguous, given by a clear affirmative action, and limited to the data necessary for the stated purpose. Withdrawal must be as easy as giving it. The second is a closed list of certain legitimate uses in section 7, such as data a person voluntarily provides for a specified purpose without objecting, employment, medical emergencies and state functions. There is nothing like GDPR's open-ended legitimate interest, which is why "we may use your chats to improve our services" does not stretch as far in India.
Duties that shape the architecture
The duties that shape architecture are these. Notices must be understandable on their own and itemise the personal data and the specified purpose, with a way to withdraw, exercise rights and complain to the Data Protection Board of India. Fiduciaries must take reasonable security safeguards; the Rules list encryption, masking or virtual tokens, access control, logs with monitoring, backups for continuity and contractual terms with processors, and require those logs to be kept for at least one year. Data must be erased once the purpose is served or consent is withdrawn, unless a law requires retention, and the processor must erase too.
A personal data breach goes to each affected person without delay and to the Board without delay, with a detailed report within 72 hours. Children are anyone under 18: processing needs verifiable parental consent, and tracking, behavioural monitoring and targeted advertising aimed at children are prohibited. People have rights to a summary of their data and who it was shared with, to correction, completion and erasure, to grievance redressal and to nominate someone to act for them. Section 8(3) adds an accuracy duty when personal data is used to make a decision about the person or is disclosed to another fiduciary.
The government can notify a fiduciary as a Significant Data Fiduciary (SDF) based on volume, sensitivity and risk. SDFs need a data protection officer based in India, an independent data auditor, and an annual data protection impact assessment and audit. The Rules also make SDFs check that technical measures, including algorithmic software, are not likely to risk the rights of data principals; for a large LLM product that reads as a standing model-risk review. Cross-border transfer is open by default under section 16, with the government able to restrict named countries, and the Rules let it require certain data of SDFs to stay in India.
| Failure | Maximum penalty (Schedule) |
|---|---|
| Reasonable security safeguards not taken | Rs 250 crore |
| Breach not notified to Board or users | Rs 200 crore |
| Children's data obligations | Rs 200 crore |
| Additional SDF obligations | Rs 150 crore |
| Any other provision of the Act or Rules | Rs 50 crore |
Where LLM systems make it harder
LLM systems stress the Act in four places. Purpose drift. A user who consents to "answer my support question" has not consented to "fine-tune our model". Because consent is specific and itemised, reuse of prompts for training, evaluation or analytics needs its own purpose and its own grant, and data collected without it must never reach those pipelines. Copies everywhere. One prompt lands in the request log, the trace store, a vector index if memory is enabled, an evaluation sample, the provider's abuse-monitoring store and possibly a labelling queue. Erasure is only as good as your inventory of those copies.
Weights are not a database. Once personal data has shaped fine-tuned weights, you cannot delete one person's contribution with a query. Machine unlearning research exists, but it does not give an auditable guarantee today. The practical answer is upstream: keep personal data out of training sets unless there is a specific consent, record lineage from each training example to its consent record, and plan retraining windows so withdrawals take effect in the next model version. Generated statements about people. Section 3 excludes data that the person themselves made public, but that does not cover what others posted about them, and a model that confidently invents facts about a named person and feeds a decision runs into the accuracy duty. Ground such outputs in verified records or block them.
Building the consent ledger and erasure fan-out
Two components carry most of the load: a consent ledger that answers "may this data be used for this purpose now?", and an erasure fan-out that reaches every store. The ledger is append-only so you can prove what was true at the time of processing; section 6 puts the burden of proving that notice was given and consent obtained on the fiduciary.
# consent_ledger.py - append-only consent events; the current state is a fold over them
import datetime as dt
from dataclasses import dataclass
PURPOSES = {"support_answer", "chat_memory", "model_training", "product_analytics"}
@dataclass(frozen=True)
class ConsentEvent:
principal_id: str
purpose: str
action: str # "grant" or "withdraw"
notice_version: str # the exact notice text shown, by hash or version
at: dt.datetime
via: str # "app", "consent_manager:<id>", "guardian"
class ConsentLedger:
def __init__(self, store):
self.store = store # append-only table
def record(self, ev: ConsentEvent):
if ev.purpose not in PURPOSES:
raise ValueError(f"unknown purpose {ev.purpose}")
self.store.append(ev)
if ev.action == "withdraw":
enqueue_erasure(ev.principal_id, ev.purpose, ev.at)
def allowed(self, principal_id, purpose, at=None):
at = at or dt.datetime.now(dt.timezone.utc)
state = False
for ev in self.store.events(principal_id, purpose, until=at):
state = ev.action == "grant"
return state
def require_purpose(ledger, purpose):
"""Decorator for pipeline stages: no grant, no processing."""
def wrap(fn):
def inner(principal_id, *a, **kw):
if not ledger.allowed(principal_id, purpose):
raise PermissionError(f"{principal_id}: no consent for {purpose}")
return fn(principal_id, *a, **kw)
return inner
return wrapEvery write to a store then carries principal_id and purpose. The erasure worker reads a registry of stores, each with a delete handler and a deadline, and records completion per store so the job is resumable and auditable.
# erasure_fanout.py - one withdrawal, every copy
STORES = {
"chat_history": (delete_rows, "support_answer"),
"vector_memory": (delete_vectors, "chat_memory"), # filter by principal_id metadata
"trace_store": (redact_spans, "support_answer"), # keep timing, drop content
"eval_samples": (delete_rows, "model_training"),
"training_queue": (exclude_ids, "model_training"), # next fine-tune skips them
"provider": (call_vendor_delete_api, "support_answer"),
}
def run_erasure(job):
for name, (handler, purpose) in STORES.items():
if job.purpose not in (purpose, "*") or job.done(name):
continue
handler(job.principal_id)
job.mark_done(name) # evidence row: store, time, count
if job.all_done(STORES):
job.close()Security logs are the deliberate exception: the Rules require at least a year of them, so redact prompt content from logs at write time and keep identifiers and timing, rather than deleting the log line and losing your breach evidence.
Worked example: an LLM tutoring app
Take a Bengaluru-based tutoring app with an LLM tutor, 400,000 users, many of them school students. Walking the flows produces a short, concrete plan. Users under 18 are present, so the signup flow needs age declaration plus verifiable parental consent, and the product must turn off behavioural profiling and targeted advertising for those accounts. The tutor's chat memory is a separate purpose from answering questions, so it gets its own toggle, defaulting off for minors. Fine-tuning on student essays is a third purpose; the team decides not to ask for it at all and to train only on licensed and synthetic material, which removes the unlearning problem entirely.
The model runs on a hosted API, so the vendor contract must cover processing only on instructions, deletion on request and breach notice fast enough for the team to meet the 72-hour Board report. The retention table ends up as: chat history 12 months after last activity, memory until withdrawal, redacted security logs 13 months, and vendor-side retention set to the shortest the provider offers. The breach runbook gets a timer that starts when the on-call engineer confirms personal data was involved, with a pre-drafted Board report and user notice in English and Hindi. The notice itself is rewritten as three itemised purposes, each with its own grant.
Failure modes and trade-offs
Common failure modes are predictable. One checkbox for everything fails specificity and makes withdrawal all-or-nothing, so users cannot turn off training without losing the product. Untagged copies, usually traces, analytics exports and evaluation spreadsheets, survive erasure and surface in a breach. Vendor defaults keep prompts for abuse monitoring longer than your notice says. Training on everything first and asking later creates a model you cannot fix without retraining. Treating a breach clock as a ticket SLA misses the 72-hour report because nobody can query which principals were affected.
The trade-offs are real. Fine-grained purposes improve compliance and trust but add consent UX friction and lower opt-in rates for training data. Keeping personal data out of training costs model quality on domain-specific language, which you can partly recover with synthetic or licensed data. Short retention limits debugging; content-redacted traces with request IDs are the usual compromise. If you serve both the EU and India, build one consent and erasure system to the stricter of each rule rather than two parallel ones; GDPR for LLM systems and the right to erasure in LLM pipelines cover the EU side. For the rest of India's AI rulebook, read India AI policy in depth; for the leaks that trigger breach duties, PII leakage in LLMs and audit logging for LLM systems.