China's Personal Information Protection Law (PIPL) took effect on 1 November 2021, and it reaches any LLM product that serves people in mainland China, wherever the servers sit. It is not a translation of GDPR. PIPL has no legitimate-interests basis. It asks for separate consent at several points where GDPR asks for none. And its cross-border rules can decide whether you may call a model hosted outside China at all. An LLM application leaks personal information into more places than a classic web app: prompts, logs, retrieval indexes, evaluation sets, fine-tuning data and, possibly, model weights. Each of those is a processing activity PIPL expects you to justify.

This article maps PIPL onto the architecture of an LLM system. It covers which flows exist, the lawful basis for each, where separate consent is required, what a personal information protection impact assessment (PIPIA) must cover, how to honour deletion across stores that were never designed for it, and how the 2024 cross-border rules change the decision to use an offshore model. It is an engineering guide, not legal advice; confirm conclusions with PRC counsel.

What PIPL covers and when it reaches you

PIPL covers any information relating to an identified or identifiable natural person. Anonymised information, which cannot identify anyone and cannot be restored, is outside the law. De-identified information is not, so pseudonymising user IDs in logs does not take them out of scope.

Under Article 3, processing outside China is covered when it provides products or services to individuals in China or analyses their behaviour. An overseas chatbot marketed to users in China is in scope, and Article 53 requires it to set up a dedicated entity or appoint a representative in China.

The personal information handler decides purpose and method, roughly a GDPR controller. A party processing on its behalf is an entrusted party (Article 21), bound by a contract fixing purpose, method, data types and retention. A model API provider is usually your entrusted party. If its terms let it train on your traffic, it becomes a handler in its own right, and your consent obligations change.

Map the data flows of an LLM product

Compliance starts with an inventory of flows, not of databases. For each flow, record the data, its purpose, where it is stored, how long it is kept, and whether it crosses the border. A typical retrieval-augmented assistant has at least six flows.

FlowPersonal information involvedTypical purposePIPL pressure point
Prompt and responseWhatever the user types or uploadsProvide the serviceSensitive PI arrives unannounced
Request logsPrompts, outputs, account and device IDsDebugging, abuse handlingRetention limit, Art 11 of the generative AI measures
Retrieval indexChunks of documents naming peopleGrounding answersDeletion must reach embeddings
Evaluation setsSampled real conversationsQuality measurementNew purpose, needs its own basis
Fine-tuning dataCurated conversationsModel improvementNew purpose; hard to delete afterwards
Offshore inferencePrompt sent to a model abroadCapability or costCross-border transfer rules
User in Chinaprompt, filesGatewayconsent check, PI detectminimisedDomestic modelno exportPrompt logsretention clockVector storeRAG chunksTuning setseparate purposeoutside ChinaOffshore modelArt 38-39 neededexport pathRights requests: fan out deletion to every store
PIPL view of an LLM product: every arrow is a processing activity needing a basis, and the red line is where Articles 38-39 begin to apply.

Lawful basis and separate consent, flow by flow

Article 13 lists the lawful bases: consent, necessity for a contract with the individual, HR management under lawful labour rules, statutory duties, public-health emergencies, news reporting in the public interest, and information the individual already made public, within a reasonable scope. There is no legitimate-interests basis, so analytics and model-improvement work needs consent or a convincing contract-necessity argument.

That argument has limits. Answering the prompt is necessary for the service. Using the conversation to fine-tune the next model is a different purpose, and Article 14 requires fresh consent when purpose changes. Article 16 forbids refusing service because someone declines consent for processing the service does not need, so a training checkbox that blocks sign-up fails.

PIPL also requires separate consent (单独同意) for providing personal information to another handler (Article 23), publicly disclosing it (Article 25), handling sensitive personal information (Article 29) and transferring it abroad (Article 39). Separate means its own affirmative act, not a clause in the privacy notice. Build consent as data: each grant has a scope, a notice version and a timestamp, and the gateway checks it before every flow it gates.

from dataclasses import dataclass
from datetime import datetime

@dataclass(frozen=True)
class Grant:
    user_id: str
    scope: str          # "service", "sensitive", "cross_border", "training", "third_party"
    notice_version: str
    granted_at: datetime
    withdrawn_at: datetime | None = None

def allowed(grants: list[Grant], user_id: str, scope: str) -> bool:
    return any(g.user_id == user_id and g.scope == scope and g.withdrawn_at is None
               for g in grants)

def route(req, grants):
    if not allowed(grants, req.user_id, "service"):
        raise PermissionError("no consent for core service")
    if req.contains_sensitive and not allowed(grants, req.user_id, "sensitive"):
        req = redact_sensitive(req)          # or ask for separate consent in-flow
    target = "offshore" if allowed(grants, req.user_id, "cross_border") else "domestic"
    keep_for_training = allowed(grants, req.user_id, "training")
    return target, req, keep_for_training

Sensitive personal information and minors

Article 28 defines sensitive personal information as information whose leak or misuse easily harms dignity, person or property. It names biometrics, religious beliefs, specific identities, medical health, financial accounts and location tracking, plus all personal information of minors under 14. Handling it requires a specific purpose, sufficient necessity and strict protection. Article 29 requires separate consent and Article 30 requires telling the person why it is necessary and how it affects them. For under-14s, Article 31 requires a parent or guardian's consent and dedicated handling rules.

Users paste lab results, bank statements and ID photos whether you ask or not. Detect sensitive categories at the gateway, redact before logging or ask for separate consent in the flow, and keep sensitive content out of secondary stores by default.

Automated decisions under Article 24

Article 24 governs automated decision-making. It must be transparent and fair, and it must not impose unreasonable differential treatment on individuals in transaction prices or conditions. Information pushes and marketing driven by automated decisions must offer an option not targeted at personal characteristics, or an easy way to refuse. Where an automated decision significantly affects someone's rights and interests, they may ask for an explanation and may refuse a decision made solely by automated means.

When an LLM screens loan applications, triages claims or ranks candidates, store the inputs, retrieved evidence, model version and prompt template behind each decision so an explanation can be reconstructed, and provide a human review route, because the right to refuse is only real if a non-automated path exists.

The impact assessment

Article 55 requires a personal information protection impact assessment before handling sensitive information, using it for automated decision-making, entrusting it to others, providing it to other handlers or making it public, transferring it abroad, or any other handling with a significant impact on individuals. A typical LLM launch triggers several of these at once. Article 56 says the assessment must cover whether purpose and method are lawful, legitimate and necessary, the impact and risks to individuals, and whether the protection measures are lawful, effective and proportionate. The report and handling records must be kept for at least three years.

Tie the PIPIA to the flow inventory: a new eval set, vendor or region should fail review until its entry exists.

Rights requests across logs, indexes and weights

Articles 44 to 47 give rights to know and decide, access and copy, portability where the regulator's conditions are met, correction and deletion. Article 47 requires deletion when the purpose is achieved, retention ends or consent is withdrawn, among other triggers. Where deletion is technically difficult, the handler must stop all processing except storage and necessary security measures. Article 50 requires a convenient request mechanism.

In an LLM system a deletion request has to reach every store in the flow map. Logs and evaluation sets are rows you can find by user ID. Retrieval indexes need chunk-level provenance, so that the vectors derived from a document can be found and removed. Fine-tuned weights are the hard case, and we know of no official guidance on how Article 47 applies to them. The defensible approach is avoidance: train only on data with training consent, record which data went into which model version, and plan retraining or retirement when that data has to go.

def erase_user(user_id, stores, lineage, audit):
    report = {}
    for store in stores:                      # logs, eval sets, vector index, caches
        report[store.name] = store.delete_where(user_id=user_id)
    for model in lineage.models_trained_on(user_id):
        report[model] = "flagged: stop further training use; schedule retrain or retire"
    audit.append({"user": user_id, "result": report})   # keep evidence of handling
    return report

Cross-border transfers and the 2024 thresholds

Article 38 allows export only through a security assessment by the Cyberspace Administration of China (CAC), certification, a contract on the CAC standard form, or another condition set by law. Article 39 requires telling the individual about the overseas recipient, with separate consent where consent is the basis. Article 40 requires critical information infrastructure operators (CIIOs), and handlers above a CAC-set volume, to store data in China and pass a security assessment before export.

The CAC's Provisions on Promoting and Regulating Cross-Border Data Flows, in force from 22 March 2024, relaxed this. For a handler that is not a CIIO, the volume bands count individuals from 1 January of the current year. A security assessment is needed when exporting non-sensitive personal information of 1 million or more individuals, or sensitive personal information of 10,000 or more. Between 100,000 and 1 million non-sensitive, or fewer than 10,000 sensitive, the standard contract or certification applies. Below 100,000 non-sensitive, none of the three is needed. Important data always needs a security assessment, and CIIOs exporting personal information generally do too; practitioners differ on whether the purpose exemptions below also cover CIIOs, so confirm with counsel. Some transfers are exempt from all three mechanisms regardless of volume, including those truly necessary to perform a contract with the individual, cross-border HR management under lawful rules, and emergencies protecting life, health or property. The Article 39 notice duty, and separate consent where consent is the basis, still apply to an exempt transfer.

def export_mechanism(is_ciio, important_data, non_sensitive_ytd, sensitive_ytd, exempt_purpose):
    # Counts are distinct individuals since 1 January of the current year.
    if important_data:
        return "security assessment"
    if is_ciio:
        return "security assessment (ask counsel whether an exemption applies)"
    if exempt_purpose:                 # contract necessity, cross-border HR, emergency
        return "exempt (Art 39 notice still applies)"
    if non_sensitive_ytd >= 1_000_000 or sensitive_ytd >= 10_000:
        return "security assessment"
    if non_sensitive_ytd >= 100_000 or sensitive_ytd >= 1:
        return "standard contract or certification"
    return "no mechanism needed (Art 39 notice still applies)"

Whether an offshore inference call is truly necessary for the contract is a legal judgement. A prompt routed abroad when a domestic model would also work is hard to call necessary.

The generative AI measures on top

Public-facing generative AI services in China are also governed by the Interim Measures for the Management of Generative Artificial Intelligence Services, in force since 15 August 2023. Article 7 requires lawful training data, with consent or another legal basis for personal information. Article 11 bars collecting unnecessary personal information, unlawfully retaining identifiable input and usage records, and unlawfully providing users' input to others, and requires prompt handling of access, copy, correction and deletion requests. It turns prompt-log retention into a compliance decision.

Worked example: a support assistant

Take a hypothetical support assistant for a retailer serving users in China. It answers order questions from a RAG index of order records. The original design sent every prompt to an offshore model, kept logs for a year and trained on all chats.

The flow map shows four problems. Order records carry addresses and phone numbers abroad. Training on chats is a new purpose with no consent. Users paste payment screenshots, which are sensitive. A year of identifiable logs is hard to justify under Article 11. The revised design routes traffic to a domestic model by default, offers the offshore model only after cross-border separate consent, and counts exported individuals against the year-to-date bands. Logs are redacted at the gateway and kept for 30 days. Training uses only chats with a training grant, and each model version records its training-set manifest. Every flow has a PIPIA entry, and the deletion job fans out to logs, the index, eval sets and the lineage registry.

Failure modes

  • Bundled consent. One checkbox for service, training and export fails the separate-consent rules and Article 16's ban on refusing service.
  • Pseudonymised means out of scope. De-identified data is still personal information. Only anonymised data is outside PIPL.
  • Vendor quietly becomes a handler. A model API whose terms allow training on your traffic changes the Article 21 and 23 analysis.
  • Deletion stops at the database. Embeddings, caches and eval exports keep the data after the row is gone.
  • Forgetting Article 53. An overseas provider serving China with no local representative breaches PIPL before any data moves.

Trade-offs

Domestic-only inference removes most cross-border work but may limit model choice. Offshore inference with separate consent keeps capability for consenting users, at the cost of a two-path architecture and volume tracking. Aggressive redaction protects users but can degrade answers that need the redacted detail. Short log retention reduces exposure but makes incident forensics harder. The stakes are real: Article 66 allows fines of up to RMB 50 million or 5% of the previous year's turnover for serious violations, plus personal fines for responsible managers.

What to do next

  1. Draw the flow map for your LLM product: prompts, logs, indexes, eval sets, training data and every offshore call.
  2. Assign an Article 13 basis to each flow and mark where separate consent is required.
  3. Enforce versioned consent at the gateway, and detect sensitive categories before logging.
  4. Write PIPIA entries for each flow and block launches that add a flow without one.
  5. Build the deletion fan-out with chunk provenance and a model-lineage registry.
  6. Count exported individuals from 1 January of the current year and alert before the 100,000 and 1 million bands.
  7. Review the design with PRC counsel. Related reading: GDPR for LLM applications, PII detection and redaction architecture, PII leakage through memorisation, LLM audit logging and the EU AI Act engineering guide.
Key takeaway: Under PIPL every LLM data flow needs its own lawful basis, and there is no legitimate-interests fallback. Build consent as versioned data, keep sensitive content out of secondary stores, make deletion reach indexes and model lineage, and decide offshore inference with the 2024 volume bands and separate consent in view.