Brazil's Lei Geral de Proteção de Dados (LGPD, Law 13.709/2018) governs personal data of people in Brazil, and LLM products process a great deal of it: prompts, uploaded documents, retrieved records, transcripts, logs and evaluation sets. LGPD is often described as Brazil's GDPR. The structure is similar, but the details that matter for engineering differ. There are ten legal bases instead of six; sensitive data has its own closed list of bases in Article 11 instead of a general basis plus an exception; the review right for automated decisions covers any solely automated decision affecting a person's interests but no longer requires a human reviewer; incident deadlines are counted in business days; and Brazil has its own standard contractual clauses for transfers abroad.

This builder's guide covers data flows, legal bases, rights across embeddings and fine-tunes, Article 20, transfers to model providers and the incident clock. It is engineering guidance, not legal advice. For the GDPR baseline, see GDPR and LLMs.

Scope and the regulator

LGPD applies to processing carried out in Brazil, to processing aimed at offering goods or services to individuals in Brazil, and to data collected in Brazil (Article 3). A US or European LLM product with Brazilian users is therefore in scope even with no servers in Brazil. Personal data is any information about an identified or identifiable natural person. Anonymized data is outside the law unless the anonymization can be reversed with reasonable effort (Article 12), and that exception is the lever for training data discussed below.

The regulator is the ANPD. In September 2025, Provisional Measure 1.317 turned the Autoridade Nacional de Proteção de Dados into the Agência Nacional de Proteção de Dados, a regulatory agency with functional, technical, decision-making, administrative and financial autonomy. Law 15.352 of 25 February 2026 converted the measure into law and gave the agency duties under the Digital Statute for Children and Adolescents (Law 15.211/2025).

The ANPD has acted on AI directly. In July 2024 it suspended Meta's use of Brazilian users' data for AI training, citing reliance on legitimate interest, weak transparency, obstacles to rights and children's data, and lifted the measure at the end of August 2024 after Meta signed a compliance plan.

Map the flows first

Personal data flows in an LLM product serving users in BrazilTitular (user)prompts, files, voiceLGPD gatewaybasis, CPF masking, purposeModel providerpossibly outside BrazilTransfer checkArt. 33 mechanismRAG indexembeddings + metadataLogs and tracesretention clockEval / fine-tune setsbasis re-checkedToken vaultkept in BrazilRights serviceArt. 18 requestsDecision logArt. 20 reviewIncident triage3 business daysmasked prompttokensdelete / correctdecisionsEvery green store is personal data with its own legal basis, retention period and deletion path.The purple services are how you answer data subjects, reviewers and the ANPD.
Personal data enters through prompts and files, crosses a gateway that records the legal basis and masks identifiers, may leave Brazil to reach a model provider, and lands in indexes, logs and evaluation sets. Rights requests, automated-decision reviews and incident triage must reach every one of those stores.

Every LGPD obligation attaches to a processing operation, so map each store and transfer first. The usual blind spots are:

  • Embeddings. Vectors of personal text are personal data; they link back to source rows and usually sit beside the text.
  • Traces and prompt logs. Observability tools capture full prompts by default, with long retention and broad access.
  • Caches. Unscoped semantic cache keys serve one user's completion to another.
  • Evaluation and fine-tuning sets. Copying conversations into them is a new purpose needing its own basis.
  • Provider-side retention. Inputs kept for abuse monitoring are processing by your operator (processor).

The LGPD calls the controller the controlador and the processor the operador, and it requires both to keep records of processing operations (Article 37). An LLM data map is that record.

Legal bases, flow by flow

Article 7 lists ten legal bases for ordinary personal data: consent, legal or regulatory obligation, public administration, research bodies, performance of a contract, exercise of rights in proceedings, protection of life, health protection by professionals, legitimate interest, and credit protection. Article 11 gives a narrower list for sensitive data (racial or ethnic origin, religion, political opinion, union membership, health, sex life, genetic and biometric data). That list has no legitimate interest and no credit protection, though it does include fraud prevention and security in identification and authentication. Assign one basis per flow, not per product:

FlowTypical basisEngineering consequence
Answering the user's own requestContract performanceProcess only what the request needs; no reuse without another basis
Abuse and safety monitoring of promptsLegitimate interest (unavailable once sensitive data is involved)Document a balancing test, minimise fields, short retention
Product analytics on conversationsLegitimate interestAggregate or pseudonymize; honour opposition
Fine-tuning or evaluation on conversationsConsent, or legitimate interest with strong safeguardsSeparate opt-out flag; exclude sensitive data and children; consider anonymization
Health assistant processing symptomsSpecific consent (Art. 11); health protection only for procedures by health professionals or servicesLegitimate interest is unavailable; segregate stores

Legitimate interest under Article 10 requires a concrete purpose, data strictly necessary for it, transparency, and respect for the data subject's legitimate expectations. The ANPD can ask for a data protection impact report (RIPD, Article 38) when processing relies on it. Write the balancing test before launch; the Meta case shows training on users' content under legitimate interest is what the regulator questions first. Consent must be free, informed, unambiguous and purpose-specific, and children's data must be processed in the child's best interest under Article 14, with specific consent from a parent or guardian where consent is the basis.

A gateway with CPF-aware masking

Put one gateway in front of every model and retrieval call. It attaches the flow's legal basis and purpose to the request, refuses calls without one, applies minimisation, masks direct identifiers, and records an audit event without the body. The most common Brazilian identifier in free text is the CPF, the individual taxpayer number. Masking it reliably needs its check digits: without them you either redact every 11-digit order number or, by requiring punctuation, miss CPFs typed without it. This listing validates the two mod-11 check digits before replacing a match with a keyed token whose mapping stays in a vault in Brazil:

import re, hmac, hashlib

CPF = re.compile(r"\b(\d{3})\.?(\d{3})\.?(\d{3})-?(\d{2})\b")

def cpf_valid(digits):
    """Check the two CPF check digits (mod-11 scheme); reject repeated digits."""
    if len(digits) != 11 or len(set(digits)) == 1:
        return False
    nums = [int(ch) for ch in digits]
    for n in (9, 10):
        total = sum(d * w for d, w in zip(nums[:n], range(n + 1, 1, -1)))
        check = total * 10 % 11 % 10
        if check != nums[n]:
            return False
    return True

def mask_cpfs(text, key, vault):
    """Replace valid CPFs with keyed tokens; keep the mapping in a vault in Brazil."""
    def swap(m):
        digits = "".join(m.groups())
        if not cpf_valid(digits):
            return m.group(0)          # not a CPF: leave order numbers alone
        token = "CPF_" + hmac.new(key, digits.encode(), hashlib.sha256).hexdigest()[:12]
        vault[token] = digits
        return token
    return CPF.sub(swap, text)

# mask_cpfs("Cliente 529.982.247-25 pediu revisao; pedido 123.456.789-00.", key, vault)
# -> "Cliente CPF_<12 hex> pediu revisao; pedido 123.456.789-00."

The listing was run against 5,000 generated CPFs and their one-digit corruptions. Valid numbers were masked and corrupted ones were left alone. In the example, 123.456.789-00 has the shape of a CPF but fails the check, so it stays. Masked text is still personal data: the vault can reverse it and the surrounding text still identifies people, so masking reduces exposure but does not anonymize. Anyone holding the HMAC key can confirm a guessed CPF, so keep the key with the vault. And regexes miss identifiers in images, audio and odd spellings, so measure the miss rate on real samples. General detection techniques are in LLM PII leakage.

Data subject rights across an LLM stack

Article 18 gives data subjects (titulares) rights to confirmation of processing, access, correction, anonymization, blocking or deletion of unnecessary or unlawful data, portability, deletion of data processed on consent, information about sharing, information about the consequences of refusing consent, and revocation of consent. Under Article 19, confirmation and access must be provided immediately in simplified form, or within 15 days as a complete statement. The hard part is reach:

  1. Key everything by subject. Store a subject identifier on every prompt log row, vector, cache entry and evaluation record at write time. You cannot delete what you cannot find, and searching free text for a person afterwards is unreliable.
  2. Delete through indexes. Deleting a source document must delete its chunks and vectors, and should trigger index compaction where the vector store retains tombstoned data.
  3. Decide your fine-tune policy in advance. Removing one person's influence from trained weights is not practical with current methods. Keep identifiable data out of training sets, or keep a documented retraining cadence after which deleted subjects' data is gone, and say which one you do in your privacy notice.

Article 20: automated decisions

Article 20 gives data subjects the right to request review of decisions made solely on the basis of automated processing that affect their interests, including profiling decisions about personal, professional, consumer or credit matters. On request, the controller must give clear information about the criteria and procedures used, subject to commercial and industrial secrets, and the ANPD may audit for discriminatory aspects when that information is withheld. The original text said the review had to be carried out by a natural person. Law 13.853/2019 removed that requirement, so the statute does not guarantee human review. Many teams still provide it, because it is the most defensible way to meet the review right.

When an LLM ranks applicants, triages claims or flags accounts, log each decision so it can be reviewed and explained:

{
  "decisionId": "claim-2026-10-04-118734",
  "subjectRef": "subj_7f3a",
  "outcome": "REFER_TO_FRAUD_TEAM",
  "automatedOnly": true,
  "model": {"provider": "<vendor>", "name": "<model id>", "promptVersion": "triage-v14"},
  "inputsUsed": ["claim_amount", "policy_age_days", "prior_claims_12m", "free_text_summary"],
  "criteria": "risk score >= 0.8 from rubric triage-v14",
  "retrievedDocs": ["policy:PX-2231#s4"],
  "reviewChannel": "https://example.com.br/revisao",
  "retainUntil": "2031-10-04"
}

Keep decision criteria in a versioned rubric outside the prompt, so you can state them without exposing the prompt, and never let a free-text summary be the only record of a decision.

International transfers to model providers

Most model APIs run outside Brazil, so sending a prompt is usually an international transfer under Article 33. Permitted mechanisms include adequacy decisions by the ANPD, standard contractual clauses, specific contractual clauses, global corporate rules, and some narrower cases. ANPD Resolution CD/ANPD 19/2024 (August 2024) regulated transfers and approved Brazil's own standard contractual clauses. Organisations that rely on contractual clauses had until August 2025 to incorporate the ANPD's clauses, so EU SCCs alone do not cover Brazilian data.

  • Record each provider, its processing regions, and the mechanism used, including the provider's own sub-processors.
  • Prefer provider regions in Brazil where offered and adequate, and keep the token vault and raw logs in Brazil in any case.
  • Disable provider-side retention and training on your inputs where the contract allows, and record the setting as evidence.

Incidents and sanctions

Under Resolution CD/ANPD 15/2024, a security incident that may cause relevant risk or damage to data subjects must be reported to the ANPD and to the affected data subjects within three business days of the controller learning that personal data were affected. Unlike GDPR's 72 hours to the authority, the clock counts business days, and the same deadline covers notice to data subjects, which GDPR requires only for high-risk breaches and without a fixed deadline. Rehearse LLM-specific incidents: a cache serving one user's completion to another, injection that exfiltrates retrieved records, an exposed trace store. Scoping affected subjects quickly works only if the subject keys above exist.

Sanctions under Article 52 range from warnings and publicizing the infraction to blocking or deleting the data concerned, suspending processing, and fines of up to 2% of the group's revenue in Brazil in the prior year, capped at R$50 million per infraction.

The AI bill on the horizon

Brazil's AI bill, PL 2338/2023, passed the Senate in December 2024. It would add risk classification for AI systems and rights to explanation and human review, with the data protection authority coordinating oversight. As of a search on 2026-10-04, it was still awaiting a committee opinion in the Chamber of Deputies and is not law. Design for LGPD now; the decision log above would also serve the bill. For how a risk-based AI law works in practice, compare EU AI Act compliance.

Failure modes

  • One basis for the whole product. Training, analytics and serving have different purposes. Mitigation: a basis per flow, enforced by the gateway.
  • Calling masking anonymization. Tokenized prompts are still personal data. Mitigation: treat every store as personal data unless anonymization has been assessed against Article 12.
  • Unkeyed stores. Rights requests and incident scoping fail. Mitigation: subject references written at ingestion.
  • Sensitive data under legitimate interest. Health or biometric content in prompts makes this basis unavailable. Mitigation: detect and route sensitive flows separately.
  • Silent new transfers. A fallback model in another region. Mitigation: route allow-lists with a recorded transfer mechanism per route.

What to do next

  1. Build the data map: every prompt, file, vector, cache, trace, evaluation and provider-retention store, each with operator, region and retention.
  2. Assign and document a legal basis per flow; write the legitimate-interest balancing tests and isolate sensitive-data flows.
  3. Put a gateway in front of every model and retrieval call that records basis and purpose, masks CPFs and other identifiers, and logs without bodies.
  4. Add subject references to every store and test an Article 18 deletion end to end, including vector indexes.
  5. Inventory automated decisions, log them as above, and publish a review channel.
  6. Check each provider route for an Article 33 mechanism and update contracts to the ANPD clauses.
  7. Add LLM scenarios to your incident runbook with the three-business-day clock, and rehearse one.
  8. Track PL 2338/2023 and ANPD guidance on AI; see PII in LLM systems for detection depth.
Key takeaway: LGPD looks like GDPR but differs where it matters for LLM engineering: ten legal bases, a separate closed list of bases for sensitive data, a broad review right for automated decisions that no longer requires a human, Brazil-specific contractual clauses for transfers abroad, and a three-business-day incident clock. Map every store your LLM stack creates, assign a basis per flow, route all model calls through a gateway that masks identifiers such as CPFs, key every store by data subject so rights and incidents can be handled, and log automated decisions so they can be reviewed.