The California Privacy Rights Act is not a separate law you comply with alongside the CCPA. It is a 2020 ballot measure, Proposition 24, that rewrote the CCPA, and most of its changes became operative on 1 January 2023. The statute is still cited as the CCPA. Its additions are why an LLM privacy review now asks harder engineering questions: can a consumer correct what your assistant believes about them, can they stop you using their email contents beyond the task they asked for, and is everything you keep proportionate to a disclosed purpose.
This article covers those additions only, from the point of view of the team that builds and runs the system. The baseline obligations (notices, access, deletion, the sale opt-out, service provider classification of model vendors) are covered in CCPA and LLM Applications, and the way California differs from other states is in US State Privacy Laws Patchwork. This is engineering guidance, not legal advice; your counsel decides how the statute applies to you.
What CPRA changed, and what it did not
Start by separating what changed from what did not. The coverage test is the same shape: a for-profit entity doing business in California that meets one of three thresholds. CPRA raised the consumer count threshold to 100,000 consumers or households, added sharing to the 50 percent revenue test, and made the revenue figure adjust for inflation every odd-numbered year. For 2025 the agency set it at $26,625,000 and set administrative fines at up to $2,663 per violation and $7,988 per intentional violation or violation involving known minors under 16. Check the agency's current figures before you quote them, because they move every two years.
| CPRA addition | Statute section | What it means for an LLM system |
|---|---|---|
| Right to correct | 1798.106 | Inaccurate facts in profiles, memories and derived records must be fixable and stay fixed |
| Right to limit sensitive personal information | 1798.121 | A per-consumer flag that narrows which purposes may read sensitive fields |
| Sharing for cross-context behavioral advertising | 1798.120, 1798.140 | Opt-out covers ad sharing even when no money changes hands |
| Minimization and purpose limitation | 1798.100(c) | Collection, use and retention must be reasonably necessary and proportionate |
| Retention disclosure | 1798.100(a)(3) | Notice must state how long each category is kept, or the criteria |
| Contract terms for processors | 1798.100(d) | Written terms with service providers, contractors and third parties |
| A dedicated agency | 1798.199.10 onward | The California Privacy Protection Agency writes rules and enforces administratively |
CPRA also removed the automatic 30-day cure period, so fixing an issue after a regulator's letter no longer erases it, and extended the breach private right of action to an email address plus password.
Sensitive personal information in an LLM pipeline
Sensitive personal information (SPI) is a defined list, and LLM products touch more of it than their teams expect. It includes government identifiers, account login plus password, financial account plus access credentials, precise geolocation (within a radius of 1,850 feet), racial or ethnic origin, religious or philosophical beliefs, union membership, genetic data, biometrics processed to identify someone, health information, sex life and sexual orientation, and, after later amendments, citizenship or immigration status and neural data. One entry deserves a slow read: the contents of a consumer's mail, email and text messages, unless the business is the intended recipient of the communication.
That entry catches assistants: an email triage agent or a summarizer reading forwarded threads processes message contents where your company is not the recipient, and a health question typed into a chatbot is health information. Classification therefore has to happen on content, at ingestion, not on which form field the data came from, and the tags have to travel with the data into every store the request touches.
Inferences count too: if the assistant concludes from a run of questions that someone is managing diabetes and writes that conclusion to memory, it now holds health information.
The right to limit as a purpose gate
The right to limit lets a consumer restrict your use of their SPI to a set of permitted purposes. The statute and the agency's regulations list what remains allowed: performing the service the consumer reasonably expects when requesting it, security and integrity, short-term transient use, maintaining the quality and safety of the service, and a few others. The right applies only where SPI is collected or processed for the purpose of inferring characteristics about the consumer; if you never use it that way you may not need the link, but you need to be able to show that. In practice the clean design is to model this as a purpose gate that every read of a tagged field passes through.
from dataclasses import dataclass
SPI = {"health", "precise_geo", "message_contents", "biometric", "gov_id",
"religion", "ethnicity", "sexual_orientation", "union", "immigration", "neural"}
# Purposes the business may keep using even after a consumer limits SPI.
# Map your internal purposes onto the regulation's list with counsel; keep the list short.
PERMITTED_WHEN_LIMITED = {"fulfil_request", "security", "transient_display", "service_quality"}
@dataclass(frozen=True)
class Field:
consumer_id: str
category: str # set by the content tagger at ingestion
value: str
def may_read(f: Field, purpose: str, prefs) -> bool:
if f.category not in SPI:
return True
if prefs.limited_spi(f.consumer_id):
return purpose in PERMITTED_WHEN_LIMITED
return purpose in prefs.disclosed_purposes(f.category) # still bounded by your notice
def build_context(fields, purpose, prefs, audit):
allowed = []
for f in fields:
ok = may_read(f, purpose, prefs)
audit.write(f.consumer_id, f.category, purpose, ok) # evidence for the regulator
if ok:
allowed.append(f)
return allowedThe gate keys on purpose, so the same email can enter the prompt that drafts a requested reply yet be refused by a nightly profiling or fine-tuning job. Preferences are read per request, so a limit set at 10:00 binds the 10:01 batch, and every decision is logged, because a log of refusals is the convincing answer when a regulator asks whether limited data reached training.
Correction when facts live in derived stores
Correction is the CPRA right most foreign to LLM systems. Deletion removes a row; correction asks you to assert a new truth and keep it true. The regulations ask a business to consider the totality of the circumstances, including documentation the consumer supplies, and to use commercially reasonable efforts to correct. They also allow deleting the contested information instead, where that does not harm the consumer or the consumer agrees. They expect service providers to be told to correct as well.
A wrong fact can live in the profile, a memory store, conversation summaries, their embeddings and analytics tables. Most are derived, so the next sync or summarization pass recreates the error. The fix is a correction ledger that sits between sources and derived stores and is re-applied on every rebuild.
def apply_corrections(record, ledger):
"""Called by every job that writes a derived store (profile, memory, index)."""
for corr in ledger.for_consumer(record.consumer_id):
if corr.field in record and record[corr.field] == corr.old_value:
if corr.action == "replace":
record[corr.field] = corr.new_value
elif corr.action == "suppress": # deletion chosen instead of correction
del record[corr.field]
return record
def handle_correction(req, stores, ledger, providers):
ledger.add(req.consumer_id, req.field, req.old_value, req.new_value, req.action)
for store in stores: # memory, profile, vector index, caches
store.reprocess(req.consumer_id, transform=lambda r: apply_corrections(r, ledger))
for p in providers: # processors holding copies
p.notify_correction(req.consumer_id, req.field)Model weights are the difficult case. A fine-tune trained on customer records may reproduce an old fact. The practical position most teams take is to keep personal data out of weights in the first place, correct the retrieval and memory layers that feed the prompt (which dominate what the user sees), and add an output check for corrected fields. Erasure from weights is discussed at length in GDPR Right to Erasure Applied to LLMs; the same techniques apply when correction is the request.
Minimization, purpose limitation and retention
Section 1798.100(c) says collection, use, retention and sharing must be reasonably necessary and proportionate to the purpose for which data was collected, or to another disclosed compatible purpose. For an LLM product this turns into four concrete questions an engineer can answer: does the prompt include fields the task does not need, are full transcripts kept when a summary would serve, is the retention period for each category written down and enforced, and is a new purpose (training, evaluation, product analytics) being added to old data without a notice update.
The notice must state how long each category is kept, and that statement is testable. Make the retention schedule a table of category, store and time-to-live that both the notice generator and the deletion job read. Include the less obvious stores: prompt logs at the model gateway, traces in the observability backend, evaluation datasets copied from production, and the vector index, whose chunks often outlive the document they came from. Redaction before logging, covered in LLM PII Detection and Redaction Architecture, shrinks the problem at the source.
Contracts with model and tool vendors
CPRA requires a written contract with every service provider, contractor and third party you hand personal information to, and it prescribes content: the limited, specified purposes; a ban on selling or sharing; a ban on combining with other data except as the rules permit; an obligation to provide the same level of protection; notice if they can no longer comply; and your right to take reasonable steps to stop unauthorized use. For an AI stack the list of counterparties is long: model API providers, embedding services, vector database hosts, observability vendors receiving traces, labeling firms, and any agent tool that receives user data.
The engineering duty is to make the contract inventory match the data flow; a flow you cannot back with a signed agreement is what the agency cited in its first order (below). Keep a register per counterparty recording the contract, default training use, retention and allowed purposes, and generate it from gateway configuration so a new model endpoint cannot ship without a row.
Worked example: an inbox assistant
Take an inbox assistant sold to consumers. It reads a user's mailbox, drafts replies, and keeps a memory of preferences. Walk one message through it. A forwarded thread from the user's doctor arrives. The tagger marks it as message contents (the company is not the recipient) and as health. The user asks for a reply draft: purpose is fulfil_request, so the gate passes both fields into the prompt. The nightly job that mines threads for product-interest signals asks with purpose personalization_analytics; if the user has used the limit link, the gate refuses and logs it, and if they have not, the job is still bounded by what the notice disclosed.
A week later the assistant names the wrong clinic because a memory summary captured an outdated address. The user files a correction; the ledger records it, the memory store and vector index are reprocessed, and the gateway provider is notified. At the next weekly re-summarization the ledger is applied again, so the old address does not return. Raw message contents leave the cache after 7 days and summaries after 12 months, and the notice says so because it is generated from the same table.
Enforcement and the evidence trail
The agency's first administrative order, announced in March 2025, fined American Honda $632,500. None of the findings involved exotic technology: the company asked for too much identity verification on requests, such as opt-outs, that do not require it; its cookie tool made declining harder than accepting; and it could not produce contracts with the advertising technology vendors that received personal information. The order also required changes to business practices. The lesson for AI teams is that enforcement starts from ordinary, checkable artifacts: request flows, choice symmetry and contracts.
Regulations approved in September 2025 took effect on 1 January 2026 with staggered deadlines: risk assessment attestations by 1 April 2028, cybersecurity audits from 2028 to 2030 by revenue, and automated decisionmaking rules from 1 January 2027 (see the CCPA article above). The logs from your purpose gate and correction ledger are the raw material for those filings.
Failure modes
- Tagging by field, not content. The form says 'message', so nobody tags the health details inside it, and the analytics job reads them freely.
- Correction that does not stick. The profile is fixed, but the memory summarizer re-derives the old value from last year's conversations.
- Shadow copies. Evaluation sets and traces copied from production carry SPI with no retention entry and no gate.
- Contract drift. A team adds a new embedding vendor through a config change; the data flows before any agreement exists.
- Over-verification. Demanding identity documents for a limit or opt-out request, which the rules say should not require verification.
Trade-offs
A purpose gate costs latency and engineering time and weakens personalization for users who limit; never using SPI for inference is simpler but rules out features. Classifier tagging over-tags and under-tags; over-tagging is the safer error. Keeping personal data out of weights makes correction tractable at the price of retrieval infrastructure. One California-grade control set everywhere is usually cheaper than per-state branches.
What to do next
- List every store an LLM request writes to, including gateway logs, traces, memories, vector indexes and evaluation sets.
- Add content-based SPI tagging at ingestion, with message contents and health as the first two categories.
- Put a purpose-keyed gate in front of every read of tagged fields, read preferences per request, and log every decision.
- Build a correction ledger and make every derived-store rebuild apply it; test that a correction survives a full re-summarization.
- Turn the retention schedule into a table that drives both the deletion job and the privacy notice.
- Reconcile the vendor contract register against the gateway configuration and block endpoints that have no row.
- Review request flows for symmetry and for verification you do not need, the two issues behind the first agency fine.
- Diary the 2027 and 2028 regulatory dates and decide which team owns the evidence for each.