A chatbot that answers order questions looks like one feature. Under the California Consumer Privacy Act it is a set of data flows: the user's message, the account data pulled in to answer it, the retrieved documents, the model provider's API, the prompt log, the evaluation set someone exported last quarter, and perhaps a fine-tuned model. Each flow raises its own questions: who receives the data, for what purpose, for how long, and whether a consumer can see, correct or delete it.

This article maps the CCPA, as amended by the CPRA and later bills and regulations, onto an LLM application's architecture: what to inventory, which vendor relationships change your obligations, how to fan a deletion request out across stores nobody designed for deletion, and what the 2026 automated-decisionmaking rules add. GDPR is covered in GDPR for LLM systems. This is not legal advice; use it to prepare questions for counsel and to build systems that can answer them.

The data flows in one picture

Where personal information flows in an LLM applicationConsumerchat, voice, formsApp backendGPC, consent, routingPrompt assemblyaccount data, RAGLLM providerservice provider?APIVector storeembedded documentsPrompt and output logsretention clockAnalytics, ad tagssale or sharing?Eval and training setscopies of logsexportFine-tuned weightsAB 1008 questiontrainRequest handlerknow, delete, correct
Personal information enters through the chat, is enriched during prompt assembly, crosses to a model provider, and is copied into logs, vector stores, evaluation sets and possibly weights. The request handler must reach every one of them.

Who is covered, and what counts

The CCPA applies to a for-profit business that does business in California and meets one of three tests: annual gross revenue above a threshold that is adjusted for inflation every two years, $26,625,000 in the adjustment that took effect on 1 January 2025; buying, selling or sharing the personal information of 100,000 or more consumers or households a year; or deriving at least half of annual revenue from selling or sharing personal information. Consumers are California residents, and since 2023 that includes employees and business contacts, which matters for internal copilots as much as customer chatbots.

Personal information is broad: anything that identifies, relates to, describes or could reasonably be linked to a consumer or household, including inferences used to build a profile. A model's guess that a user is pregnant or in debt is personal information even if the user never typed it. AB 1008, signed on 28 September 2024, added that personal information can exist in "abstract digital formats", including artificial intelligence systems capable of outputting personal information. How far that reaches into model weights is contested, but it removes the easy argument that a trained model is outside the law by definition.

Sensitive personal information is a narrower list: government identifiers, account credentials, precise geolocation, racial or ethnic origin, religious beliefs, union membership, genetic and biometric data used for identification, health, sex life or sexual orientation, neural data after a 2024 amendment, and the contents of mail, email and text messages unless the business is the intended recipient. One reading is that a message to your own chatbot is addressed to you and so not automatically sensitive, but what the user writes can still be health information. Treat that reading as a question for counsel, not a safe harbour.

Inventory by store, not by feature

Every CCPA obligation starts with knowing where the data is. LLM products scatter copies in places a traditional data map misses, so inventory by store, not by feature.

StoreTypical personal informationUsual gap
Request and response logsraw prompts, outputs, user IDs, IP addresseskept indefinitely for debugging; no deletion path
Conversation memorysummaries and facts the assistant remembersinferences nobody disclosed in the notice at collection
Vector storeembedded tickets, emails, CRM noteschunks keep no link back to the consumer, so deletion misses them
Evaluation and red-team setsreal conversations copied for testingcopies outlive the source log's retention
Fine-tuning data and weightsexamples drawn from logsno lineage from example to model version
Vendor sideprovider logs, abuse-monitoring copiesretention set by the vendor's terms, not yours
Analytics and ad tags on chat pagespage events, sometimes message textmay count as sharing for cross-context behavioural advertising

For each row record categories, purpose, retention and recipients; the privacy notice must disclose retention by category. A store you cannot list is a store a request cannot reach.

Model providers: service provider or third party

When you send a prompt containing personal information to a model provider, the CCPA question is what that provider is to you. A service provider or contractor processes personal information on your behalf for a business purpose under a written contract. Disclosure to it is not a sale. The contract must limit its use to the specified purposes and forbid it from selling or sharing the data, retaining or using it outside the direct relationship, or combining it with data from other sources except as the regulations allow. A recipient that can use the data for its own purposes, for example to train general models offered to other customers, starts to look like a third party. Disclosure to a third party for value can be a sale, which triggers opt-out rights.

In practice: use business terms stating inputs are not used for training, check retention for abuse monitoring, and file the contract with your data map. Apply the same test to embedding APIs, vector databases, observability tools that capture prompts, and labelling vendors. An advertising pixel on the chat page is the classic way an AI feature turns into "sharing" without anyone deciding it should.

Consumer requests across the stack

Consumers can ask to know, delete and correct their personal information, opt out of sale or sharing, and limit the use of sensitive personal information. Under the regulations a business confirms receipt of a request to know or delete within 10 business days and responds within 45 calendar days, extendable once by another 45 with notice. Opt-out requests must be honoured as soon as feasible and within 15 business days. Deletion must also be passed on to service providers and contractors, and to third parties the data was sold or shared with, unless that proves impossible or involves disproportionate effort.

So build an orchestrator that fans out to every store and records proof. Each store's handler finds records by a stable consumer key, which must be written into logs, chunks and examples at ingestion; retrofitting it is the expensive part.

# request_fanout.py - one consumer request, every store, with receipts.
import datetime as dt
import json

class Store:
    name = "base"
    def find(self, consumer_id): raise NotImplementedError
    def delete(self, consumer_id): raise NotImplementedError

class PromptLogs(Store):
    name = "prompt_logs"
    def __init__(self, db): self.db = db
    def find(self, cid):
        return self.db.query("SELECT id, ts, prompt, output FROM llm_log WHERE consumer_id = %s", cid)
    def delete(self, cid):
        return self.db.execute("DELETE FROM llm_log WHERE consumer_id = %s", cid)

class VectorChunks(Store):
    name = "vector_store"
    def __init__(self, index): self.index = index
    def find(self, cid):
        return self.index.query(filter={"consumer_id": cid}, include_metadata=True)
    def delete(self, cid):
        return self.index.delete(filter={"consumer_id": cid})

def handle(request_type, consumer_id, stores, vendors):
    receipt = {"type": request_type, "consumer": consumer_id,
               "received": dt.datetime.utcnow().isoformat(), "stores": {}}
    for s in stores:
        try:
            if request_type == "know":
                receipt["stores"][s.name] = {"records": len(s.find(consumer_id))}
            elif request_type == "delete":
                receipt["stores"][s.name] = {"deleted": s.delete(consumer_id)}
        except Exception as e:                      # a failed store must be visible, not skipped
            receipt["stores"][s.name] = {"error": repr(e)}
    if request_type == "delete":
        for v in vendors:                           # pass deletion to service providers
            receipt["stores"]["vendor:" + v.name] = v.request_deletion(consumer_id)
    with open(f"receipts/{consumer_id}-{request_type}.json", "w") as f:
        json.dump(receipt, f, default=str)          # store receipts without the data itself
    return receipt

Deletion has exceptions, such as security, fraud prevention and legal obligations, so record a reason whenever something is kept. An evaluation set exported from logs must carry the consumer key or be de-identified first, or deletion silently misses it.

Fine-tunes, weights and AB 1008

Deleting a row is easy; deleting a person from a fine-tuned model is not. Whether a deletion request reaches weights is unsettled, and AB 1008 makes it harder to assume it never does. The defensible position has three parts. Keep personal information out of training data by default, and prefer retrieval, where deletion means removing a chunk, over fine-tuning for customer-specific facts. Keep lineage: which consumers' examples went into which model version. And retrain on a schedule from a cleaned dataset, so deletions have a known path into the next model, with output filters for known identifiers meanwhile. How LLMs memorise and leak personal data explains why the risk is real for duplicated and rare strings.

Opt-outs, GPC and sensitive information

If any flow counts as selling or sharing, consumers must be able to opt out, and the regulations require you to treat an opt-out preference signal such as Global Privacy Control as a valid opt-out for that browser and, where known, the consumer. Browsers that support it send the header Sec-GPC: 1. Honouring it is a small piece of middleware, but it must reach every downstream decision, including whether ad tags load on the chat page.

def privacy_context(request, user):
    gpc = request.headers.get("Sec-GPC") == "1"
    if gpc and user is not None and not user.opted_out:
        user.opted_out = True                      # persist for the known consumer
        user.save()
    return {
        "allow_sale_or_sharing": not (gpc or (user and user.opted_out)),
        "limit_sensitive": bool(user and user.limit_sensitive),
    }

The right to limit sensitive personal information confines its use to what is necessary to provide the service the consumer requested. For an assistant this means a detected health or financial detail can be used to answer this question, but not copied into long-term memory for personalisation or sent to analytics. Route sensitive spans through the same PII detection used for redaction, described in PII detection and redaction for LLMs, and tag them so downstream stores refuse them.

Automated decisionmaking and risk assessments

The California Privacy Protection Agency finalised regulations in September 2025 on automated decisionmaking technology, risk assessments and cybersecurity audits. They took effect on 1 January 2026 and phase in. ADMT means technology that processes personal information and uses computation to replace, or substantially replace, human decision making. The rules apply when it is used for a significant decision: providing or denying financial or lending services, housing, education, employment or compensation, or healthcare services. From 1 January 2027, about three months after this article's date, businesses using ADMT that way must give a pre-use notice in plain language, honour opt-outs and answer access requests about how the technology affected the decision. The main opt-out exception is a human appeal: a reviewer with authority to overturn the decision.

An LLM that ranks job candidates or triages patient messages may qualify if its output substantially replaces a human's judgement; an order-status chatbot does not. Risk assessments are required before processing that presents significant risk, including selling or sharing, processing sensitive personal information, using ADMT for significant decisions, and using personal information to train ADMT for those decisions. Existing processing must be assessed by 31 December 2027, and summaries go to the Agency from 1 April 2028. Cybersecurity audit reports are due from 1 April 2028, 2029 or 2030 depending on revenue. Make design reviews for AI features produce the risk-assessment record as a by-product.

Security and the breach exposure

The CCPA's private right of action is narrow but costly. If nonencrypted and nonredacted personal information is breached because the business failed to maintain reasonable security, consumers can claim statutory damages, $107 to $799 per consumer per incident in the 2025 adjustment, or actual damages if higher. Administrative fines run to $2,663 per violation and $7,988 per intentional violation or one involving consumers known to be under 16. Prompt logs are the usual weak point: plaintext, widely readable, copied into tickets. Encrypt them, restrict access, redact at write time, keep audit trails as in audit logging for LLM systems, and shorten retention.

Worked example: a retail support assistant

A retailer with 2 million California customers launches a support assistant. The inventory finds request logs kept for 400 days, a memory feature, a vector store of 3 million past tickets, an evaluation set of 20,000 real conversations, the model provider, and a marketing tag on the help page. Logs are now redacted at write and kept for 30 days, disclosed in the notice. Ticket chunks get the customer key in metadata, so deletion is a filtered delete. The evaluation set is rebuilt from de-identified conversations. The provider moves to terms that bar training and cap retention. The marketing tag leaves pages that render conversations. A test deletion completes across four internal stores in minutes, with the vendor's confirmation in the receipt. No significant decisions are made, but a planned refund-eligibility feature is flagged for a risk assessment.

Failure modes

  • Copies outside the map. Notebooks and evaluation exports hold prompts no handler knows about.
  • No consumer key at ingestion. Chunks and examples cannot be found later, so deletion becomes a full rebuild or a false statement.
  • Vendor terms drift. A new tool captures prompts, and a service-provider relationship quietly becomes a third-party disclosure.
  • GPC honoured on the website only. The chat widget's embedded tags ignore it.
  • Assuming weights are out of scope. After AB 1008 that position needs evidence such as de-identification or lineage, not assertion.

Trade-offs

Short retention reduces risk but removes debugging history; redacted logs with a short raw-text window are the usual compromise. Retrieval beats fine-tuning for deletability but costs latency and context. Zero-retention vendor modes may disable features you rely on. The human appeal route keeps automation in significant decisions but needs staffed, empowered reviewers. For the governance structure around these choices, see data governance for LLM systems.

What to do next

  1. Check the three applicability tests, and include employees and business contacts if you run internal copilots.
  2. Inventory every store that holds prompts, outputs, chunks, memories, evaluation copies and weights, with purpose, retention and recipients.
  3. Classify each AI vendor as service provider, contractor or third party, and file the contract terms that justify it.
  4. Write a consumer key into logs, chunks and training examples at ingestion time.
  5. Build the request fan-out with receipts, and test it end to end, including the vendor deletion request.
  6. Honour Sec-GPC everywhere ads or analytics load, including chat pages and apps.
  7. Tag sensitive spans and keep them out of memory, analytics and training.
  8. List features that could be ADMT for significant decisions before 1 January 2027 and schedule their risk assessments.
  9. Encrypt, redact and shorten prompt-log retention, then disclose the new periods.
Key takeaway: Treat CCPA compliance for an LLM product as a data-flow problem: inventory every store of prompts, chunks, evaluation copies and weights; keep model providers as service providers under contracts that bar training; write a consumer key into everything so requests fan out with receipts; honour Global Privacy Control; keep sensitive details out of memory and training; and find any significant-decision automation before 1 January 2027.