The United States has no general federal privacy law for consumer data. Instead, roughly twenty states have passed comprehensive consumer privacy statutes. California's came first; Virginia's 2021 law became the template most others copied, and each copy changed a threshold, a definition or a right. An AI product with users across the country is therefore subject to many overlapping rule sets at once. They agree on the broad shape and differ in the details that decide whether a given data flow is lawful.

This article is about the engineering problem that patchwork creates, not about any one statute. The California-specific mechanics are covered in the CCPA and LLM systems article. Here you will learn the common skeleton the state laws share, the dimensions on which they diverge, and how to build a policy engine that turns statutes into data your LLM stack can enforce: consent, opt-out signals, consumer rights, assessments and training-data controls. The specifics below were checked against public summaries in October 2026. They are illustrations, not legal advice; have counsel confirm current text before relying on them.

The skeleton every state law shares

Strip away the differences and almost every state law has the same five parts. Once you see them, a new statute becomes a diff against a known model rather than a fresh research project.

  1. Applicability. A business is covered if it does business in the state or targets its residents and crosses a threshold, usually a count of consumers whose data it processes in a year, sometimes combined with a share of revenue from selling data.
  2. Roles. A controller decides why and how data is processed. A processor acts on its behalf under a contract with required terms. Your foundation-model vendor is usually your processor; when you host a model for enterprise customers, you are usually theirs.
  3. Consumer rights. Access, correction, deletion, portability, and opt-outs from targeted advertising, sale and certain profiling, with an appeal process when a request is refused.
  4. Sensitive data. A defined list, typically health, precise geolocation, race or ethnicity, religion, sexual orientation, biometric and genetic data, and data about known children, with stronger controls than ordinary personal data.
  5. Duties. Purpose limitation, data minimisation, reasonable security, a privacy notice, and a written data protection assessment before high-risk processing such as targeted advertising, sale, sensitive-data processing or risky profiling.

Enforcement is mainly by state attorneys general; California adds a dedicated agency, and its only private right of action covers certain data breaches. Many of the cure periods that let businesses fix a violation after notice have expired, so a control promised on a roadmap does not count.

Where the states diverge

The differences that cost engineering time fall into a few dimensions. The table shows representative settings, not a complete survey; the point is to see how far apart the rows are.

DimensionRepresentative settingsEngineering consequence
ThresholdVirginia: 100,000 consumers, or 25,000 plus over 50% of revenue from sale. Colorado: the same counts, but any revenue or discount from sale qualifies. Maryland and Rhode Island: 35,000, or 10,000 plus over 20% of revenue from sale. Texas: no count; small businesses under the federal SBA definition are exempt from most duties. California: a revenue test indexed for inflation, among othersCount distinct consumers per state per year; a mid-size product crosses the low thresholds first
Sensitive dataOpt-in consent in the Virginia model (Virginia, Colorado, Connecticut, Texas and others). Notice and opt-out in Utah and Iowa. California: a right to limit use. Maryland: collection only where strictly necessary, and no sale at allYour consent model must express both default-deny and default-allow, plus outright prohibition
Universal opt-out signalRequired in about a dozen states, including California, Colorado, Connecticut, Texas, Oregon, Montana, New Jersey, Delaware, Minnesota and Maryland; not required in several othersHonour Global Privacy Control everywhere is simpler than per-state branching
MinorsTeen protections beyond federal COPPA's under-13 rule: for example, Oregon bars sale of data for known under-16s, and Connecticut and Colorado restrict processing for known minors under 18Age signals become policy inputs, not just a sign-up checkbox
ProfilingOpt-out of profiling for decisions with legal or similarly significant effects; Minnesota adds a right to question the result and learn the reasonsModel-driven decisions need explanations and a human review path
ScopeCalifornia covers employee and business-contact data; most others exclude itHR copilots and sales assistants are in scope in one state and out in another

Two divergences deserve special attention for AI teams. Maryland's minimisation standard ties collection to what is reasonably necessary for the requested product or service, which is much stricter than notice and consent. A blanket choice to log every prompt for future training is exactly what it questions. And laws are being amended, not just added: Connecticut's 2025 amendments lower its threshold to 35,000 consumers and pull in anyone who processes sensitive data or sells personal data, starting July 2026.

Architecture: statutes as data

The tempting design is one if per state scattered across services. It fails within a year, because a single amendment then means a hunt through the codebase. The design that survives treats the statutes as a versioned rules matrix. A small engine reads that matrix and answers one question per action: given this person, this data and this purpose, what is required? The diagram shows the pieces.

One policy engine, many statutes: rules as data, obligations computed per requestRequest or eventuser, signals, purposeResidency signalsaccount, billing, IPJurisdiction resolverset of states, confidenceRules matrixversioned YAML per stateObligation mergerper-state or strictestConsent storeopt-in, opt-outSignal handlerGPC, UOOMRights routeraccess, delete, appealLLM pipelinetrain, retrieve, logAssessmentsDPA registerEvidence logrule version + decision for every actionStatutes change several times a year; code reads the matrix, so an amendment is a data change plus a test.
Residency signals resolve to a set of states; the merger combines their rules into obligations; each enforcement point asks the engine and logs the rule version it applied.

The jurisdiction resolver ranks signals (billing address over declared state over IP geolocation) and returns a set of states, because a Maryland resident travelling in Texas may be both. The rules matrix is reviewed by counsel and versioned like code, and the evidence log records which rule version produced each decision, which is what you show a regulator later.

The rules matrix and merger in code

A compact version of the matrix and merger in Python follows. The rule values are illustrative and must come from counsel's reading of current law; the structure is what you keep.

from dataclasses import dataclass
from enum import IntEnum

class Sensitive(IntEnum):          # ordered from least to most restrictive
    NOTICE_OPT_OUT = 1
    LIMIT_RIGHT = 2
    OPT_IN = 3
    STRICTLY_NECESSARY = 4         # collect only when needed for the service; never sell

@dataclass(frozen=True)
class StateRule:
    state: str
    version: str                   # e.g. "MD-2025-10"; bump on every amendment
    consumers: int | None          # applicability count, None = no count test
    sale_share: tuple[int, float] | None   # (consumers, revenue share) alternative test
    sensitive: Sensitive
    honour_uoom: bool
    minor_age: int                 # protections apply below this known age
    profiling_opt_out: bool

def applies(rule, n_consumers, sale_revenue_share, sba_small=False):
    if rule.consumers is None:     # Texas-style: everyone except SBA small businesses
        return not sba_small
    if n_consumers >= rule.consumers:
        return True
    if rule.sale_share:
        n, share = rule.sale_share
        return n_consumers >= n and sale_revenue_share > share
    return False

@dataclass
class Obligations:
    sensitive: Sensitive = Sensitive.NOTICE_OPT_OUT
    honour_uoom: bool = False
    minor_age: int = 13            # federal COPPA floor
    profiling_opt_out: bool = False
    rule_versions: tuple = ()

def merge(rules):
    o = Obligations()
    for r in rules:
        o.sensitive = max(o.sensitive, r.sensitive)
        o.honour_uoom |= r.honour_uoom
        o.minor_age = max(o.minor_age, r.minor_age)
        o.profiling_opt_out |= r.profiling_opt_out
        o.rule_versions += (r.version,)
    return o

The key decision is in merge: every dimension is ordered, so combining states means taking the maximum. That ordering is the hard intellectual work and belongs in review with counsel. Dimensions that do not order cleanly, such as disclosure wording, are met state by state. Applicability needs a consumer count per state, which most analytics stacks lack until the day it is needed.

Worked example: one feature, three states

Take a writing assistant with 60,000 users in Maryland, 40,000 in Virginia and 8,000 in Utah, no data sales, and a feature that remembers health conditions a user mentions so that it can tailor advice.

  • Applicability. Maryland applies (over 35,000). Virginia does not yet (under 100,000, no sales), and this company misses Utah's revenue test. Growth in Virginia will change that, so re-evaluate monthly.
  • The memory feature. Remembered health conditions are sensitive data. Under Maryland's standard the question is not whether the user clicked consent but whether keeping the data is strictly necessary for the service the user asked for. A defensible design keeps memory off by default, scopes it to the user's explicit request, stores it apart from training corpora, and deletes it on request across the primary store, vector index and caches.
  • Training. Prompts from Maryland users that contain sensitive data are excluded from fine-tuning sets unless counsel signs off on a necessity argument, which for general model improvement is weak. The filter runs at ingestion and records the rule version.
  • Signals. A browser sending Global Privacy Control is treated as an opt-out of targeted advertising and sale for that user everywhere, not only in states that require it. That is cheaper than branching and removes one class of mistake.

Where the laws touch an LLM stack

State statutes were written for databases and ad tech, but their definitions reach model systems. Map each to a control.

  • Inferences are personal data. Several laws count inferences or profiles derived from a person as personal data. A model's summary that a user is probably pregnant is sensitive data, even if the user never said so. Classify model outputs that are stored, not only inputs.
  • Prompts and logs. Raw prompts often contain names, health details and other people's data. Set retention per purpose, redact before long-term storage, and make logs searchable by user so that access and deletion requests can actually be fulfilled. Techniques are in PII handling for LLMs.
  • Retrieval stores. Embeddings derived from a person's documents are copies of their data for rights purposes. Index them by data subject so that deletion is a query, not a re-embedding project.
  • Weights. Deleting a record from training data does not delete what a model memorised. The practical controls are upstream: filter at ingestion, keep lineage from dataset to checkpoint, and test for memorisation, as described in PII leakage from language models.
  • Profiling decisions. If a model scores applicants, tenants or borrowers, the profiling opt-out and assessment duties apply. Build a human-review path and store the features and explanation that produced each decision.

Operating the program

The engine is only as good as the process around it. Four routines keep it honest.

  1. Statute watch. Assign an owner to track bills and regulations, using a tracker your counsel trusts. Every enacted change becomes a ticket that edits the matrix, adds a test and records an effective date. The engine selects rule versions by date, so you can ship early and switch on time.
  2. Data map. Keep an inventory of stores, purposes and data categories, including vector indexes, evaluation sets and observability tools. Rights requests and assessments both depend on it.
  3. Assessment register. Write a data protection assessment for each high-risk activity: targeted ads, sale, sensitive data, risky profiling. Reuse one template for every state, with a section per state for its additions. Assessments for GDPR, covered in GDPR for LLM systems, and for the EU AI Act, covered in the EU AI Act guide, can share much of the same evidence.
  4. Rights SLOs. Most state laws require a response within 45 days, extendable in some cases. Measure time to completion across every store and alert well before the deadline.

Failure modes

The recurring failures are predictable, which means each can be tested for.

  • Single-state residency. The resolver returns one state and the strictest obligation is lost. Return sets and merge them.
  • Stale thresholds. A threshold is evaluated once, at launch, and never again. Re-evaluate on a schedule from real per-state counts.
  • Signal handled in one channel. Global Privacy Control is honoured on the website but not in the mobile app or API, or it is honoured for ads but not passed to data partners. Test every channel.
  • Deletion that misses copies. The primary row goes; the embedding, the cached completion and the analytics export stay. Drive deletion from the data map and verify it with a search afterwards.
  • No evidence. The control worked but nothing recorded which rule version applied. Without the log you cannot prove compliance, which in practice equals not complying.

Trade-offs

The central choice is between a national baseline and per-state precision. A baseline applies the strictest rule on each dimension to everyone. It is simple and ages well, because new laws tend to raise the floor, but it gives up data you could lawfully use in laxer states. Per-state precision keeps that value but multiplies test cases and depends on residency signals that are often wrong. Most teams take a hybrid: a baseline for anything a user can see, such as signals, rights and notices, and per-state handling only for a few high-value flows where counsel has signed off on how residency is determined. Consent and rights-automation vendors track statutes for you, but rarely understand vector stores, prompt logs or training lineage, so those controls stay yours.

What to do next

Use this checklist to turn the patchwork into a system you can run.

  1. Produce per-state consumer counts and sale-revenue shares from analytics, and re-run them monthly.
  2. Write the rules matrix for every state where you have users, with counsel, as versioned data with effective dates.
  3. Make the jurisdiction resolver return a set of states with confidence, and log its inputs.
  4. Honour Global Privacy Control in every channel and propagate it to processors and partners.
  5. Classify stored model outputs and inferences, not only inputs, against each state's sensitive-data list.
  6. Index prompt logs, embeddings and caches by data subject, and test end-to-end deletion with a search.
  7. Add an ingestion filter that keeps restricted data out of training sets and records the rule version.
  8. Assign a statute-watch owner, and schedule a quarterly review of the matrix against new laws and amendments.
Key takeaway: State privacy laws share a skeleton of applicability, roles, rights, sensitive data and assessments, but differ on thresholds, consent models, opt-out signals, minors and scope. Encode each statute as versioned data, resolve users to sets of states, merge obligations by a reviewed ordering, enforce at every store including prompts, embeddings and training sets, and log the rule version behind every decision.