A lender adds a chat assistant that explains a customer's loan balance. A tax preparer summarises uploaded returns with a model. A wealth app lets an agent read account history to answer questions. Each of them sends nonpublic personal information about consumers into a language model pipeline, and in the United States the Gramm-Leach-Bliley Act (GLBA) decides how that information must be protected, who it may be shared with, and what happens when it leaks.

GLBA is not one rule. It is a privacy rule, a safeguards rule enforced by several regulators, and a ban on obtaining customer information by false pretences. This article maps each part onto an LLM architecture: which rulebook applies to you, how the Safeguards Rule's required elements become concrete controls on prompts, logs and vendors, why a model provider that trains on your prompts can break the Privacy Rule's service-provider exception, and how the breach notification clocks work. It is engineering guidance, not legal advice.

What GLBA covers

GLBA covers financial institutions, defined broadly as businesses significantly engaged in financial activities. Banks and credit unions are obvious; so are mortgage brokers, consumer lenders, payday lenders, tax preparers, debt collectors, many fintech apps and car dealers that arrange financing. The protected data is nonpublic personal information (NPI): personally identifiable financial information a consumer gives you, that results from a transaction, or that you otherwise obtain in providing a financial product, plus lists derived from it. Even the fact that someone is your customer is NPI.

Three parts of the law matter for an LLM system. The Privacy Rule, implemented for most institutions by the CFPB's Regulation P at 12 CFR Part 1016, requires privacy notices and gives consumers a right to opt out before NPI is shared with nonaffiliated third parties, subject to exceptions. The Safeguards Rule requires an information security program. The pretexting provisions make it unlawful to obtain customer information by impersonating the customer, which is exactly the attack a conversational assistant invites.

Which rulebook applies to you

The safeguards obligations come from different regulators depending on what kind of institution you are. Get this right first, because the notification deadlines differ.

InstitutionSafeguards sourceIncident notice to regulator
Non-bank financial institutions under FTC jurisdiction (lenders, brokers, tax preparers, many fintechs)FTC Safeguards Rule, 16 CFR Part 314, as amended in 2021Notify the FTC of a notification event involving at least 500 consumers, no later than 30 days after discovery (effective May 13, 2024)
Banks and their regulators (OCC, Federal Reserve, FDIC)Interagency Guidelines Establishing Information Security StandardsComputer-security incident notification rule: tell the primary regulator within 36 hours of determining a notification incident occurred
Broker-dealers, investment advisers, investment companies, transfer agentsSEC Regulation S-P, amended in 2024Notify affected individuals of sensitive customer information incidents; compliance by December 3, 2025 for larger entities and June 3, 2026 for smaller ones

State breach notification laws apply on top of all three, with their own definitions and clocks. The FTC amended its rule in late 2021; most of the new requirements took effect on June 9, 2023, after a delay, so programmes written against the older rule are often missing elements.

Reference architecture

Where GLBA obligations attach in an LLM servicing assistantCustomerapp, chat, voiceLLM gatewayMFA, authz, tokenize NPIOrchestratortools, retrievalModel providerservice providertokensCore systemsaccounts, loansscoped APIToken vaultencrypted, keyedLogs and tracesencrypted, 2-year purgeIncident runbook500 consumers, 30 daysGreen and grey: Safeguards Rule controls you operate. Red: Privacy Rule service-provider terms.Amber: where customer information piles up and where notification duties start.
A servicing assistant with the GLBA attachment points marked. Customer information should reach the model only as tokens unless the task truly needs the value.

The gateway authenticates the customer and resolves which accounts they may see. The orchestrator fetches only the fields a tool call needs through scoped APIs to the core systems. Before anything goes to the model provider, account numbers, SSNs and similar identifiers are replaced with tokens from a vault, and the gateway restores them in the response only for an authorised viewer. Logs and traces, which in LLM systems tend to capture whole prompts, sit in the same encrypted, retention-limited store as other customer information. The incident runbook watches the places NPI accumulates, because those are where a notification event will start.

The Safeguards Rule, element by element

The FTC rule at 16 CFR 314.4 lists the elements of an information security programme. Here is how each maps onto an LLM deployment.

ElementLLM translation
Qualified individual oversees the programmeName who owns LLM risk; the model platform is in their scope, not a side project
Written risk assessmentCover prompt injection, data exfiltration through tool calls, prompt and trace retention, and vendor training on inputs
Access controlsThe model acts with the customer's or employee's permissions, never a service account that sees every account
Inventory of data and systemsList every store of prompts, completions, embeddings, eval sets and caches
Encryption in transit over external networks and at restIncludes vector indexes and log pipelines, not only the core database
Secure development practicesThreat-model new tools; red-team prompts before release
Multi-factor authentication for anyone accessing any information systemCovers internal agent consoles and prompt-debugging tools
Disposal no later than two years after last use, unless neededPurge prompt logs and conversation memory on schedule
Change managementPrompt templates, model versions and tool definitions are changes; review them
Monitoring and logging of authorised usersLog who asked what about which customer, with tamper protection
Continuous monitoring, or annual penetration tests and vulnerability assessments every six monthsInclude the assistant and its tools in scope
Oversee service providersSelect, bind by contract and periodically reassess the model and vector database vendors
Written incident response plan; annual report to the boardAdd LLM-specific scenarios and report on them

Institutions that hold customer information on fewer than 5,000 consumers are exempt from four of these elements under 314.6: the written risk assessment, the testing cadence, the written incident response plan and the annual board report. The rest still apply.

The Privacy Rule and your model provider

Sending NPI to a model provider is a disclosure to a nonaffiliated third party. Without an exception, that would require notice and an opt-out. Most institutions rely on the service-provider exception in 12 CFR 1016.13, which applies when the third party performs services for you and you have a contract that prohibits it from disclosing or using the information other than to carry out the purposes for which you disclosed it. The exception removes the opt-out, not the initial privacy notice. The reuse limits in 1016.11 then bind what the recipient may do with the data.

That condition is where LLM terms of service matter. If a provider may use your inputs to train models it sells to others, or keep them for its own product purposes, it is using NPI beyond the purpose you disclosed it for. Enterprise and API terms from major providers generally exclude training on customer inputs, but defaults, abuse-monitoring retention and feature-specific terms vary, so read the contract and the account settings rather than the marketing page. The same applies to vector database, observability and evaluation vendors that receive prompts.

Tokenising NPI before the model call

Tokenising NPI before the model call is the single control that does the most work: it shrinks what the provider sees, what lands in traces and what a breach can expose. A minimal version looks like this.

import re, hmac, hashlib

PATTERNS = {
    "SSN":  re.compile(r"\b\d{3}-\d{2}-\d{4}\b"),
    "ACCT": re.compile(r"\b\d{10,16}\b"),
}

class Vault:
    """Deterministic tokens per tenant; values stored encrypted with a separate key."""
    def __init__(self, tenant_key: bytes, store):
        self.key, self.store = tenant_key, store

    def tokenize(self, kind, value):
        tag = hmac.new(self.key, value.encode(), hashlib.sha256).hexdigest()[:12]
        token = f"[{kind}_{tag}]"
        self.store.put_encrypted(token, value)
        return token

def redact(text, vault):
    for kind, pat in PATTERNS.items():
        text = pat.sub(lambda m: vault.tokenize(kind, m.group(0)), text)
    return text

def restore(text, vault, viewer):
    if not viewer.may_see_full_identifiers:
        return text                      # tokens stay, or show only the last four digits
    return re.sub(r"\[(SSN|ACCT)_[0-9a-f]{12}\]", lambda m: vault.store.get_decrypted(m.group(0)), text)

Regexes catch structured identifiers; they do not catch a balance described in words or a name. Combine them with the more important control, which is not fetching data the task does not need. A question about a payment date needs the due date and amount, not the full transaction history. Deterministic tokens let the model refer to the same account consistently across turns without seeing the number.

Worked example: a lender's servicing assistant

Consider a consumer lender under FTC jurisdiction with 40,000 borrowers that launches a servicing assistant. A caller on the voice channel says they are a borrower, gives a name and date of birth, and asks for the account's payoff amount and the bank account on file.

The pretexting provisions are why the assistant must not treat conversational identity as authentication. The gateway requires the same step-up the human call centre uses, such as a one-time code to the phone on file, before the orchestrator gets any account scope. After verification, the payoff tool returns the amount; the bank account on file goes back to the model only as a token, and the response shows the last four digits.

Six months later the team discovers that a debug flag sent full prompts, before tokenisation, to a third-party tracing service, and that an access key to that service was exposed in a public repository. Logs show the bucket was downloaded. The prompts contain unencrypted names, dates of birth and loan balances for 1,800 borrowers. That is a notification event: unencrypted customer information acquired without authorisation. Because it involves at least 500 consumers, the lender must notify the FTC through its online form as soon as possible and no later than 30 days after discovery, and must separately work through each state's consumer notification law. The rule treats encrypted data as unencrypted if the key was also accessed, so encryption with a key stored next to the data would not have helped.

The post-incident fixes follow the elements table: tokenise before tracing, add the tracing vendor to the service-provider inventory, purge traces on the disposal schedule, and add debug flags to change management.

Failure modes

  • Service-account retrieval. The assistant queries core systems with one broad credential and relies on the prompt to stay in scope. One injected instruction reads another customer's account.
  • Traces outside the programme. Prompt logs in an observability tool that is not in the data inventory, not encrypted at rest and never purged.
  • Vendor terms drift. A new feature, such as stored conversations or fine-tuning, changes the provider's retention, and nobody reassesses the service-provider exception.
  • Conversational authentication. The assistant accepts knowledge-based answers that an impersonator can find, enabling exactly the pretexting the law prohibits.
  • Clock confusion. The team knows the state deadline but not the FTC 30-day or bank 36-hour one.

Trade-offs and further reading

Tokenisation costs some answer quality, since the model cannot reason about a value it never sees; route the few tasks that genuinely need values to a tighter path with stronger logging. Zero-retention endpoints may rule out some providers or features. Purging traces after the disposal window makes long-term debugging harder; keep tokenised traces longer and raw ones briefly. Strict step-up authentication adds friction, but the alternative is an assistant that hands account data to anyone with a birth date.

For deeper treatments of individual controls, read LLM security in finance, third-party LLM risk, encryption in LLM systems, audit logging and LLM incident response.

What to do next

  1. Confirm which safeguards rulebook applies to you and write down every notification deadline that could apply.
  2. Add prompt logs, traces, embeddings, eval sets and caches to the data inventory, encrypted at rest.
  3. Tokenise NPI at the gateway before model calls and before any tracing or analytics hop.
  4. Make tools act with the caller's permissions; remove broad service-account access to core systems.
  5. Review each model and tooling vendor against the service-provider exception: no training, no reuse, stated retention.
  6. Require step-up authentication before any account scope on chat and voice channels.
  7. Set purge jobs within the two-year disposal window and include the assistant in penetration tests.
  8. Add an LLM leak scenario, including the 500-consumer FTC threshold, to the incident response plan and board report.
Key takeaway: For an LLM system, GLBA means three things: treat every store of prompts and traces as customer information under your safeguards programme, keep model and tooling vendors inside the service-provider exception by contract and configuration, and know which regulator's notification clock starts when that data leaks.