If a US healthcare provider, health plan or clearinghouse, or a vendor working for one, sends patient information to a large language model, HIPAA applies to every place that information goes: the prompt, the completion, the embeddings in the vector store, the cache, the traces and the evaluation set. The compliance basics are covered in HIPAA for LLM systems: whether HIPAA applies, business associate agreements (BAAs), minimum necessary, de-identification and the Security Rule safeguards. This article is the engineering companion. It shows how to build an LLM application that handles protected health information (PHI) defensibly.

The approach is a single choke point, a PHI gateway, in front of every model and retrieval call, plus a map of where PHI hides in an LLM stack. This is engineering guidance, not legal advice; your privacy officer and counsel decide your obligations.

Where PHI hides in an LLM stack

Teams usually protect the prompt and forget the rest. Treat each of these as PHI unless you can show otherwise:

WhereWhy it is PHIControl
Prompts and completionsThey contain patient facts; completions restate themGateway, BAA-covered endpoint, encrypted body store
Embeddings and vector storeDerived from PHI text; inversion research shows text can be partly recovered from embeddingsSame safeguards as the source documents; patient metadata for filtering
Prompt and semantic cachesA cache hit can serve one patient's answer to another userKey by patient and user scope, or disable for PHI
Traces and observabilityTracing tools capture full prompts by defaultLog metadata and hashes; send bodies only to a covered store
Evaluation and fine-tuning setsBuilt from real conversationsDe-identify properly or keep under the same safeguards and BAAs
Tool-call arguments and resultsContain identifiers such as record numbersAuthorize each tool call against the user's patient access
User feedback and support ticketsUsers paste screenshots and transcriptsRoute through the same classification as prompts

Every vendor in that table that creates, receives, maintains or transmits PHI on your behalf is a business associate and needs a BAA, and subcontractors need one with them. That includes the observability vendor that stores traces. Check that the specific product and configuration you use is covered; BAA coverage is often limited to particular services and settings.

Architecture: one gateway for every model call

A PHI gateway between clinical apps and modelsClinical appuser, patient, purposePHI gateway1. authenticate user, check purpose2. authorize patient relationship3. minimum-necessary context4. pseudonymize identifiers (vault)5. route only to BAA-covered endpoint6. check output, re-identify7. write audit event (no bodies)Model endpointunder a BAARetrievalpatient-scoped filterToken vaultrandom tokens, keyed accessrequestAudit storewho, which patient, purpose, hashesBody store (optional)encrypted, short retention, restrictedEverything right of the gateway, and both stores, still hold PHI: they need safeguards and BAAs.The gateway reduces how much PHI flows; it does not make the flow non-PHI.
Every model call passes through one service that enforces purpose, authorization, minimum necessary and routing, and records an audit event without copying prompt bodies into logs.

The gateway is a small internal service that every model call must go through. Network policy enforces that: application workloads have no egress route to model providers except via the gateway, so a developer cannot bypass it with a direct SDK call. The gateway does seven things in order:

  1. Authenticate and capture purpose. Who is calling, on whose behalf, and why: treatment, payment or operations. Purpose drives what is allowed.
  2. Authorize the patient relationship. The user must have access to this patient, by the same rules as the clinical record itself. The model is never the access-control layer.
  3. Apply minimum necessary. Build context from the fields this task needs, not the whole chart.
  4. Pseudonymize direct identifiers that the task does not need, replacing them with random tokens kept in a vault.
  5. Route only to endpoints on an allowlist of BAA-covered services, with provider-side retention configured as agreed.
  6. Check the output for identifiers that do not belong to this patient, then restore tokens for display.
  7. Audit the call with metadata, not bodies.

The gateway in code

A minimal gateway core in Python shows the shape. Production code adds authentication, a real vault, a stronger detector and retries, but the order of checks is the point.

import hashlib, re, secrets, time
from dataclasses import dataclass

ALLOWED_ENDPOINTS = {            # reviewed list: endpoint -> BAA reference
    "llm-prod-east": "BAA-2025-014",
}
ALLOWED_PURPOSES = {"treatment", "operations"}
MRN = re.compile(r"\bMRN[:\s]*\d{6,10}\b")
PHONE = re.compile(r"\b\d{3}[-.\s]\d{3}[-.\s]\d{4}\b")


@dataclass
class Call:
    user_id: str
    patient_id: str
    purpose: str
    endpoint: str
    text: str


def pseudonymize(text, vault):
    """Replace identifiers with random tokens; the mapping stays in the vault."""
    def swap(m):
        token = "ID_" + secrets.token_hex(4)   # random, not derived from the value
        vault[token] = m.group(0)
        return token
    return PHONE.sub(swap, MRN.sub(swap, text))


def reidentify(text, vault):
    for token, value in vault.items():
        text = text.replace(token, value)
    return text


def handle(call, can_access, build_context, model, foreign_ids, audit):
    if call.purpose not in ALLOWED_PURPOSES:
        raise PermissionError("purpose not allowed")
    if not can_access(call.user_id, call.patient_id):
        raise PermissionError("no relationship to patient")
    if call.endpoint not in ALLOWED_ENDPOINTS:
        raise PermissionError("endpoint not BAA-covered")
    vault = {}
    prompt = pseudonymize(build_context(call.patient_id, call.text), vault)
    started = time.time()
    output = model(call.endpoint, prompt)
    leaked = [i for i in foreign_ids(call.patient_id) if i in output]
    audit({
        "ts": started, "user": call.user_id, "patient": call.patient_id,
        "purpose": call.purpose, "endpoint": call.endpoint,
        "baa": ALLOWED_ENDPOINTS[call.endpoint],
        "prompt_sha256": hashlib.sha256(prompt.encode()).hexdigest(),
        "output_sha256": hashlib.sha256(output.encode()).hexdigest(),
        "tokens_replaced": len(vault), "blocked": bool(leaked),
    })
    if leaked:
        raise RuntimeError("output references another patient; withheld for review")
    return reidentify(output, vault)

Three details matter. The authorization check uses the same function the clinical application uses, so model access can never exceed record access. The audit record lets you answer "who sent what about which patient, where, and why" without storing the prompt; the hashes let you match a record to a body held in a restricted store. And the output check compares against identifiers known to belong to other patients in the context, such as other names in a shared retrieval index, which catches the most damaging leak: one patient's data shown in another patient's workflow.

Regular expressions catch structured identifiers like record numbers and phone numbers. They miss names, addresses and dates in free text, so production systems add a trained PHI detector. Treat any detector as exposure reduction with a measured miss rate, never as a guarantee. The PII protection guide compares detection approaches.

Pseudonymization is not de-identification

Pseudonymized prompts are still PHI, so the model endpoint still needs a BAA. HIPAA recognizes two de-identification methods in 45 CFR 164.514. Safe Harbor requires removing 18 categories of identifiers about the individual and their relatives, employers and household members, including names, geographic units smaller than a state (with a narrow three-digit ZIP exception), all date elements except year, phone and record numbers and full-face photos, and having no actual knowledge that the remainder could identify the person. Expert Determination requires a qualified expert to determine and document that the risk of re-identification is very small.

Clinical free text rarely meets Safe Harbor after regex redaction: admission dates, ages over 89, rare diagnoses and phrases like "the mayor's wife" survive. Two further rules catch engineers out. A re-identification code may be kept only if it is not derived from information about the individual, so a hash or HMAC of the record number is not an acceptable token, which is why the code above uses random tokens. And the mapping must not be disclosed for other purposes. So use pseudonymization to minimize what the vendor sees, and use formal de-identification only for data that genuinely leaves the HIPAA boundary, such as analytics or model training sets shared outside your covered entity.

Retrieval and caching without cross-patient leaks

Retrieval-augmented generation is where most real leaks happen, because the retriever, not the model, decides what PHI enters the prompt. Store patient and encounter identifiers as metadata on every chunk and apply the filter inside the vector query, before similarity ranking, never by asking the model to ignore other patients. Separate indexes by tenant or facility when organizations must not see each other's data. Include document sensitivity labels (behavioral health, substance use disorder records, which carry additional federal rules, and reproductive health) and exclude them unless the purpose allows.

Re-check authorization at the time of the query, not when the index was built: care relationships end. When a patient's record is amended or a document is withdrawn, delete or re-embed its chunks, and make the vector store part of your deletion and retention procedures. Caches need the same scoping: a semantic cache keyed only on the question text will happily return Patient A's answer to a question about Patient B.

Audit logging and retention

HIPAA's audit-controls standard requires mechanisms that record and examine activity in systems containing electronic PHI. For an LLM application, the audit event from the gateway is that record: user, patient, purpose, endpoint, BAA reference, timestamps, hashes and policy decisions. Keep it in an append-only store with restricted access; the LLM audit logging guide covers tamper evidence.

Prompt and completion bodies are a separate decision. Keeping them helps quality review and incident investigation but creates a large PHI repository. If you keep them, encrypt them, restrict access to a named group, log every access, and set a short retention period. Security Rule documentation, such as policies, risk analyses and procedures, must be retained for six years. Decide audit-log retention with your compliance team, and make sure debug logging in every library is configured so bodies never reach general-purpose logs.

LLM-specific incidents and breach triage

LLM systems add new ways to make an impermissible disclosure: a prompt injection in a retrieved document that makes the assistant reveal another patient's data, a cache serving a cross-patient answer, a developer pointing traffic at an endpoint not covered by a BAA, or traces shipped to an uncovered vendor. Under the Breach Notification Rule, an impermissible use or disclosure of unsecured PHI is presumed to be a breach unless a documented risk assessment shows a low probability that the PHI has been compromised. The assessment considers the nature of the data, who received it, whether it was actually viewed, and how far the risk was mitigated.

If it is a breach, individuals must be notified without unreasonable delay and no later than 60 days after discovery. Breaches affecting 500 or more people also require notice to HHS within the same window and, when 500 or more residents of a state or jurisdiction are affected, to prominent media. Smaller breaches are logged and reported to HHS annually. Business associates must notify the covered entity. Your gateway audit trail is what makes the risk assessment fast: it tells you exactly which patients, which users and which endpoints were involved. The PII leakage article covers injection-driven exfiltration patterns.

Worked example: drafting portal replies

A health system wants an assistant that drafts replies to patient portal messages for nurses to review. Walk one message through:

  1. A nurse opens a message from a patient asking whether a new rash could be from a recently started antibiotic. The app calls the gateway with purpose "treatment".
  2. The gateway confirms the nurse is on the patient's care team. It builds context from the message, the active medication list and allergies only: no full history and no notes from other specialties.
  3. The record number and phone number in the message are replaced with random tokens. The patient's name stays, because the draft must address them; this is why the endpoint must be BAA-covered regardless.
  4. Retrieval pulls the health system's own medication-reaction guidance, an index that contains no patient data, so no patient filter is needed there.
  5. The draft comes back, passes the foreign-identifier check, has its tokens restored, and appears in the nurse's queue marked as AI-drafted. The nurse edits and sends it; the final message lives in the clinical record as usual.
  6. The audit event records nurse, patient, purpose, endpoint, BAA reference and hashes. No prompt body is kept beyond a seven-day encrypted review store.

Note what the design avoids: no PHI in the retrieval index, no PHI in general logs, no path to any non-covered endpoint, and a human between the model and the patient.

The regulatory horizon

In January 2025, HHS proposed a major update to the Security Rule. Among other changes, it would remove most of the distinction between required and addressable safeguards and would expect encryption of electronic PHI at rest and in transit, multi-factor authentication, an asset inventory and network map, and regular vulnerability scanning. As of mid-2026 the update had not been finalized; check the Federal Register for its current status. A gateway design like the one above already meets most of these expectations, so building to the proposed bar is cheap insurance. Also track state health-privacy laws, which can be stricter than HIPAA, and pair the program with a framework such as SOC 2 for vendor assurance.

Failure modes

FailureHow it happensPrevention
Direct SDK calls bypass controlsA team adds a provider SDK for a prototypeEgress policy that only the gateway can pass
Traces full of PHI at an uncovered vendorDefault tracing captures bodiesMetadata-only tracing; BAA for any body store
Cross-patient answerUnscoped retrieval or cachePatient filter inside the query; scoped cache keys
Hash of the record number used as a tokenAssumed to be de-identificationRandom tokens in a vault
Eval set built from production chatsCopied to a shared bucketFormal de-identification or same safeguards
Over-broad contextWhole chart sent for convenienceTask-specific context builders and reviews

What to do next

  1. Inventory every place prompts, completions, embeddings, caches, traces and eval data are stored, and the vendor and BAA for each.
  2. Put a gateway in front of all model and retrieval calls and block direct egress to providers.
  3. Reuse your clinical access-control function for model calls; add a purpose field to every request.
  4. Write task-specific context builders that select only the fields each task needs.
  5. Replace derived tokens with random vault tokens, and measure your PHI detector's miss rate on real samples.
  6. Move patient filters inside vector queries and scope cache keys by patient and user.
  7. Switch tracing to metadata and hashes; set an explicit retention period for any body store.
  8. Add LLM scenarios (injection, cache leak, uncovered endpoint) to your incident runbook and rehearse the breach risk assessment.
Key takeaway: Building a HIPAA-ready LLM application is mostly about controlling flows. Route every model and retrieval call through one gateway that checks purpose and patient access, sends only the minimum necessary, replaces identifiers with random tokens, reaches only BAA-covered endpoints, checks outputs and writes audit events without bodies. Remember that pseudonymized prompts, embeddings, caches and traces are still PHI, and plan incident triage before you need it.