HIPAA does not mention language models, and it does not need to. It regulates protected health information wherever it goes, and an LLM application moves PHI to more places than a typical clinical system: into prompts, retrieval indexes, caches, application logs, evaluation sets, vendor-side logs and sometimes training data. Each of those copies is in scope, and most compliance failures in healthcare AI come from copies the team forgot existed rather than from the model itself.
This article explains, from first principles, which parts of HIPAA bind an LLM application and how to turn them into architecture: the business associate chain, the minimum necessary standard applied to prompt assembly, de-identification for data that leaves the clinical boundary, the Security Rule technical safeguards, audit logging and breach assessment. It ends with a worked example and a checklist. It is an engineering guide, not legal advice; your privacy officer and counsel decide how the rules apply to your organisation. General PII handling is covered in PII in LLM systems and leakage paths in LLM PII leakage.
Does HIPAA apply to your application?
HIPAA applies to covered entities, which are health plans, health care clearinghouses and health care providers that conduct standard electronic transactions, and to their business associates: anyone who creates, receives, maintains or transmits PHI on their behalf. PHI is individually identifiable health information held by those parties, in any form. If you build an LLM feature for a hospital, an insurer or a clinic, you are almost certainly a business associate, and every vendor you pass PHI to becomes a subcontractor business associate.
A consumer wellness app that collects health data directly from users, with no covered entity involved, is often outside HIPAA. That does not make it unregulated: the FTC Health Breach Notification Rule and state health privacy laws may apply instead. Settle this question first, in writing, because it decides which of the controls below are legal obligations and which are good practice.
Map every place PHI lands
The first technical deliverable is a data flow map. Start from the user request and follow PHI through each component, writing down where it is stored, for how long and who can read it. In an LLM application the list is longer than people expect.
The Security Rule requires an accurate and thorough risk analysis of the confidentiality, integrity and availability of ePHI (45 CFR 164.308(a)(1)). This map is its backbone. A useful rule: if a component can see a prompt or a completion, it handles PHI until proven otherwise. That includes observability tools that capture request bodies, prompt-management platforms, evaluation services and the model provider's abuse-monitoring logs.
The business associate chain
Before PHI reaches any third party, you need a business associate agreement with that party (164.502(e) and 164.504(e)). For an LLM application that usually means the model API provider, the cloud hosting the vector store, the logging and tracing vendor, and any annotation service that sees real records.
Three points trip teams up. First, a BAA covers specific services, so confirm that the exact model endpoint and configuration you call is in scope, not just the vendor in general. Second, read the data handling terms alongside the BAA: how long prompts and completions are retained, whether they are used for abuse monitoring or training, and in which regions they are processed. Third, HHS guidance on cloud computing says a provider that stores ePHI is a business associate even if the data is encrypted and it holds no key; the conduit exception is narrow and covers transmission-only services, not a service that keeps logs of your prompts. Vendor programs change often, so verify the current terms for each service yourself rather than relying on a list.
Minimum necessary, applied to prompts
The minimum necessary standard (164.502(b)) requires that uses and disclosures of PHI be limited to what is reasonably needed for the purpose. It does not apply to disclosures to a provider for treatment, but most LLM features also serve other purposes, such as summarising for billing, drafting patient messages or analytics, and even for treatment, sending less data is lower risk.
In an LLM application, minimum necessary becomes a property of the prompt builder. Each use case declares which fields it may read, and the builder refuses everything else. This is far easier to audit than a policy that tells the model to ignore what it does not need.
USE_CASES = {
# use case -> FHIR resource types and fields the prompt may include
"discharge_summary_draft": {
"Encounter": ["period", "reasonCode", "hospitalization.dischargeDisposition"],
"Condition": ["code", "clinicalStatus"],
"MedicationRequest": ["medicationCodeableConcept", "dosageInstruction"],
},
"appointment_reminder": {
"Appointment": ["start", "serviceType"],
},
}
def build_context(use_case, user, patient_id, fhir):
allowed = USE_CASES[use_case] # KeyError = unknown use case
if not fhir.user_can_access(user, patient_id): # same check the EHR applies
raise PermissionError("no treatment relationship")
out = {}
for rtype, fields in allowed.items():
for res in fhir.search(rtype, patient=patient_id):
out.setdefault(rtype, []).append({f: dig(res, f) for f in fields})
audit.record(user=user.id, patient=patient_id, use_case=use_case,
resources=sorted(allowed))
return outRetrieval needs the same discipline. A vector index of clinical notes must filter by patient and by the caller's authorisation before similarity search, not after, and the filter must come from the authenticated session rather than from the prompt. An index that returns another patient's note because it was semantically close is an impermissible disclosure, whatever the model then does with it.
De-identification for data that leaves the clinical boundary
Evaluation sets, fine-tuning corpora, analytics and debugging copies should not contain PHI unless there is a strong reason. HIPAA offers two de-identification methods (164.514(b)). Safe Harbor requires removing 18 categories of identifiers of the individual and of relatives, employers and household members, including names, geographic units smaller than a state (with a limited three-digit ZIP exception), all date elements except the year, ages over 89, contact details, record and account numbers, device identifiers, biometrics, full-face photographs and any other unique identifying number or code, and requires that you have no actual knowledge that the remainder could identify the person. Expert Determination instead has a qualified expert apply statistical or scientific methods and document that the risk of re-identification is very small; the rule sets no numeric threshold.
Structured fields are easy to strip. Clinical free text is not: names appear in note bodies, dates hide in phrases such as "two days after her 2019 surgery", and rare diagnoses combined with a small town can identify someone with no identifier present. Automated redaction, whether rule-based, NER-based or LLM-based, has false negatives. Treat it as one layer: redact, then sample and review, measure recall on a labelled set, and prefer Expert Determination for large free-text corpora. A limited data set under a data use agreement (164.514(e)) is a third option when dates or towns are needed for research.
def deidentify_for_eval(note, redactor, reviewer_queue, sample_rate=0.05):
red = redactor.redact(note.text) # replaces spans with [NAME], [DATE] ...
residual = residual_checks(red) # regexes: MRN formats, phones, emails, dates
if residual:
return None # fail closed: never export on a hit
if random.random() < sample_rate:
reviewer_queue.put((note.id, red)) # human recall check
if note.patient_age > 89: # Safe Harbor: ages over 89 aggregate
return {"text": red, "age_band": "90+"} # and drop the year that reveals age
return {"text": red, "year": note.date.year} # otherwise keep year only
Security Rule safeguards mapped to LLM components
| Technical safeguard (164.312) | What it means in an LLM application |
|---|---|
| Access control | Unique user IDs; the model service account cannot read more than the calling user; retrieval filtered by authorisation |
| Audit controls | Record who accessed which patient through which feature, including model-assisted reads |
| Integrity | Protect stored prompts, outputs and indexes from improper change; version the prompt templates |
| Person or entity authentication | SSO with strong authentication for users; workload identity for services calling the model |
| Transmission security | TLS on every hop, including to the model API and the vector store |
Under the current rule, encryption at rest and in transit are addressable specifications: you must implement them or document why an equivalent measure is reasonable. In practice, encrypt everything. A Security Rule update proposed by HHS in January 2025 would make many addressable items, including encryption and multi-factor authentication, mandatory; as of this writing it has not been finalised, so build to the current rule and track the rulemaking. Keep the documentation of policies, risk analyses and decisions for six years (164.316(b)(2)). For logging patterns, see audit logging for LLM systems.
Audit logs without copying PHI into them
Audit controls create a tension: you must record access, but logs full of prompt text become a second, less protected PHI store. Log identifiers and metadata, not content: user, patient identifier, use case, resource types read, model and prompt-template version, token counts, latency and outcome. If you need content for debugging, store it in a separate encrypted store with its own access controls, short retention and its own audit trail.
{"ts": "2026-10-01T09:14:22Z", "event": "llm.generate",
"user": "u-48213", "role": "hospitalist", "patient": "p-0091772",
"use_case": "discharge_summary_draft", "resources": ["Condition","Encounter","MedicationRequest"],
"template": "dsd-v7", "model": "endpoint-A", "in_tokens": 2310, "out_tokens": 412,
"outcome": "draft_shown", "content_ref": "vault://llm-content/7f3c..."}
Model-specific risks
- Memorisation: a model fine-tuned on PHI can reproduce fragments of it to other users. Fine-tune on de-identified data, or not at all.
- Cross-patient leakage: shared caches, conversation memory and retrieval indexes must be keyed by patient and user, never global.
- Prompt injection: text inside a record or uploaded document can instruct the model to reveal other data or call tools. Limit what tools can reach, as described in LLM deployment hardening.
- Hallucinated clinical content is a patient safety issue rather than a privacy one, but it belongs in the same risk analysis: drafts need clinician review before they enter the record.
When something goes wrong
An impermissible use or disclosure of unsecured PHI is presumed to be a breach unless a documented risk assessment shows a low probability that the PHI was compromised (164.402). That assessment considers four factors: the nature and extent of the PHI, the unauthorised person who received it, whether it was actually acquired or viewed, and how far the risk has been mitigated. A model that showed one patient's medication list to another clinician without a treatment relationship is a candidate breach and needs this assessment.
Individuals must be notified without unreasonable delay and no later than 60 days after discovery (164.404). Breaches affecting 500 or more people also require notice to HHS and, for 500 or more residents of a state or jurisdiction, to prominent media; smaller breaches are logged and reported to HHS annually (164.406, 164.408). A business associate must notify the covered entity, also within 60 days (164.410). Your audit logs are what make the assessment possible: without them you cannot say who saw what.
Worked example: a discharge summary assistant
A hospital wants an assistant that drafts discharge summaries for hospitalists. The team maps flows: EHR via FHIR, prompt builder, model API, application logs, tracing and an evaluation set. The model provider, the tracing vendor and the cloud running the service each sign BAAs; tracing is reconfigured to drop request bodies. The prompt builder uses the discharge_summary_draft allowlist above and checks the treatment relationship through the EHR. Audit events go to the existing access-audit pipeline so privacy staff can review AI-assisted access alongside ordinary chart access. The evaluation set is built from 300 notes de-identified by redaction plus residual checks, with 5 percent reviewed by hand; a reviewer finds a nurse's first name in a free-text field, so the redactor is retrained and the sample rate raised until recall is confirmed. Drafts are shown for editing and are never written to the chart automatically.
Trade-offs
| Decision | Option A | Option B |
|---|---|---|
| Model hosting | Vendor API under a BAA: fast, depends on vendor terms | Self-hosted model: full control, you own the security work |
| Debug content | Store prompts in a secured vault: debuggable, more PHI to protect | Store metadata only: less risk, harder debugging |
| De-identification | Safe Harbor: mechanical, weak on free text | Expert Determination: suits free text, needs an expert and documentation |
What to do next
- Decide in writing whether you are a covered entity, a business associate or outside HIPAA.
- Draw the PHI data flow map, including logs, caches, traces, evaluation data and vendor-side retention.
- Confirm a BAA covers every third-party service on the map, and read each one's retention terms.
- Implement per-use-case field allowlists in the prompt builder and authorisation filters in retrieval.
- Switch application logs to identifiers and metadata only; move any content to a secured, short-retention store.
- De-identify evaluation and fine-tuning data, measure redaction recall, and update the risk analysis with the LLM components.