Hospitals and health-tech companies hold exactly the data that makes clinical AI useful: millions of notes, labs, medications and outcomes. Turning that data into a model is a different act from using it to treat a patient. It is a secondary use, it needs its own legal basis, and the dataset you build outlives the project: once patient text is inside a fine-tuned model's weights, you cannot simply delete it.
This article is about that build step. Other pages on this site cover protecting patient data while an LLM application runs, such as gateways, logging and retrieval authorization. Here the question is how to assemble a training or evaluation dataset from patient records: choosing a legal pathway, removing identifiers from structured fields and free-text notes, measuring the re-identification risk that remains, honouring opt-outs, and recording lineage so a model can be traced back to the data that shaped it. It focuses on the US HIPAA framework with notes on the EU, and it is engineering guidance, not legal advice; the pathway decision belongs with your privacy office.
Choose the legal pathway first
Under HIPAA, a covered entity can use protected health information (PHI) for its own treatment, payment and health care operations. Building a model may fit operations when it is internal quality improvement, but developing a product or generalisable knowledge usually does not. For everything else there are four common pathways, and the choice decides what your dataset may contain.
| Pathway | What the data may contain | Main obligation | Fits |
|---|---|---|---|
| De-identified (Safe Harbor) | No listed identifiers; dates reduced to year | Remove 18 identifier types, no actual knowledge of re-identification | Broad sharing, vendor training |
| De-identified (Expert Determination) | Whatever an expert finds carries very small risk | Documented statistical analysis, conditions of release | Data needing dates, finer geography |
| Limited data set | Dates, city, state, ZIP; no direct identifiers | Data use agreement with the recipient | Research needing timelines |
| Authorization or IRB waiver | Identified data if justified | Patient authorization or privacy board waiver | Studies where identifiers are essential |
In the EU, health data is a special category under GDPR Article 9, research use relies on Article 9(2)(j) with Article 89 safeguards or national law, and pseudonymised data is still personal data. The European Health Data Space regulation adds a permit-based route for secondary use that applies in stages over the coming years. The practical difference from HIPAA is that "de-identified" in the US sense often remains personal data in Europe, so design for the stricter reading if you operate in both.
The pipeline
Three properties make this pipeline defensible. First, purpose and pathway are checked before any extraction, so the cohort query itself is limited to the fields the purpose needs, the minimum necessary idea applied to datasets. Second, opt-outs are applied at the source, before copies multiply. Third, de-identification is followed by measurement rather than trusted: the pipeline produces numbers (equivalence class sizes, note-scrubbing recall) and a release fails if they miss the threshold set by your privacy office or expert.
Structured fields: Safe Harbor and date shifting
Safe Harbor lists 18 identifier types: names; geographic subdivisions smaller than a state; all elements of dates except the year, plus ages over 89; telephone and fax numbers; email addresses; Social Security, medical record, health plan, account, certificate and licence numbers; vehicle and device identifiers; URLs and IP addresses; biometric identifiers; full-face photographs; and any other unique identifying number or code. The first three ZIP digits may stay only if the area they cover holds more than 20,000 people; otherwise they become 000, using current Census data.
import hashlib, hmac
from datetime import date, timedelta
def load_restricted_zip3() -> set[str]:
"""Return 3-digit ZIP areas with 20,000 or fewer people, built from current Census data."""
raise NotImplementedError("load from your Census extract")
RESTRICTED_ZIP3 = load_restricted_zip3()
def safe_harbor_row(r):
"""Structured fields only. Free text is handled separately."""
age = r["age"]
return {
"age": "90+" if age >= 90 else age,
"sex": r["sex"],
"zip3": "000" if r["zip"][:3] in RESTRICTED_ZIP3 else r["zip"][:3],
"admit_year": r["admit_date"].year, # month and day removed
"dx_codes": r["dx_codes"],
"labs": r["labs"],
# no MRN, no encounter id: link rows with a fresh random study id instead
}
def shifted(d: date, patient_id: str, key: bytes) -> date:
"""Per-patient date shift (keeps intervals). NOT Safe Harbor: needs Expert Determination."""
digest = hmac.new(key, patient_id.encode(), hashlib.sha256).digest()
offset = int.from_bytes(digest[:4], "big") % 365 - 182
return d + timedelta(days=offset)Date handling is where projects most often mislabel themselves. Models of readmission or disease progression need intervals, and a per-patient date shift preserves them. But shifted dates still contain month and day, so a shifted dataset is not Safe Harbor; it needs Expert Determination or a limited data set agreement. Keep the shift key in the clinical zone, never in the training zone, or the shift is reversible.
Measure re-identification risk
Removing the listed identifiers does not make every record anonymous. A combination of ordinary fields (age, sex, ZIP3, admission year, a rare diagnosis) can single out one person. These are quasi-identifiers, and the standard measurement is k-anonymity: group records by their quasi-identifier values and look at the size of each group. A record in a group of size 1 is unique in the release.
import pandas as pd
QI = ["age_band", "sex", "zip3", "admit_year"]
def k_report(df: pd.DataFrame, k_min: int = 5) -> dict:
sizes = df.groupby(QI, dropna=False)["study_id"].transform("count")
return {
"records": len(df),
"unique": int((sizes == 1).sum()),
"below_k": int((sizes < k_min).sum()),
"share_below_k": round(float((sizes < k_min).mean()), 4),
"min_k": int(sizes.min()),
}
def generalise(df):
out = df.copy()
out["age_band"] = (out["age"].clip(upper=90) // 10 * 10).astype(str) + "s"
return out
# release only if report["share_below_k"] is under the threshold agreed with your privacy officeMeasure on the quasi-identifiers an adversary could plausibly know: demographics and visit timing, plus anything published elsewhere, such as a rare diagnosis reported in local news. When too many records fall below k, generalise (wider age bands, admission year instead of month), suppress the rare records, or move to a stricter pathway. k-anonymity has known gaps: a group can be large but share one sensitive diagnosis, which is what l-diversity addresses. Treat k as a floor, not a proof.
Free-text notes
Clinical notes are where identifiers hide: "Mrs Alvarez's daughter called from Fresno", a phone number in a signature, a room number, a date written as "the day after Thanksgiving". Note de-identification combines pattern rules for formatted identifiers (phones, MRNs, dates) with a named-entity model for names and places, and then replaces findings rather than masking them.
Replacement with realistic surrogates, sometimes called hiding in plain sight, matters. If every detected name becomes [NAME], the one name the detector missed stands out; if detected names become plausible fake names, a missed real name is hard to tell apart from the fakes. Keep surrogates consistent within a patient so that "Mr Lee" stays the same person across notes.
Measure it like a classifier, because it is one. Have two annotators label identifiers in a random sample of notes, stratified by note type (discharge summaries, nursing notes and scanned letters behave differently), and report recall per identifier type. Overall recall hides failures: 99% recall on dates and 85% on names is a name problem. Recall matters more than precision here, since an over-scrubbed drug name costs some model quality while a missed patient name is a disclosure.
Memorisation in trained models
A model can reproduce training text verbatim, and clinical notes contain rare, specific strings. The controls are the ones described in LLM PII leakage: deduplicate notes (templated text repeated across thousands of notes is memorised first), insert canary records and test for their extraction before release, and consider differential privacy for fine-tuning when the model will leave the organisation. The patient-data-specific point is timing: these tests run on the de-identified training set and the trained model, and their results belong in the dataset and model records.
Opt-outs, restrictions and lineage
Patients may restrict uses, withdraw from research, or fall under extra protections such as substance use disorder records, reproductive health or psychotherapy notes. Apply these as a filter in the extraction query, sourced from an authoritative opt-out table, and record which version of that table each dataset used. When a patient withdraws later, you can then answer two questions precisely: which dataset versions contained them, and which models were trained on those versions.
dataset_manifest = {
"dataset": "readmit-notes-v7",
"pathway": "expert_determination", # with the expert's report id
"expert_report": "ED-2026-031",
"cohort_query_sha256": "b41d...",
"optout_table_version": "2026-09-30T00:00Z",
"deid_pipeline_version": "notes-deid 4.2",
"risk": {"share_below_k5": 0.0031, "name_recall": 0.982, "date_recall": 0.996},
"allowed_uses": ["internal_model_training", "internal_evaluation"],
"retention_until": "2028-10-01",
}Removing a person from a trained model is hard, so policy usually takes the practical route: the next dataset version excludes them, and the model is retrained on a defined schedule. Say that in your patient-facing notices instead of promising instant removal.
Federated learning and synthetic data
Two techniques are often proposed as ways to avoid this work. Federated learning keeps records at each hospital and shares model updates instead, which removes the central copy but not the risk: gradients and updates can leak training examples, so you still need secure aggregation, differential privacy or both, plus agreements covering each site. Synthetic patient data, generated by a model trained on real records, can be shared more freely only if it does not reproduce real patients. Test that by measuring each synthetic record's distance to its nearest real record and inspecting the closest matches, especially for rare conditions; synthetic data done right covers the evaluation side. Both are useful tools; neither is a legal pathway by itself.
Worked example: a readmission model
A health system wants a model that flags patients likely to be readmitted within 30 days, using structured data plus the discharge summary. Intervals matter, so Safe Harbor's year-only dates would destroy the signal. The privacy office chooses Expert Determination with per-patient date shifting, ten-year age bands, ZIP3, and note de-identification with surrogate replacement.
The first risk report shows 2.4% of records below k = 5, mostly elderly patients in small ZIP3 areas with rare admission diagnoses. The team merges all ages above 80 into one band, groups rare diagnoses into their parent categories for the quasi-identifier check, and reruns: 0.3% remain, which the expert accepts with suppression of those records. Note de-identification recall on a 400-note annotated sample is 98% for names but 91% for addresses in scanned referral letters, so scanned letters are excluded from this version. (These figures are illustrative of a typical first pass, not benchmarks.)
The model is trained in the research zone, canary extraction tests pass, and the registry entry points to the manifest above. When an external vendor later asks to fine-tune on the same data, the answer is no: the expert's determination covered internal use only, and a new release needs a new review.
Failure modes
| Failure mode | Consequence | Control |
|---|---|---|
| Shifted dates labelled Safe Harbor | Wrong pathway, unsupported claim | Pathway field in manifest, reviewed |
| Shift key in the training zone | Dates can be un-shifted | Keep keys in the clinical zone |
| Overall note recall reported | Weak identifier types hidden | Per-type recall by note type |
| Rare-condition records | Unique individuals in release | k report on QIs including rare diagnoses |
| Opt-outs applied downstream | Withdrawn patients in copies | Filter at extraction, version the table |
| Dataset reused for a new purpose | Use beyond the determination | Allowed uses in manifest, checked at access |
Trade-offs
Every step trades utility for risk. Coarser dates and geography lower re-identification risk and weaken temporal and regional signals. Aggressive note scrubbing protects patients and removes some clinical terms. Safe Harbor is simple to apply and audit but blunt; Expert Determination keeps more signal at the cost of expert time and conditions on use. A sound habit is to train a baseline on the most protective version first and relax only the fields whose value you can show on held-out data. For deployment-time controls, continue with HIPAA for healthcare LLM applications.
What to do next
- Write down the purpose of the model and get a pathway decision from your privacy office before extracting anything.
- List the fields the purpose needs and build the cohort query from that list only.
- Source opt-outs and special-category restrictions from an authoritative table and filter at extraction.
- Run structured de-identification, then a k report on realistic quasi-identifiers; generalise until it passes your threshold.
- Annotate a stratified note sample and report per-type recall before training on notes.
- Write a dataset manifest and link it from the model registry entry.
- Before deployment, add runtime controls from the HIPAA LLM gateway guide.