A membership inference attack asks whether a specific record was in a model's training data. For a model fine-tuned on patient notes, support tickets or HR files, a confident "yes" is itself the leak: it discloses that a person was a patient, a customer or an employee, even if none of the record's text is ever reproduced. The membership inference architecture article lays out the layered picture of where defenses sit. This article is the working companion: how to measure the leak properly, what each defense does to that measurement, which defenses give a guarantee and which only raise the attacker's cost, and how LLMs change the picture.
The single most important idea is that you cannot defend what you do not measure, and that the usual measurement, accuracy or AUC of an attack, hides exactly the cases that matter. So we start with the audit.
What a defense has to reduce
An attack computes a score for a candidate record and thresholds it. Averaged metrics are misleading because leakage is concentrated: most records are barely memorised, while a few outliers, unusual or duplicated records, are memorised strongly. An attack can have an AUC near 0.5 overall and still identify a handful of records with near certainty. Carlini and co-authors made this argument in "Membership Inference Attacks From First Principles" (IEEE S&P 2022) and proposed reporting the true positive rate at a low false positive rate, for example TPR at 0.1% FPR: of the non-members, only 1 in 1,000 is wrongly flagged; what fraction of members is correctly flagged?
That is the number to drive down. A defense that cuts average accuracy but leaves TPR at 0.1% FPR unchanged has protected the easy cases and left the people actually at risk exposed. Report it per data slice too, because rare groups tend to be outliers and leak more. And size the non-member set for the FPR you report: a stable estimate needs roughly 10/FPR non-members, so 10,000 or more at 0.1% FPR. With fewer, report TPR at 1% FPR with a bootstrap confidence interval.
Build the audit before the defense
The audit needs a member set (records you trained on), a non-member set drawn from the same distribution and never trained on, and one or more scoring functions. The code below computes TPR at fixed FPRs from any score, and four scores that are standard for language models.
import numpy as np, zlib
def roc_tpr_at(scores_in, scores_out, fprs=(0.001, 0.01)):
"""Higher score = more likely member. Returns TPR at each target FPR."""
out_sorted = np.sort(scores_out)[::-1]
result = {}
for f in fprs:
k = max(int(f * len(out_sorted)), 1)
thresh = out_sorted[k - 1] # at most f of non-members at or above this
result[f] = float(np.mean(scores_in > thresh))
return result
def llm_scores(nll_per_token, text, ref_nll_per_token=None, k=0.2):
"""nll_per_token: target model's negative log-likelihood of each token of `text`."""
nll = np.asarray(nll_per_token)
s = {"loss": -nll.mean()} # plain loss attack
s["zlib"] = -nll.sum() / len(zlib.compress(text.encode())) # penalise easy, repetitive text
worst = np.sort(nll)[-max(int(k * len(nll)), 1):] # Min-K%: least likely tokens
s["min_k"] = -worst.mean()
if ref_nll_per_token is not None: # reference-model calibration
s["ref"] = np.mean(ref_nll_per_token) - nll.mean()
return s
# members: records you trained on; non_members: held out from the SAME distribution
# for name in ("loss", "zlib", "min_k", "ref"):
# print(name, roc_tpr_at(member_scores[name], non_member_scores[name]))- Loss: the mean negative log-likelihood. Members tend to have lower loss. Simple, and dominated by how intrinsically easy a text is.
- zlib ratio: total log-likelihood divided by compressed length, from the training-data extraction work by Carlini et al. (2021). Repetitive, boilerplate text is easy for every model; dividing by compressibility removes some of that.
- Min-K% Prob: the average of the least likely K% of tokens, from Shi et al., "Detecting Pretraining Data from Large Language Models" (ICLR 2024). An unseen text usually contains a few surprising tokens; a memorised one has fewer.
- Reference-model ratio: how much lower the target's loss is than a similar model trained without the candidate. This is the idea behind LiRA, which trains many shadow models with and without each record and runs a likelihood ratio test. It is the strongest and most expensive family; for a fine-tuned LLM the base model before fine-tuning is a cheap, useful reference.
The non-member set is where audits go wrong. If members are older articles and non-members are newer ones, a score can separate them by date or topic alone and you will measure distribution shift, not memorisation. Duan et al., "Do Membership Inference Attacks Work on Large Language Models?" (COLM 2024), found that on carefully constructed pretraining splits most attacks performed close to random, and attributed several earlier positive results to this kind of shift. Build non-members by random hold-out from the same pool before training; that is easy for your own fine-tuning data and nearly impossible after the fact.
Defense 1: fix the data
Duplication is the strongest single driver of memorisation. Kandpal, Wallace and Raffel, "Deduplicating Training Data Mitigates Privacy Risks in Language Models" (ICML 2022), showed that sequences repeated many times in training are regenerated far more often than sequences seen once, and that deduplication substantially reduces that. For your own corpora, run exact deduplication on normalised text and near-duplicate detection with MinHash over shingles, and deduplicate by person as well as by document: ten tickets from the same customer are ten chances to memorise that customer's details.
Filter what should never be there. Records with identifiers the task does not need, such as account numbers and addresses, should be removed or pseudonymised before training; the PII leakage article covers detection pipelines. Insert canary records, unique random strings, into the training set so the audit has records whose membership you know with certainty and whose exposure you can measure directly.
Defense 2: generalise better
Leakage tracks the gap between how the model treats training and unseen data. Anything that shrinks that gap helps: fewer epochs, early stopping on held-out loss, weight decay, dropout, data augmentation, and larger, more diverse training sets. For fine-tuning, the most common mistake is simply training too long on a small dataset until training loss approaches zero. Parameter-efficient methods such as LoRA reduce the number of trainable parameters but are not a privacy mechanism on their own; a low-rank adapter can still memorise a small dataset.
Research defenses target the gap directly. Adversarial regularisation (Nasr, Shokri and Houmansadr, 2018) trains the model against an inference adversary. RelaxLoss (Chen et al., 2022) stops pushing training loss below a target level so members and non-members end with similar loss distributions. These are useful, but they are empirical: they reduce the attacks you tested, and they give no bound against a stronger attack you did not.
Defense 3: differential privacy, the one with a guarantee
Differential privacy bounds how much any one training record can change the distribution of trained models. The DP-SGD article covers per-example clipping, noise and accounting. What matters here is the translation into membership terms. For an (epsilon, delta)-DP training procedure, any membership test satisfies TPR at most e^epsilon times FPR plus delta. Compute what that means at the FPRs you report:
import math
for eps in (0.5, 1, 2, 4, 8):
for fpr in (0.001, 0.01):
print(eps, fpr, round(math.exp(eps) * fpr, 4)) # plus delta; above 1 means no guarantee| epsilon | TPR bound at FPR 0.1% | TPR bound at FPR 1% |
|---|---|---|
| 0.5 | 0.0016 | 0.016 |
| 1 | 0.0027 | 0.027 |
| 2 | 0.0074 | 0.074 |
| 4 | 0.055 | 0.55 |
| 8 | 2.98, so no guarantee | 29.8, so no guarantee |
Two lessons. At small epsilon the guarantee is strong and holds against every attack, including ones not yet invented. At epsilon 8, a value commonly reported for private training of large models, the worst-case bound says nothing at low FPR, even though empirical attacks against such models usually do much worse than the bound. So a DP model at a large epsilon still needs the empirical audit.
The second lesson is about the unit of privacy. Standard DP-SGD protects one training example. If one person contributes fifty documents, their protection degrades with the number of contributions. For people-level claims, use user-level DP, where clipping and accounting are done per user, or cap the number of records per person before training.
Defense 4: distil or ensemble
If the deployed model never saw the sensitive records directly, it memorises less of them. PATE (Papernot et al., 2017) trains an ensemble of teachers on disjoint partitions of private data, aggregates their votes with noise, and trains a student on public data labelled by that aggregate; the student is what ships, with a DP guarantee on the labels. Softer variants train a student on a teacher's outputs over reference data, as in distillation for membership privacy (Shejwalkar and Houmansadr, 2021), or have each ensemble member withhold its prediction for records it trained on, as in SELENA (Tang et al., 2022). These are more practical for classifiers than for large generative models, but the principle carries over: a model trained on a teacher's outputs over public prompts inherits far less of any single private record.
Defense 5: harden the output, and know its limits
Serving controls reduce the signal an attacker can read: return labels or text rather than probabilities, round or truncate confidences, disable per-token log-probabilities on sensitive fine-tuned models, and rate-limit systematic probing. MemGuard (Jia et al., 2019) went further and perturbed confidence vectors to fool attack classifiers.
These controls do not change what the model has memorised, and label-only attacks bypass them. Choquette-Choo et al., "Label-Only Membership Inference Attacks" (ICML 2021), showed that measuring how robust a prediction is to small input perturbations recovers much of the membership signal without any confidence scores. For LLMs, generation is itself an output channel: prompting with a record's prefix and checking whether the model completes it verbatim needs no log-probabilities at all. The same applies to anyone who can obtain the weights, whether through a leak, an internal user or an extraction attack. Treat output hardening as a cost increase layered on top of real defenses, never as the defense.
Worked example: a fine-tuned support model
A company fine-tunes an open model on 40,000 resolved support tickets to draft replies. The audit first holds out 12,000 more tickets at random from the same pool and inserts 50 canary strings. After three epochs the loss score gives a TPR near zero at 0.1% FPR overall, but the reference-model score, target loss minus base-model loss, flags a small cluster of tickets with high confidence, and 12 of 50 canaries complete verbatim from their prefixes.
Investigation finds the flagged tickets are from a few enterprise customers who opened hundreds of near-identical tickets. The team deduplicates by customer, caps each customer at 20 tickets, removes account numbers, trains one epoch with early stopping, and turns off log-probabilities on the endpoint. The re-audit shows no canary completions and TPR at 0.1% FPR indistinguishable from the hold-out baseline for the reference-model score. Because the model will be offered to regulated customers, they also run a DP fine-tune with per-customer accounting, record the utility cost, and keep it as the release candidate for those deployments.
Failure modes
- Auditing with shifted non-members. You measure date or topic differences and declare leakage, or you fix a fake problem and miss the real one.
- Reporting AUC only. Hides the outliers that carry most of the risk.
- Example-level epsilon quoted as person-level protection. Repeated contributors are less protected than the number suggests.
- Output hardening as the defense. Defeated by label-only and generation attacks and by anyone holding the weights.
- Auditing once. Every retrain, new data source or extra epoch changes leakage; audit each release.
- Ignoring the base model. A pretrained model may already contain public copies of data you consider private; the reference-model score separates what fine-tuning added.
What to do next
- Before your next fine-tune, hold out enough random non-members for the FPR you report (about 10/FPR) and insert canaries.
- Implement the four scores and report TPR at 0.1% and 1% FPR, overall and per slice, for every release.
- Deduplicate by document and by person, and cap contributions per person.
- Tune epochs and early stopping for the smallest train-to-held-out gap that meets quality, and record the gap.
- Decide whether you need a guarantee; if so, pick user-level or example-level DP explicitly and compute the TPR bound at your epsilon.
- Disable log-probabilities on sensitive endpoints and add rate limits, then re-run the audit with a generation-based attack that does not need them.