A membership inference attack asks one question about a trained model: was this particular example in its training data? The answer matters well beyond curiosity. If a fine-tuned model reveals that a patient's note or a customer's ticket was in its training set, that alone can be a privacy breach. Publishers want to know whether their articles were used to train a model, benchmark maintainers want to know whether test sets leaked into pretraining, and privacy teams need a number that says how much a model leaks before it ships.

This page explains the attacks as tools for auditing models you own or are authorised to test: what the signals are, why they work, and, most importantly, why many published results on large language models measured something other than membership. It ends with an audit design you can run on a fine-tuned model. Defenses, from data deduplication to differential privacy, are covered in membership inference defenses and the serving-side architecture in membership-inference defense architecture.

Advertisement

What a member is, and what the attacker can see

For a classifier, a member is a training example. For a language model the unit is fuzzier, and the choice changes the answer. It can be an exact sequence (this paragraph, verbatim), a document (this article, perhaps reformatted), a person (any text about them), or a whole dataset (this collection of books). Attacks on short sequences are hardest; attacks that aggregate evidence over many documents from one source are much easier, because many weak signals add up.

The second axis is access. With only generated text, an auditor can test whether the model completes a prefix verbatim, which is extraction rather than inference. With per-token log-probabilities, which many APIs expose for the output and some for the prompt, every score on this page is possible. With weights, the auditor can also compute gradients and train shadow copies. Finally, the auditor needs non-members: text from the same distribution that was certainly not used for training. That last requirement turns out to be the whole game.

A membership audit: the split decides whether the number means anythingCandidate poolone source, one periodrandom splitMembers (train)Held out (never train)Fine-tune or train targetplus injected canariesScore every exampletarget and reference modelScore functionsloss, ref ratio, zlib, Min-K%, Min-K%++MetricsTPR at 0.1% and 1% FPR, log-scale ROCBlind baselinetext-only classifiermust be near chanceReport and gate releasewith canary exposureIf the blind baseline separates the two sets, the audit measured a distribution shift, not memorisation.
An honest audit draws members and non-members from one pool by a random split, scores both, and checks that a model-free classifier cannot separate them.

The core signal: loss, and why raw loss is weak

Training lowers the loss on the examples the model sees. If the model fits its training set better than unseen text, members have lower loss on average, and thresholding the loss is an attack (Yeom, 2018). Everything else on this page is a refinement of that idea.

Raw loss is a weak signal because difficulty varies far more between examples than membership moves it. A boilerplate licence header has low loss whether or not it was trained on; a paragraph of rare jargon has high loss even if it was. A single threshold therefore mostly sorts text by how predictable it is. The refinements all calibrate for difficulty: they ask whether this example is easier for this model than it should be.

Advertisement

Calibrated scores

  • Reference model ratio. Score an example by how much lower its loss is under the target model than under a reference model trained on similar data without it (Carlini, 2021, used a smaller model from the same family). Predictable text is easy for both, so the difference isolates what the target learned specifically. For a fine-tuned model the natural reference is the base model it started from, which makes this the strongest cheap score in that setting.
  • zlib ratio. With no reference model, use a compressor as a crude one: divide the loss by the zlib-compressed length of the text (Carlini, 2021). Repetitive text compresses well and also has low loss, so the ratio discounts it.
  • Neighbourhood comparison. Generate small rewrites of the example (swap a few words with a masked language model) and compare its loss with theirs (Mattern, 2023). A member sits in a sharp dip that its neighbours do not share.
  • Min-K% probability. Average the log-probabilities of only the k% least likely tokens (Shi, 2023). Unseen text tends to contain a few very surprising tokens; seen text has fewer outliers.
  • Min-K%++. Standardise each token's log-probability against the mean and spread of the model's own next-token distribution at that position, then take the lowest k% (Zhang, 2024). It asks whether the actual token is unusually likely compared with the alternatives the model was considering.

The function below computes all of these except the neighbourhood score for one text. It expects a Hugging Face causal language model and a reference that shares its tokenizer; with different tokenizers, compare per-character or per-byte losses instead.

import zlib
import torch

@torch.no_grad()
def token_logprobs(model, tok, text, max_len=512):
    ids = tok(text, return_tensors="pt", truncation=True, max_length=max_len).input_ids.to(model.device)
    logp = torch.log_softmax(model(ids).logits[0, :-1].float(), dim=-1)  # predictions for tokens 1..L-1
    tgt = ids[0, 1:]
    return logp.gather(1, tgt[:, None]).squeeze(1), logp

def membership_scores(target, reference, tok, text, k=0.2):
    """Higher score = more likely a member. `reference` must share the tokenizer."""
    lp, logp = token_logprobs(target, tok, text)
    loss = -lp.mean().item()
    n = max(1, int(k * lp.numel()))
    # Min-K%: mean of the k% least likely tokens
    min_k = lp.sort().values[:n].mean().item()
    # Min-K%++: standardise each token against the model's own next-token distribution
    p = logp.exp()
    mu = (p * logp).sum(-1)
    sigma = ((p * logp ** 2).sum(-1) - mu ** 2).clamp_min(1e-6).sqrt()
    min_kpp = ((lp - mu) / sigma).sort().values[:n].mean().item()
    ref_lp, _ = token_logprobs(reference, tok, text)
    return {
        "loss": -loss,
        "ref_ratio": -ref_lp.mean().item() - loss,      # how much easier for target than reference
        "zlib": -loss / len(zlib.compress(text.encode("utf-8"))),
        "min_k": min_k,
        "min_k_pp": min_kpp,
    }

Shadow models and LiRA

The strongest attacks learn what a member looks like for each individual example. Shokri (2017) trained shadow models that imitate the target on data the attacker controls, then trained a classifier on their outputs. The likelihood ratio attack, LiRA (Carlini, 2022), sharpens this: train many shadow models on random halves of a candidate pool, so each example is in about half of them. For each example, fit one Gaussian to its scores from the models that included it and another to the models that did not. The target's score is then compared with both, which gives a per-example test rather than one global threshold.

For pretraining-scale models this is unaffordable. For fine-tuning it is practical: train eight or sixteen LoRA adapters on random halves of your fine-tuning set and you have the IN and OUT distributions for every example. That makes LiRA the gold-standard auditor for the case most teams actually face, a fine-tune on private data.

Measuring an attack: the low false-positive rate

Area under the ROC curve is the number most papers lead with and the least useful one. An attack can reach a respectable AUC by being slightly better than chance on everyone while identifying nobody with confidence. Privacy harm comes from confident identification of a few examples, so the metric that matters is the true positive rate at a very low false positive rate: of all members, how many can be flagged while wrongly flagging only one non-member in a thousand (Carlini, 2022). Plot ROC curves on log-log axes so this region is visible.

import numpy as np
from sklearn.metrics import roc_curve, roc_auc_score
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_predict

def tpr_at_fpr(y, s, target_fpr):
    fpr, tpr, _ = roc_curve(y, s)
    return tpr[np.searchsorted(fpr, target_fpr, side="right") - 1]

def report(y, s):
    return {"auc": roc_auc_score(y, s),
            "tpr@1%": tpr_at_fpr(y, s, 0.01),
            "tpr@0.1%": tpr_at_fpr(y, s, 0.001)}

def blind_baseline(texts, y):
    """No model access at all. If this separates members from non-members,
    your split leaks a distribution shift and the attack numbers are inflated."""
    X = TfidfVectorizer(min_df=2, ngram_range=(1, 2)).fit_transform(texts)
    s = cross_val_predict(LogisticRegression(max_iter=2000), X, y, cv=5,
                          method="predict_proba")[:, 1]
    return report(y, s)

To estimate TPR at 0.1% FPR you need thousands of non-members, or the threshold is set by a handful of examples and the number is noise. Report the count alongside the rate.

The distribution-shift trap

Auditing a pretrained model is hard because you rarely know its training set, so researchers built benchmarks from a proxy. Members were text published before the model's cut-off date (old Wikipedia articles, for example) and non-members were text published after it. The attacks scored well on these benchmarks. Then Duan (2024) evaluated on a model whose training data was public, with members and non-members drawn from the same distribution, and found the attacks barely beat chance for most settings. Das (2024) went further: "blind" classifiers that never query the model, using only features of the text such as dates and topical words, matched or beat published attacks on several benchmarks. Meeus (2024) surveyed the field and reached the same conclusion.

The lesson is not that membership inference does not work. It is that an attack evaluated on a non-random split measures how different the two sets are, and text written in different years differs in topics, names and vocabulary. The model had seen older topics, so their loss was lower whether or not those exact documents were in training. The fix is to make the split random and verify it: train the blind baseline from the code above, and if it separates members from non-members, discard the benchmark.

Dataset inference and canaries

Single-sequence membership on large pretrained models is close to undetectable with current methods, but a set of documents leaves a stronger trace. Dataset inference (Maini, 2024) combines several weak scores across many documents from one source with a statistical test against held-out documents from the same source, giving a calibrated answer to "was this dataset used?" The same caution applies: the held-out documents must come from the same distribution, ideally a random split made before training.

If you control training, the cleanest audit is to plant evidence. Insert random canary sequences (for example "the access code is" followed by random digits) a known number of times, train, and measure how highly the model ranks each true canary against many random alternatives; Carlini (2019) called this exposure. Canaries are cheap, have perfect ground truth and give a release-to-release trend. They are related to, but used differently from, the leak tripwires described in canary tokens.

Worked example: auditing a support-ticket fine-tune

A team fine-tunes an open-weights model on 40,000 customer support tickets and wants to know, before release, whether the model reveals which customers' tickets it saw. The numbers below are illustrative, not measurements.

  1. Split first. From one month of tickets, 44,000 are drawn; a random 4,000 are held out and never trained on. Fifty random canaries are inserted, ten copies each.
  2. Blind baseline. The TF-IDF classifier on members versus held-out tickets gives AUC near 0.5. The split is sound. Had the held-out tickets come from the following month, a product launch in that month would have let the classifier separate them.
  3. Scores. Each of 4,000 members (a random sample) and 4,000 held-out tickets is scored with the base model as reference. Raw loss is weak; the reference ratio is much stronger at 1% FPR, as expected for a fine-tune.
  4. LiRA. Eight LoRA shadows on random halves of the pool confirm which examples are most exposed: long tickets repeated with small edits, and tickets containing unusual names.
  5. Decision. The team deduplicates near-identical tickets, reduces epochs from four to two, re-runs the audit and canary exposure, and records both numbers as release criteria.

The audit result is attached to the model card. If customer personal data was involved, the team also checks direct leakage with prefix completion, which PII leakage covers.

Failure modes

  • Non-random splits. Members and non-members from different times or sources. Run the blind baseline every time.
  • Reporting only AUC. Hides whether anyone can be identified confidently.
  • Too few non-members. TPR at 0.1% FPR from 500 negatives is a guess.
  • Reference model trained on the same data. The ratio cancels the very signal you want.
  • Tokenizer mismatch. Comparing per-token losses across tokenizers is meaningless; normalise per byte.
  • Treating a weak attack as proof of safety. A failed attack is a lower bound on leakage, not a guarantee. Only differential privacy gives an upper bound.
  • Truncation. Scoring only the first 512 tokens of long documents misses late memorised spans; slide a window.

What to do next

  1. Decide the unit: sequence, document, person or dataset, and write it into the audit plan.
  2. Hold out a random sample from the same pool before training, and plant canaries.
  3. Run the blind baseline; do not proceed until it is near chance.
  4. Score with the reference ratio (base model for fine-tunes), Min-K%++ and zlib; add LiRA with LoRA shadows when the model handles sensitive data.
  5. Report TPR at 1% and 0.1% FPR with counts, plus canary exposure, and gate release on them.
  6. Choose defenses from the defenses guide and re-run the same audit to confirm they worked.
Key takeaway: Membership inference measures how much easier a model finds its training examples than comparable unseen text. Calibrate for difficulty with a reference model, Min-K%++ or zlib, use LiRA with cheap shadow adapters for fine-tunes, and judge results by true positives at a very low false positive rate. Above all, draw members and non-members by a random split and prove it with a blind baseline; without that, the audit measures a distribution shift and its number means nothing.