A property inference attack learns a statistic of a model's training data that its owner never meant to publish: the share of patients with a diagnosis, the gender balance of a hiring dataset, the fraction of emails written by one company, whether a fine-tuning set came mostly from one region. No individual record is revealed. What leaks is a fact about the dataset, and that fact can be commercially or legally sensitive in its own right.
The attack has a ten-year history in classical machine learning and was shown in 2025 to work against fine-tuned LLMs. This article explains the threat model from first principles, formalises it as a distinguishing game, walks through the shadow-model, black-box, federated and poisoning variants, gives code for a generation-based attack and an audit harness, and explains why ordinary differential privacy, often suggested as the fix, does not address it.
Property versus membership versus attribute inference
Three inference attacks are easy to confuse. Each asks a different question, and defences differ accordingly.
| Attack | Question | Unit of secrecy | Typical defence |
|---|---|---|---|
| Membership inference | Was this record in the training set? | One record | Differential privacy, regularisation |
| Attribute inference | Given part of a record, what is its hidden attribute? | One record's field | Limit correlations, output restriction |
| Property inference | What is a global statistic of the training set? | The whole distribution | Distribution shaping, access limits, property-level audits |
Membership and attribute inference are covered in depth in separate articles linked at the end. Property inference is the distribution-level member of the family. Its harm model differs: the victim may be an organisation rather than a person, for example a hospital whose case mix reveals a business strategy, or a company whose fine-tuning set reveals which customer segment it serves. It can also be a population, when the property is the prevalence of a condition in a community.
The distinguishing game
The cleanest formalisation, used by Suri and Evans in their work on distribution inference, is a distinguishing game. The defender has a data distribution and a training recipe. Fix a property function, say the fraction of records with sex = F, and two values t0 and t1. The challenger draws a training set from the distribution conditioned on either t0 or t1 by a hidden coin flip, trains a model and gives the adversary whatever access the threat model allows. The adversary guesses the coin. Its advantage is accuracy minus one half.
This framing forces precision about three things. Access: white-box weights, black-box probabilities, text generations only, or gradients as a federated participant. Knowledge: the adversary usually needs auxiliary data from the same domain and knowledge of the training recipe. Ratio gap: distinguishing 30 percent from 70 percent is easy; 48 from 52 percent is hard. A useful report states advantage as a function of the gap, and can convert it into an estimate of how many records' worth of change the attack can detect.
White-box attacks: shadow models and meta-classifiers
Ateniese and colleagues introduced the meta-classifier approach in 2015 against support vector machines and hidden Markov models. Ganju and colleagues extended it to fully connected neural networks in 2018, solving a real obstacle: hidden units can be permuted without changing the function, so raw weight vectors make poor features. Their meta-classifier sorts or set-encodes neurons so the representation is permutation-invariant.
def property_inference_whitebox(aux_data, recipe, prop, t0, t1, n_shadow=200):
X, y = [], []
for i in range(n_shadow):
t, label = (t0, 0) if i % 2 == 0 else (t1, 1)
D = resample_with_property(aux_data, prop, ratio=t, size=recipe.train_size)
shadow = recipe.train(D, seed=i) # same architecture and hyperparameters
X.append(featurise(shadow)) # permutation-invariant weight summary
y.append(label)
meta = train_classifier(X, y) # e.g. small DeepSets network
return lambda target: meta.predict_proba(featurise(target))[1]The cost is training hundreds of shadow models, which is cheap for small tabular models and prohibitive for frontier LLMs. For LLMs, the practical shadow is a fine-tune of a public base model, because the sensitive property usually lives in the fine-tuning set rather than in pretraining.
Black-box and generation-based attacks
Query-based black-box attacks. Without weights, the adversary builds probe inputs whose loss or confidence is sensitive to the property and uses the vector of outputs as the meta-classifier's features. Zhang and colleagues showed in 2021 that dataset properties leak this way in multi-party settings even when the property is not a model feature at all.
Generation-based attacks on LLMs. Huang and colleagues introduced PropInfer, presented at NeurIPS 2025, a benchmark built on the ChatDoctor medical dialogue dataset, with properties such as patient gender ratio and disease prevalence. They proposed two attacks: a black-box generation attack that samples many completions and counts how often the property appears, and a shadow-model attack using word-frequency features. The generation attack was especially effective for models fine-tuned in chat-completion mode, the shadow attack for question-answering mode. The intuition is simple: a fine-tuned model reproduces the distribution it was trained on, so its outputs are a biased sample of its training set.
import random
PROMPTS = ["Patient: I have had a headache for three days.",
"Patient: My knee has been swelling after running."]
FEMALE = {"she", "her", "woman", "female", "pregnant"}
MALE = {"he", "his", "man", "male"}
def generation_estimate(model, n=2000, temperature=1.0):
f = m = 0
for _ in range(n):
out = model.generate(random.choice(PROMPTS), max_tokens=120,
temperature=temperature).lower().split()
f += any(w in FEMALE for w in out)
m += any(w in MALE for w in out)
return f / max(f + m, 1) # raw estimate, biased by the base model
def calibrated(target, shadows):
"""shadows: list of (known_ratio, model) fine-tuned from the same base.
Fit a line from raw estimate to true ratio, then apply it to the target."""
xs = [generation_estimate(s) for _, s in shadows]
ys = [r for r, _ in shadows]
a, b = linear_fit(xs, ys)
return a * generation_estimate(target) + bThe calibration step is what turns a vague signal into a number. The base model has its own gender skew, so the raw count is meaningless until a few shadow fine-tunes with known ratios map it onto the true scale.
Federated and poisoning-amplified variants
Federated and collaborative learning. Melis and colleagues showed in 2019 that a participant in collaborative training sees aggregated updates that reveal properties of other participants' batches, for example that a particular person appears in photos, or when a property starts appearing in another party's data. Updates are computed on small batches, so they carry far more distributional signal than a finished model.
Poisoning-amplified attacks. Mahloujifar, Ghosh and Chase showed in 2022 that an adversary who can contribute a small fraction of the training data can craft poison that makes the model's behaviour on chosen queries depend strongly on the target property, turning a weak signal into a reliable one. Chaudhari and colleagues made this far more efficient with SNAP in 2023. For any system that fine-tunes on user-contributed data, this is the variant to worry about.
Worked example: a hospital's case mix
A hospital network fine-tunes an open-weight model on 40,000 de-identified patient conversations and exposes it to partner clinics as a chat assistant. A competitor wants to know the network's case mix, in particular whether its oncology share is high enough to signal an expansion. The steps below are illustrative, following the PropInfer pattern.
- The attacker downloads the same base model and a public medical dialogue dataset as auxiliary data.
- They fine-tune six shadow models on subsets with oncology shares of 5, 10, 15, 20, 25 and 30 percent, copying the hospital's published fine-tuning recipe.
- They sample 2,000 completions from each shadow from symptom prompts and count oncology-related terms, fitting a calibration line from raw frequency to true share.
- Through ordinary partner access, they sample 2,000 completions from the target and apply the line. The estimate lands near 22 percent, against a regional norm of 12.
- No patient record was extracted and the model's output filter saw nothing unusual, because every single completion was a normal-looking medical answer.
The hospital's de-identification, membership-inference testing and PII filters all passed. None of them measured what this attack measures.
Why differential privacy is not the fix
The stub this article replaces suggested that differentially private training helps. As a general claim that is wrong, and the reason is worth understanding. Differential privacy with parameter epsilon bounds how much any one record can change the distribution of trained models. Group privacy extends the bound to k records, but the guarantee degrades to roughly k times epsilon. Moving a property from 10 to 20 percent of a 40,000-record set changes 4,000 records, so even a very strong per-record epsilon yields a vacuous bound on the property. DP training may add enough noise to blunt weak attacks in practice, but it provides no guarantee for dataset-level statistics, and those statistics are exactly what a useful model is supposed to learn.
That tension is fundamental. A model that predicts disease well must reflect disease prevalence. Defences therefore aim at properties the model does not need, or at limiting how precisely an outsider can read them.
Defences
- Decide which properties are secret. Write them down, as you would list sensitive fields. If the property is irrelevant to the task, it can be shaped away.
- Rebalance the training distribution. Resample or reweight so the sensitive ratio matches a public reference, which removes the signal at its source. This costs some fidelity when the property is correlated with the task.
- Restrict access. Generation and logit access is far more revealing than a classification label. Limit sampling volume per client, avoid exposing logits, and cap temperature where the product allows.
- Monitor query patterns. Thousands of near-identical sampling requests from one partner are an anomaly worth investigating.
- Guard the training pipeline. Against poisoning-amplified attacks, vet and rate-limit contributions to any fine-tuning set built from user data.
- Audit with your own attack. Before release, run the generation attack against your model with shadows at known ratios and report how precisely the property can be estimated.
def property_leakage_audit(base, recipe, aux, prop, ratios, target, trials=3):
shadows = [(r, recipe.finetune(base, resample_with_property(aux, prop, r), seed=s))
for r in ratios for s in range(trials)]
est = calibrated(target, shadows)
# Leave-one-out error tells you how precise an outside attacker could be.
errs = [abs(calibrated(m, [x for x in shadows if x[1] is not m]) - r) for r, m in shadows]
return {"estimate": est, "attacker_mae": sum(errs) / len(errs)}
Failure modes
| Failure | Consequence | Mitigation |
|---|---|---|
| Assuming DP covers it | False assurance; property still readable | Audit property leakage separately |
| Testing only membership inference | Dataset-level leak goes unmeasured | Add a property audit to the release gate |
| Publishing the fine-tuning recipe | Makes shadow models cheap and accurate | Publish less detail when the property is sensitive |
| Unlimited high-temperature sampling | Gives the generation attack its sample | Per-client quotas, anomaly alerts |
| Fine-tuning on open user contributions | Enables poisoning-amplified inference | Contribution vetting and caps |
Trade-offs
Fidelity versus secrecy. Rebalancing removes the leak and the information; acceptable for properties the task does not need, harmful for those it does.
Openness versus attack cost. Open weights and documented recipes help science and also make white-box attacks practical. Decide per model, based on which properties are sensitive.
Audit cost. A dozen small fine-tunes per release is real compute, but it is the only way to know your exposure, and far cheaper than learning it from a competitor's press release.
What to do next
- List the dataset-level statistics of each fine-tuning set that would harm you or a community if revealed.
- For each, decide whether the task needs it; rebalance or reweight the ones it does not.
- Add a property-leakage audit, like the harness above, to the model release checklist.
- Remove logit access and cap sampling volume per client on externally exposed models.
- Alert on bulk, repetitive sampling patterns from a single key.
- Vet any user-contributed fine-tuning data against poisoning.
- Read the PropInfer paper (arXiv 2506.10364) and the distribution inference work of Suri and Evans.
- Continue with membership inference, attribute inference, DP-SGD, model extraction and federated learning security.