Differential privacy gives a mathematical promise: anything released about a dataset would have been almost as likely if any one person's data had been left out. The size of the promise is the privacy budget, written as epsilon and delta. A release with a small epsilon reveals very little about any individual. Every further release reveals a little more, and the budget is the total you are willing to let it add up to.

In practice the budget behaves like money in an account. Every query, model and dashboard that touches the data spends it, the total depends on how you do the arithmetic, and once spent it cannot be earned back. This article explains what the numbers mean, how composition and accounting work, and how to run a budget ledger, with a worked allocation for a year of releases.

How the budget is spent inside a single training run (per-example clipping, noise and the training accountant) is covered in DP-SGD, in depth. This page is about the budget across everything you release.

What epsilon and delta promise

A randomised mechanism M is (epsilon, delta)-differentially private if, for any two datasets D and D' that differ in one unit of privacy, and any set of outputs S, Pr[M(D) in S] <= eepsilon Pr[M(D') in S] + delta. Epsilon bounds how much any outcome's probability can change when one unit is added or removed. Delta is a small allowance for that bound to fail.

Take an attacker trying to decide whether a target is in the data. Any test they run has a true-positive rate of at most eepsilon times its false-positive rate, plus delta. At epsilon 1 that ratio is about 2.7; at epsilon 8 it is about 2,981, which is a much weaker promise. Delta should be far smaller than one over the number of units. A mechanism that publishes one randomly chosen record satisfies (0, 1/n)-DP, so a delta near 1/n allows exactly that kind of leak. The measured side of this, running membership attacks against a model, is covered in membership inference attacks.

The unit of privacy

The budget is meaningless until you say what one unit is. Record-level privacy protects one row, such as a single prompt or support ticket. User-level privacy protects everything one person contributed. For LLM data the difference is large: a user who wrote 400 prompts is protected far less by a record-level guarantee than its epsilon suggests.

The tool that connects the two is group privacy. A pure epsilon-DP guarantee for one record gives k times epsilon for any group of k records. For approximate DP, a group of k gets (k epsilon, k e(k-1)epsilon delta), and that delta grows quickly. So a record-level epsilon of 0.5 over a user with 20 records is, at best, a user-level epsilon of 10.

The practical answer is contribution bounding: before any mechanism runs, keep at most m records per user, or clip each user's total influence on a statistic. Noise is then calibrated to that bound, and the guarantee is user-level by construction. Bounding lives in the data pipeline, which is why it is often missing.

Composition: how costs add up

Composition is how per-release costs combine into a total. Four rules cover most of the arithmetic.

  • Basic (sequential) composition. k mechanisms on the same data, with guarantees (epsiloni, deltai), together satisfy (sum of epsiloni, sum of deltai). It is always valid and often very loose.
  • Advanced composition. k mechanisms each (epsilon, delta) together satisfy (epsilon sqrt(2k ln(1/delta')) + k epsilon (eepsilon - 1), k delta + delta') for any delta' you choose. The total grows roughly with the square root of k rather than linearly.
  • Parallel composition. Mechanisms run on disjoint subsets of units cost only the maximum of their epsilons, not the sum. Counts per region, where each user is in exactly one region, cost one release, not one per region.
  • Post-processing. Anything computed from a DP output without touching the raw data is free. Rounding, plotting, caching and re-serving a noisy number cost nothing more.
ReleasesBasic totalAdvanced total (delta' = 1e-6)
50 at epsilon 0.15.04.24
1,000 at epsilon 0.0110.01.76

Advanced composition only pays off for many small releases. For a few large ones, the second term dominates and basic composition can even be tighter. Neither is the best available for the Gaussian noise that most modern systems use, which is why accountants moved to Renyi DP.

Renyi DP and zCDP accounting

Renyi differential privacy (RDP) describes a mechanism by a curve: for each order alpha > 1, a bound epsilon(alpha) on the Renyi divergence between outputs on neighbouring datasets. RDP curves compose by adding them order by order. And any RDP curve converts to an (epsilon, delta) guarantee: epsilon = min over alpha of epsilon(alpha) + ln(1/delta)/(alpha - 1).

The Gaussian mechanism with L2 sensitivity Delta and noise standard deviation sigma has the curve epsilon(alpha) = alpha Delta2 / (2 sigma2). Zero-concentrated DP (zCDP) captures the same thing with a single number, rho = Delta2 / (2 sigma2), which also adds under composition and converts as epsilon = rho + 2 sqrt(rho ln(1/delta)).

The difference is not cosmetic. Take 50 Gaussian releases with sensitivity 1 and sigma 10. Each has rho = 0.005, so together rho = 0.25, which converts to epsilon 3.97 at delta 1e-6; an RDP accountant over a grid of orders gives the same 3.97. Convert each release to (epsilon, delta) first and add them with basic composition, and you get 50 times 0.60, or 30.0. The noise and the data are identical; only the accounting changed. Compose in RDP or zCDP and convert once, at the end.

Subsampling amplification

Sampling makes privacy cheaper. If a mechanism runs on a random subset that includes each unit independently with probability q, and the attacker does not know who was sampled, an epsilon-DP mechanism becomes ln(1 + q(eepsilon - 1))-DP. With q = 0.01 and epsilon = 1, the amplified cost is 0.017. DP-SGD relies on exactly this: each step touches a small random batch, and the accountant uses the RDP curve of the subsampled Gaussian rather than of the full-data mechanism.

Amplification holds only if the sampling really happened as stated. A loader that shuffles and takes fixed consecutive batches is not Poisson sampling, and if sample membership leaks through logs or a fixed seed, the amplification is gone.

A budget ledger

A budget that lives in a spreadsheet is a budget nobody enforces. A ledger is a small service that every consumer must call before running a mechanism on a protected dataset. It holds the spent RDP curve for each (dataset, unit of privacy) pair and a cap set by policy. It adds the request's curve, converts to (epsilon, delta) and approves only if the cap still holds. The charge happens before the noise is drawn and is never refunded: once a noisy value exists, it may have been seen.

A privacy budget ledger: every release is charged before it runsWeekly dashboardGaussian countsDP fine-tunesubsampled GaussianAnalyst queryad hoc, reviewedBudget ledgerRDP curve per dataset + unitrequestrequestrequestPolicy capeps 4 at delta 1e-6 per yearMechanism runsnoise drawn onceDeniedbudget would exceed capapprovedenyReleased outputpost-processing is free
Figure 1. Consumers ask the ledger before running a mechanism; it approves only if the total spent on that dataset and unit stays under the cap.
import math
import threading

# Renyi orders to track. More orders give a tighter conversion; these cover typical budgets.
ORDERS = [1.25, 1.5, 1.75, 2, 2.5, 3, 4, 5, 6, 8, 10, 12, 16, 20, 24, 32, 48, 64, 128, 256]


def gaussian_rdp(sensitivity, sigma):
    """RDP curve of one Gaussian release: alpha * sensitivity^2 / (2 sigma^2)."""
    rho = sensitivity ** 2 / (2 * sigma ** 2)
    return [a * rho for a in ORDERS]


def to_eps(rdp, delta):
    """Convert an RDP curve to (eps, delta)-DP using the best tracked order."""
    return min(r + math.log(1 / delta) / (a - 1) for r, a in zip(rdp, ORDERS))


class BudgetExhausted(Exception):
    pass


class PrivacyLedger:
    """One ledger per (dataset, unit of privacy). Persist entries in an append-only store."""

    def __init__(self, eps_cap, delta):
        self.eps_cap, self.delta = eps_cap, delta
        self.spent = [0.0] * len(ORDERS)
        self.entries = []
        self.lock = threading.Lock()

    def charge(self, name, rdp):
        with self.lock:
            trial = [s + r for s, r in zip(self.spent, rdp)]
            eps = to_eps(trial, self.delta)
            if eps > self.eps_cap:
                raise BudgetExhausted(f"{name}: eps would reach {eps:.3f} > cap {self.eps_cap}")
            self.spent = trial              # charged before the mechanism runs; never refunded
            self.entries.append((name, rdp))
            return eps

Consumers whose mechanism is not a plain Gaussian, such as a DP-SGD run, submit the RDP curve their own accountant computed at the same orders. The ledger needs only the curve, so one service can account for dashboards, queries and training runs.

Worked example: a year of releases

Suppose a product team holds a year of LLM usage logs and policy sets a user-level cap of epsilon 4 at delta 1e-6 per year. That cap is a governance decision, not a number this article can choose for you. Two consumers want budget: a weekly dashboard of counts per region, and one DP fine-tune later in the year.

Convert the cap to rho. Solving rho + 2 sqrt(rho ln(106)) = 4 gives rho = 0.254. Allocation is easiest in rho because rho adds linearly and epsilon does not.

Give the dashboard 60%. That is rho 0.152 over 52 weekly releases, or 0.00293 a week. Contribution bounding keeps each user in at most one region per week with at most one count, so the regional counts compose in parallel and the L2 sensitivity is 1. Solving 1/(2 sigma2) = 0.00293 gives sigma = 13.07. A region with 5,000 active users gets a count accurate to about plus or minus 26 at two standard deviations; a region with 40 users gets noise that swamps the signal, so the dashboard should suppress or merge small regions.

ledger = PrivacyLedger(eps_cap=4.0, delta=1e-6)
for week in range(52):
    eps_now = ledger.charge(f"dashboard-w{week:02d}", gaussian_rdp(1.0, 13.07))
print(round(eps_now, 2))   # 3.06

Read the result carefully. Sixty per cent of rho has used 3.06 of the 4.0 epsilon, because epsilon grows with the square root of rho. The remaining 40% of rho, about 0.101, is what the fine-tune's accountant must fit inside, and the ledger will reject a run whose curve does not.

Choosing the cap

No universal epsilon exists. Choose with three inputs. First, the attack bound: at your epsilon and delta, what true-positive rate could a membership test reach at a 1% false-positive rate? Second, utility: run the actual analysis at several candidate budgets on public or synthetic data. Third, horizon: the cap must cover every release the data will ever feed, not one project. Record the decision and publish the cap and unit alongside every release.

Empirical audits complement the bound. They cannot prove privacy, but a measured attack that beats the theoretical bound proves an implementation bug. The defensive side is covered in membership inference defenses.

When the budget runs out

When the ledger refuses a charge, there are only a few honest options. Serve already-released results, which costs nothing because it is post-processing. Answer from a public or synthetic proxy. Collect new data from new users, whose budget is fresh. Or raise the cap through the same governance that set it, recorded as a change to the published guarantee.

The tempting option is renewal: reset the budget every month. If the same users' data stays in the dataset, renewal does not reset anything for them; their privacy loss keeps composing. A monthly rho of 0.1 looks like epsilon 2.45 per month, but after 12 months the same users have epsilon 9.34 at delta 1e-6, and after 36 months 17.70. Renewal is defensible only when the data itself turns over, and even then a user active every month accumulates loss. Say which applies in the published guarantee.

Failure modes

  • Unbounded contributions. Sensitivity is calibrated to one record while users contribute hundreds. The noise looks right and the user-level guarantee is far weaker.
  • Accounting the wrong sampler. The accountant assumes Poisson sampling while the loader shuffles and batches. The amplification is not real.
  • Charging after release, or refunding failures. A job that drew noise and then crashed may already have logged or cached its output. Charge first; never refund.
  • Free tuning. Each hyperparameter sweep on private data consumes budget. Account for tuning or do it on public data.
  • Floating-point noise. Naive Laplace or Gaussian sampling in floating point can leak through the low-order bits of the output, a known attack on textbook implementations. Use a vetted library's samplers.
  • Ledger per copy. The same users sit in a warehouse table, a feature store and a training snapshot, each with its own ledger. The budget must follow people, not tables.

Trade-offs

A tight cap protects users and makes small populations unusable; a loose cap keeps dashboards sharp and the guarantee hard to defend. A central ledger gives one enforceable number at the cost of a service on the critical path of every private computation. User-level accounting is the honest unit for LLM data, and its contribution bounding costs utility for heavy users. Where DP is too costly for a use case, other controls such as deduplication, access restrictions and memorisation tests still matter; see membership-inference defense architecture.

What to do next

  1. Write down the unit of privacy for each protected dataset, and add contribution bounding to the pipelines that feed it.
  2. Inventory every release, model and dashboard that touches the data, and estimate each one's RDP curve or rho.
  3. Get a cap and delta approved through governance, with delta well below 1/n, and publish both with the unit.
  4. Stand up a ledger service that charges before execution, never refunds, and keys budgets to people rather than tables.
  5. Move all accounting to RDP or zCDP and convert to (epsilon, delta) once.
  6. Check that sampling in training matches what the accountant assumes.
  7. Plan for exhaustion: cache released outputs, prepare synthetic proxies, and decide in advance how renewal and new cohorts are handled.
Key takeaway: A privacy budget is a cap on the total privacy loss that every release from a dataset may add up to, stated for a specific unit such as one user. Bound each user's contribution, compose in Renyi DP or zCDP and convert to epsilon and delta once, and enforce the cap with a ledger that charges before any noise is drawn. Remember that renewing a budget does not reset it for users whose data stays.