Every RLHF pipeline optimizes a reward model that is only a proxy for what people want. Push hard enough and the policy finds outputs the proxy scores highly that people do not actually prefer. This is Goodhart's law: when a measure becomes a target, it ceases to be a good measure. Gao, Schulman and Hilton (2022) turned it into something you can measure and fit. They found that true reward, as a function of how far the policy has moved from its starting point, follows simple curves whose coefficients scale smoothly with reward model size.

The forms themselves are covered in Reward Hacking Math and RLHF Scaling Laws. This article is the working version. It covers how to reproduce the measurement, how to estimate best-of-n curves without bias, and how to fit the coefficients to your own logs and predict where the peak is. It also gives a correction to a formula many write-ups call exact, and what to do in production, where nobody hands you a gold reward model.

The gold-versus-proxy experiment

The synthetic overoptimization experimentGold RMstands in for humansSynthetic labelsgold ranks pairsProxy RMtrained on labelsOptimizerbest-of-n or PPOrewardPolicy samplesat distance dscore with goldPlot proxy and gold vs dd = sqrt(KL(policy || initial))Proxy score keeps rising with d; gold score rises, peaks and falls. The gap is Goodhart's law, measured.
A fixed gold reward model plays the human, so the true reward of any sample is cheap to compute and overoptimization can be measured at scale.

Measuring overoptimization with real humans is expensive, so the paper replaced humans with a fixed 6B-parameter gold reward model. The gold model labelled 100,000 synthetic comparisons (90,000 for training). Proxy reward models from 3M to 3B parameters were trained on those labels, and policies, mainly 1.2B parameters, were optimized against a proxy by best-of-n sampling or by PPO. At each point the policy was scored by both models. Because the gold model defines truth here, the gap between the two curves is pure proxy error, not human noise.

The x-axis is the key design choice. Steps, samples or epochs mean different things for different optimizers, so the paper measures distance as d = sqrt(KL(pi || pi_init)), the square root of the KL divergence from the initial policy. Scores are reported relative to the initial policy, so R(0) = 0.

Two functional forms and their peaks

The two optimizers trace different curves:

best-of-n:  R_bon(d) = d * (alpha_bon - beta_bon * d)
RL (PPO):   R_rl(d)  = d * (alpha_rl  - beta_rl  * ln d)

The alphas are the initial slope: how much true reward each unit of distance buys at first. The betas are the overoptimization rate: how fast proxy error eats the gains. Setting the derivative to zero gives the peak, which is the number you actually want:

Best-of-nRL
Peak distance d*alpha / (2 beta)exp(alpha / beta - 1)
Peak gold score R*alpha^2 / (4 beta)beta * d*
Score returns to zero atd = alpha / betad = exp(alpha / beta)

For RL, the derivative of alpha d - beta d ln d is alpha - beta ln d - beta, which is zero at ln d* = alpha/beta - 1; substituting back leaves R* = beta d*. The paper notes that the RL form probably does not hold near the origin, so never fit it on the first few checkpoints. It also reports that RL is far less KL-efficient than best-of-n: it spends much more KL to reach the same gold score, and to overoptimize, so d values from the two methods are not interchangeable.

Best-of-n: the KL ruler is an upper bound

Best-of-n is attractive for measurement because it needs no training: sample n completions, keep the one the proxy likes most. The KL of the best-of-n policy from the base policy is usually quoted as log n - (n - 1)/n. Beirami et al. (2024) showed this is an upper bound, not an identity. Their counterexample: with a uniform base policy over two outputs, the true KL is log 2 - h(1/2^n), where h is binary entropy, which never exceeds log 2, while the formula grows without bound. The bound is close when the base policy rarely repeats outputs, which is typical for long free-form generations, and loose when outputs collide often, as with short answers or low temperature.

nKL bound (nats)d bound
40.640.80
161.841.35
643.181.78
2564.552.13
10245.932.44

Two lessons follow. Best-of-n moves slowly: d grows like the square root of log n, so quadrupling n buys less and less distance, which is part of why its overoptimization is gentle. And when you compare best-of-n against RL on one plot, the best-of-n points may sit too far right. Use the paper's estimator or a direct Monte Carlo KL estimate if the comparison matters.

Estimating the best-of-n curve naively, by drawing n samples per prompt for each n, is wasteful and noisy. Draw N samples once per prompt, score them with both models, and compute the exact expectation over all size-n subsets. Sort by proxy score ascending; the sample of rank i (1-based) is the subset maximum with probability C(i-1, n-1) / C(N, n):

import numpy as np
from math import comb

def bon_gold(proxy, gold, n):
    """Expected gold score of best-of-n under the proxy, from N >= n samples
    of ONE prompt. Unbiased; ignores ties in proxy score."""
    order = np.argsort(proxy)              # ascending proxy
    g = np.asarray(gold)[order]
    N = len(g)
    w = np.array([comb(i - 1, n - 1) for i in range(1, N + 1)], dtype=float)
    return float((w / comb(N, n)) @ g)

def bon_curve(prompts, ns):
    """prompts: list of (proxy_scores, gold_scores); returns mean gold per n,
    relative to n = 1 so that R(0) = 0."""
    base = np.mean([np.mean(gold) for _, gold in prompts])
    return {n: np.mean([bon_gold(p_, g_, n) for p_, g_ in prompts]) - base for n in ns}

With N = 1024 samples per prompt you get the whole curve for every n up to 1024 from one sampling pass. At n = 1 the weights are uniform and the estimate is the plain mean, which is the baseline.

Fitting the coefficients to your own runs

Both forms become linear after dividing by d: R/d = alpha - beta d for best-of-n and R/d = alpha - beta ln d for RL. Ordinary least squares is enough:

import numpy as np

def fit_overopt(d, R, kind, d_min=0.5):
    d, R = np.asarray(d, float), np.asarray(R, float)
    keep = d >= d_min                       # RL form is unreliable near d = 0
    d, R = d[keep], R[keep]
    x = d if kind == "bon" else np.log(d)
    A = np.column_stack([np.ones_like(x), -x])
    (alpha, beta), *_ = np.linalg.lstsq(A, R / d, rcond=None)
    if kind == "bon":
        d_star = alpha / (2 * beta); r_star = alpha ** 2 / (4 * beta)
    else:
        d_star = np.exp(alpha / beta - 1); r_star = beta * d_star
    return dict(alpha=alpha, beta=beta, d_star=d_star,
                kl_star=d_star ** 2, r_star=r_star)

Dividing by d inflates noise at small d, so weight points by d, or fit R directly with nonlinear least squares, once you have a first estimate. Bootstrap over prompts to get an interval on d*. A point estimate of the peak without an interval invites over-confident checkpoint choices.

Worked example: finding n and the KL budget

Suppose a best-of-n sweep against a 1B-parameter proxy gives gold-score gains of 0.64, 0.89, 0.99, 1.00 and 0.95 at n = 4, 16, 64, 256 and 1024. Using the d bounds from the table above, the fit returns roughly alpha = 1.0, beta = 0.25 (illustrative numbers, not the paper's). The predicted peak is d* = 2.0, or 4 nats of KL, with R* = 1.0. Solving log n - (n-1)/n = 4 gives n of about 150, so sampling more than about 150 candidates per prompt makes outputs worse by the gold standard, while the proxy keeps reporting gains.

For a PPO run, fit the same way on checkpoints with measured KL. Illustrative coefficients alpha = 0.6, beta = 0.25 give ln d* = 2.4 - 1 = 1.4, so d* = 4.06, a KL of about 16.4 nats, and a peak gold gain of about 1.0. The gold score returns to zero at d = exp(2.4), about 11, or 121 nats. The two numbers to remember are the budget, around 16 nats, and the margin: 16 to 121 nats is a wide plateau, so stopping somewhat late costs little, but an unregularized run that drifts to hundreds of nats throws away all the gain.

What moves the curves

LeverEffect reported by Gao et al.Practical reading
Proxy RM sizealpha and beta vary smoothly, roughly logarithmically, with parameters; for RL alpha can be held constant while beta fallsBigger reward models buy a later, higher peak, with diminishing returns
RM data sizeLittle improvement below about 2,000 comparisons; more data helps above itSmall preference sets give proxies that overoptimize almost immediately
Policy sizeLarger policies gain less from optimization but do not overoptimize more; peaks sit at similar KLA KL budget fitted on a small policy is a reasonable first guess for a larger one
KL penalty in RLGold score depended only on KL distance; the penalty made it converge earlier, like early stoppingThe penalty is a brake, not a better frontier; the authors caution this may depend on hyperparameters

The paper also maps its findings onto four kinds of Goodhart effect. Regressional: proxy noise, so the highest proxy score is partly luck. Extremal: optimization pushes into regions where the proxy was never trained. Causal: the proxy keys on a correlate, such as length, rather than the cause. Adversarial: the policy actively exploits proxy flaws. With light-tailed proxy noise, regressional error alone slows the gold gains but does not make them fall. A curve that turns over therefore points to extremal, causal or adversarial effects, or to heavy-tailed proxy error, though not necessarily to deliberate gaming.

Running it without a gold model

In production there is no gold model, only expensive human judgments. The workable substitute is a family of weaker gold signals that are independent of the training proxy:

  • A held-out reward model trained on preference data the proxy never saw, ideally of a different size or architecture, scored on every checkpoint.
  • Periodic human evaluation on a fixed prompt set at a few KL milestones, enough to place the peak, not to trace the curve.
  • Cheap symptom metrics: response length, refusal rate, repeated phrases, all of which drift when proxies are exploited.
def check_checkpoint(step, kl, proxy_gain, holdout_gain, history, slope_floor=0.0):
    """Flag the run once the held-out reward stops improving per unit of d."""
    d = kl ** 0.5
    history.append((d, proxy_gain, holdout_gain))
    if len(history) >= 4:
        (d0, _, h0), (d1, _, h1) = history[-4], history[-1]
        slope = (h1 - h0) / max(d1 - d0, 1e-6)
        if slope <= slope_floor and proxy_gain > history[-4][1]:
            return f"step {step}: proxy rising, held-out flat or falling at d={d:.2f}"
    return None

Later work points the same way. Coste et al. (2023) reported that ensembles of reward models, optimized through a worst-case or uncertainty-penalized objective, reduce overoptimization in a gold-model setup like Gao's. Rafailov et al. (2024) reported similar hump-shaped curves for direct alignment methods such as DPO, which have no explicit reward model, so dropping the reward model does not drop the problem. See DPO math for that objective.

Failure modes

  • Fitting on steps instead of KL. Coefficients from a step axis do not transfer across learning rates or batch sizes. Always log KL from the reference policy.
  • Mixing best-of-n and RL distances. The best-of-n formula overstates KL when outputs repeat, and RL spends KL less efficiently, so one shared x-axis misleads.
  • Gold leakage. If the held-out reward model shares training data with the proxy, their errors correlate and the held-out curve peaks late or never.
  • Extrapolating past the data. A fit with every point left of the peak predicts the peak from curvature alone; collect at least one point near or past it.
  • Ignoring noise at small d. The R/d transform magnifies early noise; weight or drop those points.

Trade-offs

Best-of-n needs no training and overoptimizes gently, but it costs n forward passes plus n reward model calls for every request, and its reach is capped by how slowly d grows with n. RL pays once, in training compute and engineering, and serves at the cost of a single sample, but it spends KL far less efficiently and needs a monitored stopping rule. A bigger proxy reward model moves the peak later and higher for both, at the cost of more preference data and slower scoring.

What to do next

  1. Log KL(pi || pi_ref) on every checkpoint and switch your plots to sqrt(KL) on the x-axis.
  2. Train a held-out reward model on disjoint preference data and score every checkpoint with it alongside the training proxy.
  3. Run one best-of-n sweep with N = 256 or more samples per prompt and compute the curve with the unbiased estimator above.
  4. Fit alpha and beta with bootstrap intervals, and pick n and the RL KL budget from the lower end of the d* interval.
  5. Add the checkpoint monitor to training and stop or roll back when the held-out slope turns non-positive while the proxy still rises.
  6. Re-fit whenever the reward model's size or data changes, and read Reward Model Math for the training objective behind the proxy.
Key takeaway: Overoptimization is measurable: plot true and proxy reward against d = sqrt(KL), fit R = d(alpha - beta d) for best-of-n or R = d(alpha - beta ln d) for RL, and read the peak from alpha and beta. Estimate best-of-n curves with the unbiased subset estimator, treat log n - (n-1)/n as an upper bound on its KL, and in production substitute a held-out reward model and periodic human checks for the gold model, stopping when held-out reward flattens while the proxy keeps climbing.