Pretraining teaches a language model to continue text. Supervised fine-tuning teaches it to imitate good answers. Reinforcement learning (RL) does something neither can: it lets the model try answers of its own, scores them, and pushes probability towards the ones that scored better. That is how a model learns behaviour you can judge but cannot easily write down, such as being helpful without being verbose, or solving a maths problem by any valid route rather than by copying one reference solution.
This article explains the architecture of that loop as a set of design decisions. You will choose a reward signal and reward model, pick an optimiser (PPO, GRPO or DPO), place and size the KL constraint, and set up monitors that catch the loop cheating. The GPU systems side, such as memory for several models and weight sync to the inference engine, is covered in the RLHF pipeline on GPU; here the focus is the learning problem.
The loop in one picture
Every RL method for LLMs is a variant of the same step. Take a batch of prompts. Let the current policy generate one or more completions per prompt. Score each completion. Turn scores into advantages, which say how much better or worse each completion was than expected. Raise the log-probability of tokens in better-than-expected completions and lower it for worse ones, while a penalty keeps the policy close to a frozen reference model so it does not drift into gibberish that happens to score well.
In RL language, the prompt is the state, each generated token is an action, the completion is an episode, and the reward usually arrives only once, at the end. The methods below differ mostly in how they decide which tokens deserve the credit.
Choose the signal before the algorithm
The most consequential decision is what produces the reward. The algorithm matters less than whether the reward measures what you actually want.
| Signal | How it is produced | Strengths | Weaknesses |
|---|---|---|---|
| Human preferences | Annotators compare two or more answers to the same prompt | Captures taste, tone, helpfulness | Expensive, noisy, slow to refresh; needs a reward model to generalise |
| AI feedback | A strong model judges answers against a written rubric or constitution | Cheap and fast to scale; rubric is auditable | Inherits the judge's biases, such as favouring long or confident answers |
| Verifiable rewards | A program checks the answer: unit tests pass, the final number matches, the output parses | Hard to fool if the checker is right; no reward model to drift | Only works where correctness is checkable; buggy checkers are silently exploited |
Recipes often mix them, for example verifiable rewards for code plus a reward model for chat. Log each component separately: a policy that raises the total by exploiting one component is the most common failure. Constitutional AI is a worked example of the AI-feedback route.
Reward models: architecture and loss
A reward model is usually the same transformer architecture as the policy, often initialised from the SFT checkpoint, with the language-model head replaced by a linear head that outputs one scalar. The scalar is read at the last token of the prompt-plus-response sequence, so the model sees the whole answer before scoring it.
It is trained on pairs: for a prompt x, a chosen answer y_w and a rejected answer y_l. The Bradley-Terry model says the probability that y_w is preferred is the sigmoid of the difference of their scores, so the loss is the negative log of that probability.
def reward_model_loss(rm, prompt_ids, chosen_ids, rejected_ids):
# rm returns one scalar per sequence, read at the final non-pad token
r_chosen = rm(concat(prompt_ids, chosen_ids)) # shape [batch]
r_rejected = rm(concat(prompt_ids, rejected_ids)) # shape [batch]
# Bradley-Terry: P(chosen > rejected) = sigmoid(r_chosen - r_rejected)
loss = -logsigmoid(r_chosen - r_rejected).mean()
accuracy = (r_chosen > r_rejected).float().mean() # log this
return loss, accuracyThree properties matter more than the training loss. First, only differences are learned, so the absolute scale drifts; normalise rewards per batch or subtract a running mean before using them. Second, held-out pairwise accuracy is the honest quality metric, measured also on answers from the policy you are about to train. Third, reward models pick up shortcuts such as length and list formatting; check the correlation between reward and response length on a held-out set before you start RL. The maths of the pairwise objective is derived in reward model math.
Three optimiser families
Once you have a scorer, you need a way to turn scores into weight updates. Three families dominate.
PPO (Proximal Policy Optimization, Schulman et al., 2017), as used in InstructGPT (Ouyang et al., 2022), trains a separate critic, a value model that predicts the expected reward from each token position. Advantages come from the gap between actual and predicted reward, smoothed with generalised advantage estimation. The policy update multiplies each token's advantage by the ratio of new to old probability, clipped to a band such as 1 plus or minus 0.2, so one batch cannot move the policy far.
GRPO (Group Relative Policy Optimization, introduced in the DeepSeekMath paper, Shao et al., 2024) drops the critic. It samples a group of G completions for each prompt and uses the group itself as the baseline: each completion's advantage is its reward minus the group mean, divided by the group standard deviation. The clipped ratio stays; the KL term moves into the loss.
DPO (Direct Preference Optimization, Rafailov et al., 2023) drops sampling and the reward model altogether. It shows that the optimal KL-constrained policy implies a reward of beta times the log-ratio of policy to reference, and substitutes that into the Bradley-Terry loss, so the policy is trained directly on preference pairs with a classification-style loss.
| PPO | GRPO | DPO | |
|---|---|---|---|
| Networks held during training | Policy, reference, reward model, critic | Policy, reference, scorer (or a program) | Policy, reference |
| Generates during training | Yes | Yes, G per prompt | No (offline pairs) |
| Baseline | Learned critic | Group mean and std | Implicit, via reference |
| Best fit | Dense or shaped rewards, mature infrastructure | Verifiable rewards, reasoning tasks | Fixed preference data, small budgets |
| Main risk | Critic instability, many hyperparameters | Zero-variance groups waste compute | Goes stale: data is not from the current policy |
A useful rule: if you can write a checker, start with GRPO; if you only have preference pairs, start with DPO and move to an iterative, on-policy variant (see iterative DPO) when it plateaus; reach for PPO when you need per-token shaped rewards or already run it well.
Worked example: one GRPO group by hand
Take the prompt 'What is 17 x 24?' with a verifier that returns 1 when the final answer is 408 and 0 otherwise. Sample G = 4 completions. Two get 408 and two make an arithmetic slip, so rewards are [1, 0, 0, 1].
The group mean is 0.5 and the standard deviation is 0.5. Advantages are (r - 0.5) / 0.5, giving [+1, -1, -1, +1]. Every token in completions 1 and 4 receives advantage +1; every token in 2 and 3 receives -1. The update raises the probability of the reasoning paths that ended correctly and lowers the others, without anyone labelling which step was wrong.
Now the clip. Suppose a token in completion 1 had probability 0.30 under the sampling policy and, after one gradient step within the same batch, 0.40 under the current weights. The ratio is 1.33. With a clip band of 0.2 the objective uses min(1.33 x 1, 1.2 x 1) = 1.2, so the gradient through that token stops: it has already moved as far as one batch is allowed to move it.
Finally the degenerate cases. If all four completions are correct, or all four are wrong, the standard deviation is zero and every advantage is zero (implementations add a small epsilon to avoid division by zero). That prompt cost generation time and taught nothing, which is why practitioners filter prompts towards those with mixed outcomes.
Code: a GRPO-style training step
The step below is framework-neutral pseudocode in PyTorch style. It omits batching across devices and assumes the rollout engine returns the per-token log-probabilities it sampled with.
def grpo_step(policy, ref, prompts, score_fn, G=8, clip=0.2, beta=0.04, eps=1e-6):
rollouts = generate(policy, prompts, n=G) # list of (prompt, tokens, old_logp)
rewards = tensor([score_fn(r.prompt, r.tokens) for r in rollouts]).view(-1, G)
mean = rewards.mean(dim=1, keepdim=True)
std = rewards.std(dim=1, keepdim=True, correction=0) # population std
adv = ((rewards - mean) / (std + eps)).view(-1) # one advantage per completion
keep = (std.view(-1).repeat_interleave(G) > 0) # drop zero-variance groups
for r, a in zip(compress(rollouts, keep), adv[keep]):
logp = policy.token_logprobs(r.prompt, r.tokens) # current weights
ratio = exp(logp - r.old_logp)
pg = -minimum(ratio * a, clamp(ratio, 1 - clip, 1 + clip) * a)
ref_logp = ref.token_logprobs(r.prompt, r.tokens).detach()
log_rho = ref_logp - logp
kl = exp(log_rho) - log_rho - 1 # non-negative per-token estimator
loss = (pg + beta * kl).mean() # mean over this completion's tokens
loss.backward()
optimizer.step(); optimizer.zero_grad()
log(reward_mean=rewards.mean(), kl=kl.mean(), kept_groups=keep.float().mean())Two details catch people out. The ratio must use log-probabilities from the exact weights that sampled the tokens; inference engines and trainers can compute slightly different values for the same weights, which silently biases the update. And averaging per completion versus over all tokens changes how long and short answers are weighted, and so how response length evolves.
The KL constraint: where it goes and how strong
The KL term measures how far the policy has moved from the reference model. Without it the policy finds whatever text maximises the scorer. There are two places to put it.
- In the reward. InstructGPT-style PPO subtracts beta times log(pi / pi_ref) from the reward at each token. The penalty then flows through the critic and the advantages like any other reward.
- In the loss. GRPO adds beta times a per-token KL estimate directly to the loss. This keeps the reward signal clean and makes the penalty easy to read off a dashboard.
The estimator matters. The naive log-ratio can be negative on individual tokens and is noisy. The form used in the code above, exp(log rho) - log rho - 1, is always non-negative and has lower variance; it is one of the estimators described in John Schulman's note on approximating KL divergence.
Choosing beta is an empirical exercise. Too small, and KL climbs steadily while reward rises and evaluation quality falls: classic reward hacking. Too large, and the policy barely moves. An adaptive controller, which raises beta when measured KL exceeds a target and lowers it when KL is under target, was used in early work on fine-tuning language models from human preferences and is still a good default when you do not know the right value.
DPO in practice
def dpo_loss(policy, ref, x, y_w, y_l, beta=0.1):
pw, pl = policy.seq_logprob(x, y_w), policy.seq_logprob(x, y_l)
with no_grad():
rw, rl = ref.seq_logprob(x, y_w), ref.seq_logprob(x, y_l)
margin = beta * ((pw - rw) - (pl - rl))
return -logsigmoid(margin).mean()DPO is a supervised loop: no sampling, no reward model. Its weaknesses come from that simplicity. The loss only cares about the margin, so the log-probability of the chosen answer can fall during training as long as the rejected one falls faster; watch both, not just the loss. The pairs were generated by some earlier model, so as the policy improves, the data describes mistakes it no longer makes. And DPO cannot exploit a checker directly. The derivation is in DPO math.
Failure modes
| Symptom | Likely cause | Response |
|---|---|---|
| Reward climbs, human or eval scores fall | Reward hacking: the scorer has a shortcut | Inspect top-reward samples by hand; patch the scorer; raise beta |
| Responses grow longer every step | Length bias in the reward model or loss averaging | Log the reward-length correlation; length-normalise or penalise; check loss averaging |
| Entropy collapses, samples become near-identical | Over-optimisation; too-high learning rate | Lower the learning rate; stop earlier; monitor per-group reward variance |
| Many groups with zero variance | Prompts too easy or too hard | Filter prompts by recent pass rate; rebalance the curriculum |
| KL jumps suddenly | A bad batch, or a mismatch between sampler and trainer log-probs | Compare sampler and trainer log-probs on the same tokens; skip or clip outlier batches |
| Verifier pass rate near 100% but outputs are wrong | The checker accepts malformed answers | Unit-test the verifier against adversarial outputs before training |
Operational guidance and trade-offs
- Freeze a held-out evaluation suite before the run and evaluate every N steps; the training reward is not an evaluation. LLM evaluation covers building one.
- Log reward components separately, along with KL, entropy, mean response length, the fraction of zero-variance groups and the clip fraction.
- Save a sample of completions with their scores at every evaluation step. Reading twenty of them catches hacking faster than any metric.
The trade-off is control against cost: DPO is cheap but limited by its data, GRPO is simple where answers are checkable, PPO is flexible but hardest to tune. None fixes a reward that measures the wrong thing.
What to do next
- Write down the behaviour you want and decide whether a program can check it; that decides between verifiable rewards and a learned reward model.
- If you need a reward model, train it on pairs, report held-out pairwise accuracy and the reward-length correlation, and test it on samples from your current policy.
- Pick the optimiser with the table above, and start from a small model and a few hundred prompts to debug the loop end to end.
- Instrument reward components, KL, entropy, length, clip fraction and zero-variance groups before scaling up.
- Run a beta sweep, or use an adaptive KL target, and choose the setting by held-out evaluation, not by training reward.
- Read the top-scoring and bottom-scoring completions at every checkpoint, and fix the scorer whenever the policy finds a shortcut.