Pretraining teaches a language model to continue text. Supervised fine-tuning teaches it to imitate good answers. Reinforcement learning (RL) does something neither can: it lets the model try answers of its own, scores them, and pushes probability towards the ones that scored better. That is how a model learns behaviour you can judge but cannot easily write down, such as being helpful without being verbose, or solving a maths problem by any valid route rather than by copying one reference solution.

This article explains the architecture of that loop as a set of design decisions. You will choose a reward signal and reward model, pick an optimiser (PPO, GRPO or DPO), place and size the KL constraint, and set up monitors that catch the loop cheating. The GPU systems side, such as memory for several models and weight sync to the inference engine, is covered in the RLHF pipeline on GPU; here the focus is the learning problem.

Advertisement

The loop in one picture

Every RL method for LLMs is a variant of the same step. Take a batch of prompts. Let the current policy generate one or more completions per prompt. Score each completion. Turn scores into advantages, which say how much better or worse each completion was than expected. Raise the log-probability of tokens in better-than-expected completions and lower it for worse ones, while a penalty keeps the policy close to a frozen reference model so it does not drift into gibberish that happens to score well.

One RL step for an LLM: sample, score, compare against a baseline, update under a KL leashPrompt setcurated, dedupedPolicy (actor)samples K answersScorerRM / verifier / judgeReference modelfrozen copyAdvantagecritic or group meanPolicy updateclipped ratio + KLMonitorsreward, KL, length, evalsCriticPPO onlybatchcompletionsrewardslog-probsvaluesmetricsnew weightsPPO keeps four networks (policy, reference, reward model, critic); GRPO drops the critic;DPO drops sampling and the reward model and learns from fixed preference pairs.
The policy samples completions, a scorer rates them, advantages compare each rating with a baseline, and the update is limited by a clipped ratio and a KL penalty against a frozen reference. Monitors watch for the policy gaming the scorer.

In RL language, the prompt is the state, each generated token is an action, the completion is an episode, and the reward usually arrives only once, at the end. The methods below differ mostly in how they decide which tokens deserve the credit.

Choose the signal before the algorithm

The most consequential decision is what produces the reward. The algorithm matters less than whether the reward measures what you actually want.

SignalHow it is producedStrengthsWeaknesses
Human preferencesAnnotators compare two or more answers to the same promptCaptures taste, tone, helpfulnessExpensive, noisy, slow to refresh; needs a reward model to generalise
AI feedbackA strong model judges answers against a written rubric or constitutionCheap and fast to scale; rubric is auditableInherits the judge's biases, such as favouring long or confident answers
Verifiable rewardsA program checks the answer: unit tests pass, the final number matches, the output parsesHard to fool if the checker is right; no reward model to driftOnly works where correctness is checkable; buggy checkers are silently exploited

Recipes often mix them, for example verifiable rewards for code plus a reward model for chat. Log each component separately: a policy that raises the total by exploiting one component is the most common failure. Constitutional AI is a worked example of the AI-feedback route.

Advertisement

Reward models: architecture and loss

A reward model is usually the same transformer architecture as the policy, often initialised from the SFT checkpoint, with the language-model head replaced by a linear head that outputs one scalar. The scalar is read at the last token of the prompt-plus-response sequence, so the model sees the whole answer before scoring it.

It is trained on pairs: for a prompt x, a chosen answer y_w and a rejected answer y_l. The Bradley-Terry model says the probability that y_w is preferred is the sigmoid of the difference of their scores, so the loss is the negative log of that probability.

def reward_model_loss(rm, prompt_ids, chosen_ids, rejected_ids):
    # rm returns one scalar per sequence, read at the final non-pad token
    r_chosen = rm(concat(prompt_ids, chosen_ids))      # shape [batch]
    r_rejected = rm(concat(prompt_ids, rejected_ids))  # shape [batch]
    # Bradley-Terry: P(chosen > rejected) = sigmoid(r_chosen - r_rejected)
    loss = -logsigmoid(r_chosen - r_rejected).mean()
    accuracy = (r_chosen > r_rejected).float().mean()  # log this
    return loss, accuracy

Three properties matter more than the training loss. First, only differences are learned, so the absolute scale drifts; normalise rewards per batch or subtract a running mean before using them. Second, held-out pairwise accuracy is the honest quality metric, measured also on answers from the policy you are about to train. Third, reward models pick up shortcuts such as length and list formatting; check the correlation between reward and response length on a held-out set before you start RL. The maths of the pairwise objective is derived in reward model math.

Three optimiser families

Once you have a scorer, you need a way to turn scores into weight updates. Three families dominate.

PPO (Proximal Policy Optimization, Schulman et al., 2017), as used in InstructGPT (Ouyang et al., 2022), trains a separate critic, a value model that predicts the expected reward from each token position. Advantages come from the gap between actual and predicted reward, smoothed with generalised advantage estimation. The policy update multiplies each token's advantage by the ratio of new to old probability, clipped to a band such as 1 plus or minus 0.2, so one batch cannot move the policy far.

GRPO (Group Relative Policy Optimization, introduced in the DeepSeekMath paper, Shao et al., 2024) drops the critic. It samples a group of G completions for each prompt and uses the group itself as the baseline: each completion's advantage is its reward minus the group mean, divided by the group standard deviation. The clipped ratio stays; the KL term moves into the loss.

DPO (Direct Preference Optimization, Rafailov et al., 2023) drops sampling and the reward model altogether. It shows that the optimal KL-constrained policy implies a reward of beta times the log-ratio of policy to reference, and substitutes that into the Bradley-Terry loss, so the policy is trained directly on preference pairs with a classification-style loss.

PPOGRPODPO
Networks held during trainingPolicy, reference, reward model, criticPolicy, reference, scorer (or a program)Policy, reference
Generates during trainingYesYes, G per promptNo (offline pairs)
BaselineLearned criticGroup mean and stdImplicit, via reference
Best fitDense or shaped rewards, mature infrastructureVerifiable rewards, reasoning tasksFixed preference data, small budgets
Main riskCritic instability, many hyperparametersZero-variance groups waste computeGoes stale: data is not from the current policy

A useful rule: if you can write a checker, start with GRPO; if you only have preference pairs, start with DPO and move to an iterative, on-policy variant (see iterative DPO) when it plateaus; reach for PPO when you need per-token shaped rewards or already run it well.

Worked example: one GRPO group by hand

Take the prompt 'What is 17 x 24?' with a verifier that returns 1 when the final answer is 408 and 0 otherwise. Sample G = 4 completions. Two get 408 and two make an arithmetic slip, so rewards are [1, 0, 0, 1].

The group mean is 0.5 and the standard deviation is 0.5. Advantages are (r - 0.5) / 0.5, giving [+1, -1, -1, +1]. Every token in completions 1 and 4 receives advantage +1; every token in 2 and 3 receives -1. The update raises the probability of the reasoning paths that ended correctly and lowers the others, without anyone labelling which step was wrong.

Now the clip. Suppose a token in completion 1 had probability 0.30 under the sampling policy and, after one gradient step within the same batch, 0.40 under the current weights. The ratio is 1.33. With a clip band of 0.2 the objective uses min(1.33 x 1, 1.2 x 1) = 1.2, so the gradient through that token stops: it has already moved as far as one batch is allowed to move it.

Finally the degenerate cases. If all four completions are correct, or all four are wrong, the standard deviation is zero and every advantage is zero (implementations add a small epsilon to avoid division by zero). That prompt cost generation time and taught nothing, which is why practitioners filter prompts towards those with mixed outcomes.

Code: a GRPO-style training step

The step below is framework-neutral pseudocode in PyTorch style. It omits batching across devices and assumes the rollout engine returns the per-token log-probabilities it sampled with.

def grpo_step(policy, ref, prompts, score_fn, G=8, clip=0.2, beta=0.04, eps=1e-6):
    rollouts = generate(policy, prompts, n=G)          # list of (prompt, tokens, old_logp)
    rewards = tensor([score_fn(r.prompt, r.tokens) for r in rollouts]).view(-1, G)

    mean = rewards.mean(dim=1, keepdim=True)
    std = rewards.std(dim=1, keepdim=True, correction=0)  # population std
    adv = ((rewards - mean) / (std + eps)).view(-1)    # one advantage per completion

    keep = (std.view(-1).repeat_interleave(G) > 0)     # drop zero-variance groups
    for r, a in zip(compress(rollouts, keep), adv[keep]):
        logp = policy.token_logprobs(r.prompt, r.tokens)          # current weights
        ratio = exp(logp - r.old_logp)
        pg = -minimum(ratio * a, clamp(ratio, 1 - clip, 1 + clip) * a)

        ref_logp = ref.token_logprobs(r.prompt, r.tokens).detach()
        log_rho = ref_logp - logp
        kl = exp(log_rho) - log_rho - 1                 # non-negative per-token estimator

        loss = (pg + beta * kl).mean()                 # mean over this completion's tokens
        loss.backward()
    optimizer.step(); optimizer.zero_grad()
    log(reward_mean=rewards.mean(), kl=kl.mean(), kept_groups=keep.float().mean())

Two details catch people out. The ratio must use log-probabilities from the exact weights that sampled the tokens; inference engines and trainers can compute slightly different values for the same weights, which silently biases the update. And averaging per completion versus over all tokens changes how long and short answers are weighted, and so how response length evolves.

The KL constraint: where it goes and how strong

The KL term measures how far the policy has moved from the reference model. Without it the policy finds whatever text maximises the scorer. There are two places to put it.

  • In the reward. InstructGPT-style PPO subtracts beta times log(pi / pi_ref) from the reward at each token. The penalty then flows through the critic and the advantages like any other reward.
  • In the loss. GRPO adds beta times a per-token KL estimate directly to the loss. This keeps the reward signal clean and makes the penalty easy to read off a dashboard.

The estimator matters. The naive log-ratio can be negative on individual tokens and is noisy. The form used in the code above, exp(log rho) - log rho - 1, is always non-negative and has lower variance; it is one of the estimators described in John Schulman's note on approximating KL divergence.

Choosing beta is an empirical exercise. Too small, and KL climbs steadily while reward rises and evaluation quality falls: classic reward hacking. Too large, and the policy barely moves. An adaptive controller, which raises beta when measured KL exceeds a target and lowers it when KL is under target, was used in early work on fine-tuning language models from human preferences and is still a good default when you do not know the right value.

DPO in practice

def dpo_loss(policy, ref, x, y_w, y_l, beta=0.1):
    pw, pl = policy.seq_logprob(x, y_w), policy.seq_logprob(x, y_l)
    with no_grad():
        rw, rl = ref.seq_logprob(x, y_w), ref.seq_logprob(x, y_l)
    margin = beta * ((pw - rw) - (pl - rl))
    return -logsigmoid(margin).mean()

DPO is a supervised loop: no sampling, no reward model. Its weaknesses come from that simplicity. The loss only cares about the margin, so the log-probability of the chosen answer can fall during training as long as the rejected one falls faster; watch both, not just the loss. The pairs were generated by some earlier model, so as the policy improves, the data describes mistakes it no longer makes. And DPO cannot exploit a checker directly. The derivation is in DPO math.

Failure modes

SymptomLikely causeResponse
Reward climbs, human or eval scores fallReward hacking: the scorer has a shortcutInspect top-reward samples by hand; patch the scorer; raise beta
Responses grow longer every stepLength bias in the reward model or loss averagingLog the reward-length correlation; length-normalise or penalise; check loss averaging
Entropy collapses, samples become near-identicalOver-optimisation; too-high learning rateLower the learning rate; stop earlier; monitor per-group reward variance
Many groups with zero variancePrompts too easy or too hardFilter prompts by recent pass rate; rebalance the curriculum
KL jumps suddenlyA bad batch, or a mismatch between sampler and trainer log-probsCompare sampler and trainer log-probs on the same tokens; skip or clip outlier batches
Verifier pass rate near 100% but outputs are wrongThe checker accepts malformed answersUnit-test the verifier against adversarial outputs before training

Operational guidance and trade-offs

  • Freeze a held-out evaluation suite before the run and evaluate every N steps; the training reward is not an evaluation. LLM evaluation covers building one.
  • Log reward components separately, along with KL, entropy, mean response length, the fraction of zero-variance groups and the clip fraction.
  • Save a sample of completions with their scores at every evaluation step. Reading twenty of them catches hacking faster than any metric.

The trade-off is control against cost: DPO is cheap but limited by its data, GRPO is simple where answers are checkable, PPO is flexible but hardest to tune. None fixes a reward that measures the wrong thing.

What to do next

  1. Write down the behaviour you want and decide whether a program can check it; that decides between verifiable rewards and a learned reward model.
  2. If you need a reward model, train it on pairs, report held-out pairwise accuracy and the reward-length correlation, and test it on samples from your current policy.
  3. Pick the optimiser with the table above, and start from a small model and a few hundred prompts to debug the loop end to end.
  4. Instrument reward components, KL, entropy, length, clip fraction and zero-variance groups before scaling up.
  5. Run a beta sweep, or use an adaptive KL target, and choose the setting by held-out evaluation, not by training reward.
  6. Read the top-scoring and bottom-scoring completions at every checkpoint, and fix the scorer whenever the policy finds a shortcut.
Key takeaway: RL for LLMs is a loop of sampling, scoring, comparing against a baseline and updating under a KL leash. The reward signal decides most of the outcome: use verifiable rewards where you can check answers, a well-tested reward model where you cannot, and log every component. Choose GRPO for checkable tasks, DPO for fixed preference data and PPO when you need shaped rewards. Tune the KL strength by held-out evaluation, and read samples constantly, because the policy will find any shortcut the scorer allows.