Reinforcement learning from human feedback is usually drawn as one arrow: a model goes in, a helpful assistant comes out. In practice it is a pipeline of four separate artifacts, each produced by its own training or labelling job, and most RLHF failures are really failures to check one artifact before building the next on top of it. This matters even more for small language models in the 0.5 to 3 billion parameter range, where there is less spare capacity to absorb a noisy reward signal and where reward hacking shows up quickly.

This article walks through the whole pipeline for a small model, stage by stage. The GPU systems side, including co-locating models and syncing weights into an inference engine, is covered in RLHF Pipeline on GPU; the tensor-level PPO step is covered in PPO for LLMs, in depth. Here the focus is the workflow and its decisions.

Advertisement

What RLHF optimises

RLHF trains a policy, the language model, to maximise a learned reward while staying close to a reference model. Written as one objective, for prompts x drawn from a prompt set and responses y sampled from the policy, it is the expected reward r(x, y) minus beta times the KL divergence between the policy and the reference. The KL term stops the policy drifting into text the reward model has never seen, where its scores are meaningless.

Two consequences follow from that objective. First, the policy can only be as good as the reward model's judgement on the policy's own outputs, so reward model quality is the ceiling. Second, the reference model is not a detail: it is usually the SFT checkpoint, and the KL penalty is measured against it on every token. A weak SFT model therefore caps RLHF twice, once as the starting point and once as the anchor.

RLHF as four artifacts and four gates1. SFT checkpointdemonstrations2. Preference setprompt, chosen, rejected3. Reward modeltext to one scalar4. PolicyPPO with KL to SFTGate Aformat, stop, evalsGate Bagreement, balanceGate Cheld-out acc, lengthGate Dwin rate, regressionsInside stage 4: one PPO iterationPolicygenerate rolloutsReference (frozen)log-probs for KLReward model (frozen)score per sequenceValue modelbaseline per tokenPer-token reward = -beta x KL(policy, reference), plus normalised RM score on the final tokenadvantages via GAE, clipped policy update, value regressionUpdated policy, next batch of prompts
The pipeline as four artifacts, each with a gate that must pass before the next stage starts, and the four models involved in one PPO iteration.

Stage 1: the SFT checkpoint

Supervised fine-tuning on demonstrations teaches the base model the chat format, the role tokens and when to stop. RLHF does not teach format well, because a reward model rarely sees malformed text and so gives it arbitrary scores. Fix format in SFT, as described in SFT, in depth.

Gate A checks that the SFT model emits the end-of-turn token reliably, follows the template on a held-out set, and has not lost too much on your capability evals compared with the base model. Also check sampling diversity: four responses per prompt at temperature 1.0 should differ, or stage 2 pairs will carry no signal.

Advertisement

Stage 2: on-policy preference data

The reward model learns from comparisons: for a prompt, a chosen response and a rejected one. People agree far more on which of two answers is better than on absolute scores.

Sample the responses from the SFT model itself, not from a stronger model. The reward model must be accurate on the distribution the policy will produce, and a dataset of GPT-class answers teaches it to separate excellent from very good while the small policy produces adequate and broken. Draw two to four responses per prompt at a temperature high enough to give real variation, then label pairs. Prompts should match production traffic in topic and length.

Gate B checks the labels. Have a sample of ten to fifteen percent labelled twice and measure inter-annotator agreement; agreement in the low 70s percent is common for general helpfulness, and it is a soft ceiling on reward model accuracy. Check the chosen side is not systematically longer, because if it is, the reward model will learn length. Drop ties or record them separately rather than forcing a choice. If an AI judge labels, also swap response order to detect position bias. DPO alignment for small models covers pair construction in more depth.

Stage 3: training the reward model

A reward model is a language model with its vocabulary head replaced by a single linear layer that outputs one number. The score for a sequence is the head's output at the final token. Training uses the Bradley-Terry model: the probability that the chosen response beats the rejected one is the sigmoid of the difference of their scores, and the loss is the negative log of that probability.

import torch
import torch.nn.functional as F

class RewardModel(torch.nn.Module):
    def __init__(self, backbone, hidden):
        super().__init__()
        self.backbone = backbone                     # transformer without LM head
        self.head = torch.nn.Linear(hidden, 1)

    def forward(self, ids, mask):
        h = self.backbone(ids, attention_mask=mask).last_hidden_state
        last = mask.sum(dim=1) - 1                   # index of final real token (right padding)
        return self.head(h[torch.arange(len(ids)), last]).squeeze(-1)

def rm_step(rm, batch, opt):
    r_chosen = rm(batch["chosen_ids"], batch["chosen_mask"])
    r_rejected = rm(batch["rejected_ids"], batch["rejected_mask"])
    loss = -F.logsigmoid(r_chosen - r_rejected).mean()
    acc = (r_chosen > r_rejected).float().mean()
    opt.zero_grad(); loss.backward(); opt.step()
    return loss.item(), acc.item()

Train for one epoch, occasionally two. Reward models overfit preference data quickly, and an overfit model is confidently wrong on the policy's new outputs.

The reward model does not have to be the same size as the policy. During PPO it only runs forward passes, so a reward model two or three times larger than a 1B policy is often affordable and noticeably more robust. The constraint is tokenisation: if the reward model uses a different tokenizer, decode the policy's tokens to text and re-tokenise, never pass token ids across.

Gate C has four checks. Held-out pairwise accuracy should approach your inter-annotator agreement; well below it means underfitting or noisy labels, and above it on a small eval set usually means leakage. Compute the correlation between score and response length on held-out data; a strong positive correlation is a warning that PPO will inflate length. Probe it with padded and confidently wrong answers. Finally record the mean and standard deviation of scores on SFT samples, because PPO will normalise rewards with them.

Stage 4: the PPO loop

Proximal policy optimisation turns the reward into gradient updates. Each iteration samples a batch of prompts, generates responses with the current policy, scores them, and then takes several small, clipped gradient steps on that batch. Four models take part: the policy being trained, the frozen reference, the frozen reward model and a value model that estimates the expected future reward at each token so that advantages have a baseline.

for it in range(num_iterations):
    prompts = sample(prompt_pool, batch_size)
    with torch.no_grad():
        resp, logp_old = policy.generate(prompts, temperature=1.0, max_new_tokens=512)
        logp_ref = reference.logprobs(prompts, resp)        # per token
        values = value_model(prompts, resp)                 # per token
        score = reward_model(prompts, resp)                 # one per sequence
    score = (score - rm_mean) / rm_std                      # stats from Gate C
    rewards = -beta * (logp_old - logp_ref)                 # KL penalty on every token
    rewards[range(batch_size), resp_last_index] += score    # task reward on the last token
    adv, returns = gae(rewards, values, gamma=1.0, lam=0.95)
    adv = (adv - adv.mean()) / (adv.std() + 1e-8)

    for epoch in range(ppo_epochs):                         # typically 1-4
        for mb in minibatches(prompts, resp, logp_old, adv, returns):
            logp = policy.logprobs(mb.prompts, mb.resp)
            ratio = torch.exp(logp - mb.logp_old)
            pg = -torch.min(ratio * mb.adv, ratio.clamp(1 - eps, 1 + eps) * mb.adv)
            vf = (value_model(mb.prompts, mb.resp) - mb.returns) ** 2
            loss = masked_mean(pg, mb.mask) + vf_coef * masked_mean(vf, mb.mask)
            loss.backward(); clip_grad_norm_(params, 1.0); opt.step(); opt.zero_grad()

Three choices dominate behaviour. Beta, the KL coefficient, trades reward against drift; values around 0.01 to 0.1 are common starting points, and an adaptive controller that adjusts beta to hold KL near a target is a sensible default. The clip range eps, commonly 0.2, bounds how far one batch can move the policy. The number of PPO epochs per batch trades sample efficiency against staleness of the rollouts.

Gate D compares the final policy with the SFT model on a held-out prompt set judged by people or a strong judge model that is not the reward model, and reruns the capability evals from Gate A. A drop there is the alignment tax. InstructGPT reduced it by mixing a pretraining loss into the PPO objective.

Sizing it for a small model

A useful rule of thumb for full fine-tuning with Adam in mixed precision is about 16 bytes per trained parameter: two for bf16 weights, two for gradients, and twelve for fp32 master weights and the two Adam moments. Frozen models in bf16 need two bytes per parameter. This ignores activations and the KV cache for generation, which can be large with long responses.

Setup (1B policy)Trained parametersRough weight and optimiser memory
Full PPO: policy, value trained; reference and 1B RM frozenabout 2Babout 32 GB + 4 GB = 36 GB
Same, with a 3B reward modelabout 2Babout 32 GB + 8 GB = 40 GB
LoRA policy and value heads on one shared basetens of millionsabout 2 GB base + RM + small optimiser state

The LoRA setup is what makes RLHF on a single GPU realistic for small models. The reference model does not need its own copy: it is the same base weights with the adapter disabled. The value model can be a second adapter or a scalar head on the shared backbone. The cost is that LoRA limits how far the policy can move, which for RLHF is often acceptable because the KL penalty limits it anyway.

Worked example: a support assistant

Consider an illustrative project: a 1.5B model fine-tuned as a customer support assistant. SFT uses 15,000 curated conversations and passes Gate A. For stage 2, 8,000 production-like prompts each get four sampled responses; labellers rank them, giving about 24,000 pairs after removing ties. Agreement on the double-labelled sample is 74 percent, and the chosen side is longer in 61 percent of pairs, which is noted as a risk.

A 3B reward model initialised from a sibling SFT checkpoint reaches 71 percent held-out accuracy, close to the agreement ceiling, but its score has a 0.45 correlation with length. Adding pairs where a concise answer beats a padded one cuts it to 0.15.

In PPO, the mean reward rises steadily for 300 iterations while KL stays under 6 nats per sequence. Around iteration 350, reward keeps rising but mean response length jumps by 40 percent and responses begin to open with apologies. Human evaluation of checkpoints at 300 and 400 prefers 300. That pattern, reward up and held-out preference flat or down, is reward hacking,; ship checkpoint 300 and retrain the reward model on fresh pairs that include the hacked outputs. These numbers are illustrative of a typical run, not results from a specific experiment.

PPO, DPO or GRPO for a small model

PPO is not the only way to use preference data. Direct preference optimisation skips the reward model and the sampling loop, and optimises the policy directly on the pairs with a loss derived from the same KL-constrained objective. Group relative policy optimisation keeps online sampling but drops the value model, using the mean reward of a group of responses to the same prompt as the baseline; it is covered in GRPO, in depth.

MethodModels in memoryDataBest fit
PPOPolicy, reference, reward, valuePrompts plus a reward modelOpen-ended quality when you can afford tuning
DPOPolicy, reference (or precomputed log-probs)Fixed preference pairsFirst alignment pass, small budgets, stable training
GRPOPolicy, reference, reward or verifierPrompts plus a scorerVerifiable tasks such as maths or code tests

For most small-model projects, a sensible order is SFT, then DPO on on-policy pairs, then PPO or GRPO only if DPO plateaus and you have a reward signal you trust.

Failure modes

  • Reward hacking. The policy finds outputs that score well and are not better: length, hedging, flattery, repeating the question. Track held-out preference from an independent judge, plus length.
  • KL blow-up. KL climbs past your target within a few iterations, usually because beta is too small, the learning rate too high, or rewards were not normalised.
  • Entropy collapse. All responses to a prompt become identical. Watch per-token entropy.
  • Distribution shift. The reward model was trained on SFT samples and the policy has moved away from them. Refresh preference data from the current policy every few rounds.
  • Masking bugs. Padding or prompt tokens included in the loss or the KL sum. Unit-test the masks.
  • Tokenizer mismatch. Token ids passed between a policy and reward model that tokenise differently produce garbage scores without errors.

Operating the pipeline

Version every artifact and record its lineage: which SFT checkpoint sampled which preference set, which reward model trained on which set, which policy used which reward model and beta. Log a few hundred rollouts per evaluation step and read them.

Keep a best-of-n baseline: sample n responses from the SFT model and return the one the reward model prefers. If PPO cannot beat best-of-4 on human evaluation, it is not paying for itself. Dashboards should show mean reward, KL, response length, entropy, value loss and clip fraction per iteration, and a held-out win rate every few dozen iterations.

What to do next

  1. Run Gate A on your SFT checkpoint now, including the sampling diversity check, before collecting any preference data.
  2. Collect a pilot of 1,000 on-policy pairs, double-label 15 percent, and measure agreement and length bias.
  3. Train a reward model for one epoch, then record held-out accuracy, length correlation and score statistics.
  4. Try DPO on the same pairs as a baseline, and a best-of-4 sampler as a second baseline.
  5. Only then run PPO with LoRA, an adaptive KL target and a held-out judge that is not your reward model, and stop at the checkpoint where held-out preference peaks.
  6. Schedule a refresh of preference data from the latest policy before starting a second RLHF round.
Key takeaway: RLHF is a pipeline of four artifacts: an SFT checkpoint, a preference dataset sampled from it, a reward model trained on that data, and a policy optimised against the reward model with a KL anchor to the SFT model. Each stage has a gate, and skipping one moves its problem downstream where it is harder to see. For small models, sample preferences on-policy, validate the reward model for length bias, use LoRA so the reference is free, start with DPO as a baseline, and judge PPO by held-out preference rather than by its own reward curve.