Proximal Policy Optimization is the algorithm behind the original instruction-following RLHF results, and it is still the most general way to push a language model toward a learned reward: it handles dense or sparse rewards, scores from a neural reward model, and per-token shaping. It is also the most expensive and fragile post-training method: four networks, generation inside the loop, and a reward model that can be fooled. Most failed PPO projects fail outside the clipped objective: an unaudited reward model, a memory plan that did not fit, a KL penalty silently off, or a checkpoint chosen by training reward.
This page is the project plan. The mechanics of a single update are covered in PPO for LLMs in depth, and the derivations of GAE and the clipped surrogate in the PPO math walkthrough. Read those for the inside of the step; read this one before you book the GPUs.
Is PPO the right method?
Start by asking whether you need an online, critic-based method at all. Direct Preference Optimization trains on fixed preference pairs with a classification-style loss: no generation during training, no reward model at run time and two models in memory instead of four. Group-relative methods such as GRPO and RLOO generate online but replace the critic with statistics over several samples per prompt, which suits verifiable rewards such as unit tests or exact answers. PPO earns its cost when the reward is a learned model you want to keep querying on fresh samples, when responses are long enough that per-token credit assignment from a value function helps, or when you need to shape the reward inside a response rather than only at the end.
| Situation | Reasonable first choice | Why |
|---|---|---|
| Static preference pairs, small budget | DPO | Offline, two models, no rollout infrastructure |
| Verifiable reward (tests pass, answer matches) | GRPO or RLOO | No critic; group baseline is cheap and stable |
| Learned reward model, long open-ended answers | PPO | Critic gives per-token advantages; fresh samples each step |
| Reward needs intermediate shaping | PPO | Per-token rewards fit naturally into GAE |
If you cannot say what the critic buys you, run GRPO or DPO first; see the GRPO fine-tuning run plan.
The project at a glance
Tooling in 2026
Tooling moved under PPO users this year. In Hugging Face TRL, where many people first ran PPO, the top-level import of PPOTrainer stopped working in v1.10 when it moved to trl.experimental.ppo, and the trainer was then removed entirely, together with the value-head model wrappers and the PPO examples, in v1.13.0 (September 2026). The maintainers gave low usage and lack of maintenance as the reasons. Old tutorials that start with from trl import PPOTrainer will therefore fail on a current install.
The options: pin an older TRL that still ships the trainer and accept no fixes; move to a distributed online RL framework such as verl or OpenRLHF, which generate with an inference engine and train with FSDP, Megatron or DeepSpeed; or write a compact loop yourself for small models. The example uses verl keys from its quickstart; check them against your installed version.
# verl, single node, learned critic (adv_estimator=gae is the PPO path)
# Reward-model overrides are omitted: their keys live under verl/trainer/config/reward
# and have changed between releases, so copy them from your installed version.
python3 -m verl.trainer.main_ppo \
algorithm.adv_estimator=gae \
algorithm.use_kl_in_reward=True \
algorithm.kl_ctrl.kl_coef=0.02 \
data.train_files=$HOME/data/prompts/train.parquet \
data.train_batch_size=256 \
data.max_response_length=1024 \
actor_rollout_ref.model.path=/ckpt/sft-8b \
actor_rollout_ref.actor.optim.lr=1e-6 \
actor_rollout_ref.actor.ppo_mini_batch_size=64 \
actor_rollout_ref.actor.ppo_micro_batch_size_per_gpu=4 \
actor_rollout_ref.rollout.name=vllm \
actor_rollout_ref.rollout.gpu_memory_utilization=0.4 \
critic.model.path=/ckpt/rm-8b \
critic.optim.lr=1e-5 \
critic.ppo_micro_batch_size_per_gpu=4 \
trainer.n_gpus_per_node=8 trainer.nnodes=1 \
trainer.save_freq=20 trainer.test_freq=10Two details in that command are easy to miss. In the verl default configuration algorithm.use_kl_in_reward is False, so the KL coefficient you set does nothing until you switch it on; classic RLHF PPO subtracts a KL term from the reward, so turn it on deliberately and confirm it in the logs. Second, the critic is initialised from the reward model checkpoint, a common choice. How the learned reward model is plugged in differs between versions, so follow your release's configuration reference.
Prove the reward model is ready
The policy will find every weakness the reward model has, so treat the reward model as a component that must pass acceptance tests before any PPO step runs: held-out pairwise accuracy, correlation with length (which PPO will exploit within a few hundred steps), score spread (the KL coefficient only means something relative to it), and probes where filler is added to a good answer.
import numpy as np
from scipy.stats import spearmanr
def audit_reward_model(score, pairs, probes):
# pairs: list of (prompt, chosen, rejected) from a HELD-OUT collection
# probes: list of (prompt, clean, padded) where padded adds filler only
margins = np.array([score(p, c) - score(p, r) for p, c, r in pairs])
acc = float((margins > 0).mean())
texts = [(p, c) for p, c, _ in pairs] + [(p, r) for p, _, r in pairs]
scores = np.array([score(p, t) for p, t in texts])
lengths = np.array([len(t.split()) for _, t in texts])
rho = spearmanr(scores, lengths).correlation
probe_wins = np.mean([score(p, padded) > score(p, clean) for p, clean, padded in probes])
return {
"heldout_accuracy": acc, # want well above 0.5; compare with annotator agreement
"length_spearman": rho, # near 0 is healthy; 0.4+ means length is the reward
"score_std": float(scores.std()), # needed to set the KL coefficient and normalisation
"padding_win_rate": probe_wins, # should be low; high means filler is rewarded
}There is no universal pass mark; compare accuracy with how often two annotators agree on the same pairs. If length correlation is high, fix the data or add a length control before training. The reward model page covers training one.
Sizing memory and time
PPO memory is a sum of four models plus the rollout engine, and the arithmetic is worth doing on paper first. Use bytes per parameter: about 2 for bf16 weights held for inference, and about 16 for a model trained with Adam in mixed precision (bf16 weights and gradients plus fp32 master weights and two fp32 optimizer moments). Activations and the generation KV cache come on top and depend on batch and sequence length.
| Component (8B policy, 8B critic) | Trained? | Approximate memory |
|---|---|---|
| Policy (actor) | Yes, full fine-tune | 8B x 16 B = 128 GB |
| Critic (value model) | Yes | 8B x 16 B = 128 GB |
| Reference policy | No, forward only | 8B x 2 B = 16 GB |
| Reward model | No, forward only | 8B x 2 B = 16 GB |
| Total model state | about 288 GB before activations and KV cache |
Sharded with ZeRO-3 or FSDP across eight 80 GB GPUs, the trained state is roughly 32 GB per GPU and the frozen models another 4 GB, which leaves room for activations and a rollout engine sharing the cards. On fewer GPUs the levers, in order of how little they cost in quality, are: use a smaller critic or initialise it from a smaller reward model; train the policy with LoRA so the reference policy is simply the base weights with the adapter disabled, removing a full copy; offload the frozen models; and cut response length. Each lever has a price, listed in the trade-offs section below.
Generation dominates step time. Estimate it from prompts times mean response length over measured rollout tokens per second, then run a ten-step pilot and replace estimates with measurements. Colocated versus separate pools is covered in the RLHF pipeline on GPU.
Starting settings
Start from conservative, published defaults and change one thing at a time. The values below are a reasonable first run for a 1B to 8B model; treat them as a starting point, not as tuned results.
| Setting | Start | What it controls |
|---|---|---|
| Policy learning rate | 1e-6 to 3e-6 | Higher values drift from the reference quickly |
| Critic learning rate | about 10x the policy | Value fitting needs to keep up with the policy |
| Clip range | 0.2 | Size of the trust region on the probability ratio |
| PPO epochs per batch | 1 to 4 | Reuse of each rollout; more epochs raise clip fractions |
| KL coefficient | 0.01 to 0.05 | Penalty per token for leaving the reference; tune against score spread |
| Discount and GAE lambda | 1.0 and 0.95 to 1.0 | Episodes are short, so no discounting |
| Sampling temperature | 0.7 to 1.0 | Too low and there is nothing to learn from |
| Missing EOS penalty | a fixed negative score | Discourages hitting the length cap |
Log the raw score, the KL term and the combined reward separately, or you cannot tell why the reward rose.
Guarding against reward hacking
Reward hacking is the default outcome of optimising a learned reward hard enough, so build guards into the reward function and into the monitoring before you need them. The guards below are cheap and each one targets a known exploit.
def shaped_reward(prompt, response, rm_score, eos_seen, cfg):
r = rm_score
if not eos_seen: # truncated at the length cap
r -= cfg.missing_eos_penalty
n = len(response.split())
if n > cfg.soft_len: # explicit length control, not a hope
r -= cfg.len_coef * (n - cfg.soft_len) / cfg.soft_len
if cfg.banned_regex.search(response): # known exploits found in sample reviews
r -= cfg.exploit_penalty
return r # the KL term is added per token by the trainer- Score every eval checkpoint with a second, independent judge (a different reward model or a strong LLM judge with a fixed rubric) and alert when the training reward rises while the judge score falls.
- Track mean response length, the fraction of responses truncated at the cap, and the rate of repeated phrases; each tends to move before quality visibly collapses.
- Read twenty random samples at every eval interval. Numbers can be gamed; a person scanning outputs spots sycophancy, list padding and refusal drift in minutes.
- Set a KL budget for the run. If the KL to the reference passes it, stop and inspect rather than continuing on the assumption that the extra reward is real.
Worked example: a support assistant
Consider a support assistant built on an 8B SFT model. The reward model was trained on 60,000 preference pairs and reaches 71 percent held-out accuracy against an annotator agreement of 74 percent, so it is close to its ceiling. The audit shows a length correlation of 0.38, so the team adds a soft length limit at 250 words before training. A ten-step pilot on eight GPUs measures 41 seconds per step, of which 29 are generation, so they keep the colocated layout and plan 1,500 steps.
By step 600 the reward model score is still rising, but the independent judge win rate against the SFT model peaked at step 420 and is now falling, the KL to the reference has doubled since step 400, and sample reviews show answers opening with the same apology template. That is the reward model being exploited. The team selects the step 420 checkpoint, adds the template to the exploit list, and plans the next reward model round with pairs that contrast templated and direct answers. The run cost was not wasted: it produced the shipped checkpoint and the data needed to fix the reward.
Choosing the checkpoint
Choose the checkpoint on held-out evidence, never on the training reward. Plot each eval checkpoint as a point with KL to the reference on one axis and independent judge win rate on the other. The useful checkpoints lie on the upper-left frontier: the most improvement for the least drift. Among those, pick the earliest one that clears your regression suite, which should include capability benchmarks the PPO data never touched, safety and refusal tests, and formatting checks for downstream parsers. Record the reward model version, the configuration, the data snapshot and the chosen step with the artifact so the result can be reproduced or rolled back.
Failure modes
- KL never applied. A configuration flag left at its default means the policy drifts freely. Confirm a non-zero KL term in the first logged step.
- Critic lagging. Value loss stays high and explained variance stays near zero, so advantages are noise. Raise the critic learning rate, warm up the critic for a few steps before policy updates, or initialise it from the reward model.
- Generation and training disagree. The rollout engine and the trainer compute slightly different log-probabilities, so the ratio starts away from 1.0. Recompute old log-probabilities with the trainer, as described in the step-level guide.
- Stale weights in the rollout engine. A failed weight sync means the engine samples from an old policy; ratios drift and clip fractions jump. Log a weight checksum on both sides.
Trade-offs
A smaller critic saves memory but gives noisier advantages. LoRA removes a model copy and most optimizer state but limits how far the policy can move. More PPO epochs reuse rollouts but make updates more off-policy. A larger KL coefficient resists hacking and slows learning. Separate rollout pools raise throughput but add weight sync as a failure point. Measure each on held-out evaluation.
What to do next
- Write down why the critic is needed for your task; if you cannot, run DPO or GRPO first as a baseline.
- Check which TRL version your existing scripts depend on; anything importing PPO from current TRL must be pinned or ported.
- Run the reward audit: held-out accuracy, length correlation, score spread and padding probes.
- Do the memory table for your model sizes and choose the levers before renting hardware.
- Run a ten-step pilot with the KL term confirmed in logs, then size the full run from measurements.
- Set up the independent judge, length and truncation metrics, and a sample review rota.
- Choose the checkpoint from the KL versus judge frontier and archive it with full provenance.