Group Relative Policy Optimization (GRPO) is the reinforcement learning method behind most open reasoning-model recipes since 2025. It was introduced in the DeepSeekMath paper (Shao et al., 2024, arXiv 2402.03300) and used at scale for DeepSeek-R1. Its idea fits in one sentence: instead of training a critic network to predict how good a response should be, sample several responses to the same prompt and judge each one against its siblings.
That idea changes the systems problem. Dropping the critic frees a model's worth of memory and a forward and backward pass, but sampling a group of long completions for every prompt makes generation, not training, the dominant cost. This article implements the objective, works through the arithmetic, shows where GPU time and memory go, and covers the fixes the community made to the original loss. For where GRPO sits among PPO, DPO and reward models, read the RL for LLMs architecture guide first.
The objective, from first principles
Policy-gradient methods push up the probability of actions with positive advantage, meaning better than expected, and push down the rest. PPO, covered in the PPO for LLMs article, estimates what was expected with a learned value network the size of the policy. GRPO replaces that estimate with a statistic computed on the spot. For each prompt q it samples G completions from the current policy, scores each with a reward function, and sets every completion's advantage to its reward minus the group mean, divided by the group standard deviation.
Every token of completion i then receives the same advantage. The loss is PPO's clipped surrogate: the ratio of new to old token probability, multiplied by the advantage, with the ratio clipped to the interval from one minus epsilon to one plus epsilon so a single batch cannot move the policy too far. The original paper adds a KL penalty towards a frozen reference model, using the estimator ref/pi minus log(ref/pi) minus one, which is always non-negative, and averages the loss first over each completion's tokens and then over the group.
Because rewards only need to rank siblings, GRPO works naturally with verifiable rewards: a unit test that passes, a maths answer that matches, a format that parses. That is why it became the default for reasoning training.
Worked example: four samples, three outcomes
Take G = 4 and a binary correctness reward. This example uses the population standard deviation; implementations differ, and PyTorch's default is the sample version, so choose one and keep it.
| Rewards | Mean | Std | Advantages | What the update does |
|---|---|---|---|---|
| 1, 0, 0, 1 | 0.5 | 0.5 | +1, -1, -1, +1 | Raises the two correct, lowers the two wrong, equally |
| 1, 0, 0, 0 | 0.25 | 0.433 | +1.73, -0.58, -0.58, -0.58 | Strong push to the single success: a rare win is informative |
| 0, 0, 0, 0 | 0 | 0 | 0, 0, 0, 0 | No gradient: the whole group was compute spent for nothing |
The second row shows why GRPO learns well on hard problems: when one sample in many succeeds, normalisation gives it a large advantage. The third row is the cost: a group in which every sample earns the same reward, all right or all wrong, contributes nothing. On a dataset that is too easy or too hard, most of your rollout compute lands in that row. Watch the TRL metric frac_reward_zero_std, which reports this fraction, and filter or reweight prompts when it climbs.
The loss in PyTorch
Here is the full loss with the three aggregation choices that matter in practice. Masks select completion tokens only; log-probabilities are per token; old_logp comes from the policy that generated the samples.
import torch
def grpo_advantages(rewards, G, scale="group", eps=1e-4):
# rewards: [B*G], grouped contiguously by prompt
r = rewards.view(-1, G)
adv = r - r.mean(dim=1, keepdim=True)
if scale == "group":
adv = adv / (r.std(dim=1, keepdim=True, unbiased=False) + eps)
elif scale == "batch":
adv = adv / (rewards.std(unbiased=False) + eps)
return adv.view(-1) # one scalar per completion
def grpo_loss(logp, old_logp, ref_logp, adv, mask,
eps_low=0.2, eps_high=0.2, beta=0.0, agg="token", max_len=None):
# logp, old_logp, ref_logp, mask: [N, T] per-token; adv: [N]
ratio = torch.exp(logp - old_logp)
a = adv.unsqueeze(1)
clipped = torch.clamp(ratio, 1 - eps_low, 1 + eps_high)
per_tok = -torch.min(ratio * a, clipped * a)
if beta > 0: # k3 estimator of KL(pi || ref)
d = ref_logp - logp
per_tok = per_tok + beta * (torch.exp(d) - d - 1)
per_tok = per_tok * mask
if agg == "sequence": # original: mean over tokens, then over samples
return (per_tok.sum(1) / mask.sum(1).clamp(min=1)).mean()
if agg == "token": # DAPO: every token in the batch weighs the same
return per_tok.sum() / mask.sum().clamp(min=1)
if agg == "constant": # Dr. GRPO: divide by a fixed length
return per_tok.sum() / (mask.shape[0] * max_len)
raise ValueError(agg)Two details are easy to get wrong. With a single optimizer step per batch of rollouts, old_logp equals the current log-probabilities detached, so the ratio is exactly one and clipping never fires; clipping only matters when you reuse rollouts for several passes. And the aggregation is not cosmetic: under the sequence mode, a token in a 100-token answer gets ten times the weight of a token in a 1,000-token answer, which, as the next section shows, biases length.
Dr. GRPO and DAPO: what changed and why
Two 2025 papers diagnosed biases in the original objective. Dr. GRPO (Liu et al., Understanding R1-Zero-Like Training) argued that dividing each completion's loss by its own length rewards short correct answers more per token and penalises long wrong answers less per token, so models drift towards long wrong answers. It also argued that dividing by the group standard deviation over-weights prompts that are nearly always solved or nearly always failed. Its fix is to divide by a constant length and to drop the standard deviation.
DAPO (Yu et al.) made four changes. Clip-higher decouples the clip range so the upper bound is wider than the lower, which lets low-probability tokens grow and slows entropy collapse. Dynamic sampling over-samples and drops groups whose rewards are all equal, so every batch carries gradient. A token-level loss weighs every token in the batch equally. Overlong reward shaping penalises truncated completions gradually rather than scoring them as plain failures. Many recipes also set the KL coefficient to zero for verifiable-reward training, which removes the reference model entirely.
| Choice | Original GRPO | Dr. GRPO | DAPO |
|---|---|---|---|
| Loss aggregation | Per-sequence mean, then mean | Sum over a constant length | Token-level over the batch |
| Std normalisation | Yes, per group | No | Yes |
| Clip range | Symmetric | Symmetric | Upper wider than lower |
| All-equal groups | Kept, zero gradient | Kept | Filtered by dynamic sampling |
| KL to reference | Yes | Not part of the fix | Dropped |
Where the GPU time goes
Count generated tokens per step: prompts times group size times mean completion length. With 64 prompts, G = 8 and completions averaging 1,000 tokens, that is 512,000 tokens generated autoregressively, one token per sequence per decoding step. Decoding is limited by memory bandwidth, reading weights and KV cache, not by arithmetic. The training pass over the same tokens is a parallel forward and backward, limited by compute. On reasoning workloads, generation commonly takes most of the step, and it grows as the model learns to think longer, so step time tends to increase during a run.
That is why a high-throughput inference engine, as described in the vLLM on GPUs article, sits inside the training loop. Two of its features matter especially here. Prefix caching lets the G samples of one prompt share the prompt's KV cache. Continuous batching keeps the GPU busy when completions finish at different lengths, which matters because a group's longest sample sets when the group is done.
Where the memory goes
Estimate training memory from bytes per parameter. Mixed-precision AdamW keeps bf16 weights and gradients (2 bytes each), an fp32 master copy (4) and two fp32 optimizer moments (8): about 16 bytes per parameter before activations. For a 7-billion-parameter policy that is roughly 112 GB, so it must be sharded across GPUs with ZeRO or FSDP. A reference model adds 2 bytes per parameter in bf16, about 14 GB at 7B, only when the KL coefficient is non-zero. PPO's critic would add another trained network of similar size, which is the saving GRPO is known for.
Generation adds its own budget: a copy of the weights in the inference engine plus KV cache for B times G sequences of the maximum length. Long completions make KV cache the swing factor. These are planning estimates; measure peak memory on a short run before booking a cluster.
Colocated or separate rollout GPUs
There are two layouts. Colocated: the inference engine runs on the same GPUs as training and the two phases take turns, freeing memory between them. It needs no extra hardware and no network weight transfer, but each phase idles while the other runs, and memory must be shared. Separate: dedicated GPUs run the inference server and the trainer sends updated weights after each step. Each side is sized for its job, but you pay for idle time on whichever side waits, plus the transfer of the full weights every step.
Choose colocated for small models and single nodes. Choose separate when the model is large, completions are long, or you want to overlap generation of the next batch with training on the current one, which makes training slightly off-policy. Either way, the sampling engine and the trainer can compute slightly different log-probabilities for the same tokens because their kernels and precision differ; TRL logs sampling/sampling_logp_difference and importance-sampling ratios so you can see the gap.
A TRL configuration with every choice explicit
The Hugging Face TRL library implements GRPO in GRPOTrainer. Its defaults have changed between releases, including the default loss type and KL coefficient, so set every value that matters and pin the TRL version. Reward functions receive the completions plus any dataset columns as keyword arguments and return one float per completion.
import re
from trl import GRPOConfig, GRPOTrainer
def correct(completions, answer, **kwargs):
out = []
for comp, gold in zip(completions, answer):
text = comp[-1]["content"] if isinstance(comp, list) else comp
m = re.search(r"<answer>(.*?)</answer>", text, re.S)
out.append(1.0 if m and m.group(1).strip() == str(gold).strip() else 0.0)
return out
def well_formed(completions, **kwargs):
texts = [c[-1]["content"] if isinstance(c, list) else c for c in completions]
return [0.2 if re.search(r"<answer>.*?</answer>\s*$", t, re.S) else 0.0 for t in texts]
config = GRPOConfig(
output_dir="grpo-math",
num_generations=8, # G: completions per prompt
per_device_train_batch_size=8, # completions per device; global batch divisible by G
max_completion_length=1024,
beta=0.0, # no reference model loaded
epsilon=0.2,
epsilon_high=0.28, # DAPO-style clip-higher
loss_type="dapo", # token-level aggregation
scale_rewards="group",
mask_truncated_completions=True,
num_iterations=1, # one gradient pass per batch of rollouts
learning_rate=1e-6,
bf16=True,
use_vllm=True,
vllm_mode="colocate",
logging_steps=1,
)
trainer = GRPOTrainer(model="Qwen/Qwen2.5-1.5B-Instruct", reward_funcs=[correct, well_formed],
args=config, train_dataset=dataset) # dataset has prompt and answer columns
trainer.train()The format reward is deliberately small so it cannot outweigh correctness; a large format reward is a classic source of hacking, where the model learns to emit tidy tags around wrong answers. Masking truncated completions removes samples that ran out of length from the loss instead of scoring them as failures.
Failure modes and the metrics that show them
| Symptom | Metric to watch | Likely cause | Response |
|---|---|---|---|
| Reward flat, no learning | frac_reward_zero_std near 1 | Prompts too easy or too hard | Filter by pass rate; raise G for hard prompts |
| Length grows without accuracy | completions/mean_length, completions/clipped_ratio | Length bias, truncation scored as failure | Token-level or constant aggregation; mask truncation |
| Reward rises, evals fall | held-out accuracy vs reward | Reward hacking | Harden verifiers; inspect samples |
| Outputs become repetitive | entropy collapsing | Clip range too tight, learning rate too high | Clip-higher; lower learning rate |
| Training unstable after layout change | sampling log-prob difference | Engine and trainer disagree | Match dtypes; correct with importance ratios |
Reward hacking deserves its own reading: the reward hacking maths article explains why optimisation pressure finds every gap in a verifier. Read a sample of completions every few hundred steps, not just the curves.
What to do next
- Build verifiable reward functions first and test them on known good and bad answers; most GRPO failures start in the reward.
- Measure the pass rate of your current model on each training prompt, and keep prompts where it is neither 0 nor 1 for a group of your chosen size.
- Estimate generated tokens per step and memory per GPU with the formulas above, then confirm on a 20-step run.
- Start colocated with an explicit, pinned TRL configuration, token-level loss, KL off and truncation masked.
- Dashboard
frac_reward_zero_std, mean and clipped length, entropy, reward and a held-out evaluation. - Move to separate rollout GPUs only when generation time is clearly the bottleneck and the model is large enough to justify the weight transfer.