Group Relative Policy Optimization, GRPO, fine-tunes a language model with reinforcement learning by sampling several answers to the same prompt, scoring them, and pushing the model toward the answers that scored above their group's average. It needs no value network, which is why it became the default recipe for teaching models to reason on tasks with checkable answers. The objective, its loss, the Dr. GRPO and DAPO variants and the breakdown of GPU time and memory are covered in GRPO in depth. This page is the companion runbook: how to decide whether GRPO fits your task, prepare data so the GPUs are not wasted, write rewards that cannot be gamed, lay out the job on one GPU or a full node, size it before you rent it, and pick the checkpoint you ship.

The examples use Hugging Face TRL, whose GRPOTrainer is the most common entry point, as covered in the TRL deep dive. TRL moves quickly, so check the documentation for the version you install; defaults in particular have changed between releases.

Does GRPO fit your task?

GRPO learns only from differences within a group. If all G completions for a prompt receive the same reward, every advantage is zero and that prompt contributes nothing but compute. That single fact decides most of the planning.

  • You need a reward you can compute. Exact-match maths answers, unit tests for code, schema validation for structured output, or a checker for a game state. A learned reward model works but invites reward hacking; see reward models before going that way.
  • The starting model must sometimes succeed. If its pass rate on a prompt is zero, no sample in the group earns reward and nothing is learned. Supervised fine-tuning on a few hundred worked examples first is the usual fix when the model cannot even produce the right format.
  • You want a skill, not a style. For tone, refusals or preferences between two acceptable answers, preference methods such as DPO or KTO are cheaper because they need no generation during training.

The run, end to end

A GRPO fine-tuning run, from data to a shipped checkpointPrompts + answersverifiable tasksDifficulty filterkeep 0 < pass rate < 1Optional SFTformat warm-startReward testsunit + adversarialRollout GPUs (vLLM)G completions per promptRewardsCPU, sandboxedTrainer GPUsadvantages, loss, stepweight sync after each optimiser stepHeld-out evalpass@1, regressionsCheckpoint choicebest held-out, not best reward
The outer pipeline is done once; the inner loop of generate, score and update repeats every step, with fresh weights pushed to the rollout engine.

Most wasted GRPO runs fail before the first optimiser step: prompts too easy or too hard, a reward function with a parsing bug, or a completion budget that truncates every answer. The outer pipeline in the diagram exists to catch those problems on cheap inference before you spend training hours.

Filter prompts by pass rate

With a binary reward and a per-prompt pass rate p, the probability that a group of G samples has any variation is 1 - p**G - (1 - p)**G. With G = 8, a prompt at p = 0.5 yields a useful group 99.2 percent of the time; at p = 0.9 it is 57 percent; at p = 0.05 it is only 34 percent. A dataset dominated by prompts the model always or never solves burns most of the generation budget. Estimate pass rates once with fast inference and keep the middle.

from vllm import LLM, SamplingParams

def estimate_pass_rates(model_name, rows, check, k=16, max_tokens=1024):
    llm = LLM(model=model_name)
    params = SamplingParams(n=k, temperature=1.0, max_tokens=max_tokens)
    outs = llm.generate([r["prompt"] for r in rows], params)
    rates = []
    for row, out in zip(rows, outs):
        hits = sum(check(o.text, row["answer"]) for o in out.outputs)
        rates.append(hits / k)
    return rates

rates = estimate_pass_rates("Qwen/Qwen2.5-7B-Instruct", rows, check=numeric_match)
train_rows = [r for r, p in zip(rows, rates) if 0.0 < p < 1.0]

Sample at the temperature you will train with, and keep the estimates: as training progresses, prompts drift toward always solved, and a second filtering pass halfway through a long run recovers useful signal. Hold out a separate evaluation set before filtering so the evaluation stays honest.

Rewards you can test

A TRL reward function receives the batch of completions plus every extra dataset column as keyword arguments and returns one float per completion. Multiple functions are summed, optionally with weights. Keep each function narrow and testable.

import re

ANSWER = re.compile(r"\\boxed\{([^{}]*)\}")

def _text(c):
    return c[-1]["content"] if isinstance(c, list) else c   # chat or plain format

def numeric_match(completions, answer, **kwargs):
    scores = []
    for comp, gold in zip(completions, answer):
        found = ANSWER.findall(_text(comp))
        if len(found) != 1:                                 # zero or several answers: no credit
            scores.append(0.0)
            continue
        try:
            ok = abs(float(found[0].replace(",", "")) - float(gold)) <= 1e-6
        except ValueError:
            ok = False
        scores.append(1.0 if ok else 0.0)
    return scores

def length_guard(completions, **kwargs):
    return [-0.5 if len(_text(c)) > 6000 else 0.0 for c in completions]

# Reward unit tests run before any GPU is allocated.
assert numeric_match(["so \\boxed{1,024}"], answer=["1024"]) == [1.0]
assert numeric_match(["\\boxed{3} or \\boxed{4}"], answer=["4"]) == [0.0]

The rule that only one boxed answer earns credit is the kind of detail that matters: without it, a policy learns to list several candidates and collect reward whenever one is right. Before training, write adversarial completions by hand, such as an empty answer, the answer repeated, an answer hidden in code, and confirm each scores what you intend. Code rewards must run in a sandbox with time and memory limits, because the policy will eventually produce an infinite loop.

Laying out the job

There are two sensible layouts. On a single GPU, train LoRA adapters and run vLLM in colocate mode, sharing the device between generation and training. On a multi-GPU node, dedicate some GPUs to a vLLM server and the rest to training, which keeps both busy and lets each be configured for its own job. The LoRA mechanics are covered in LoRA fine-tuning.

from peft import LoraConfig
from trl import GRPOConfig, GRPOTrainer

config = GRPOConfig(
    output_dir="grpo-7b-lora",
    num_generations=8,                  # G
    per_device_train_batch_size=8,      # completions per device; global batch divisible by G
    gradient_accumulation_steps=8,
    max_completion_length=1536,
    temperature=1.0,
    learning_rate=1e-5,                 # LoRA tolerates higher rates than full fine-tuning
    beta=0.0,                           # set explicitly: the default has changed across versions
    loss_type="dr_grpo",
    mask_truncated_completions=True,
    bf16=True,
    use_vllm=True,
    vllm_mode="colocate",               # "server" on a multi-GPU node, see below
    logging_steps=1,
    save_steps=50,
)
trainer = GRPOTrainer(
    model="Qwen/Qwen2.5-7B-Instruct",
    reward_funcs=[numeric_match, length_guard],
    args=config,
    train_dataset=train_ds,             # columns: prompt, answer
    peft_config=LoraConfig(r=32, lora_alpha=64, target_modules="all-linear", task_type="CAUSAL_LM"),
)
trainer.train()

For server mode on an eight-GPU node, start the generation server on two GPUs and train on the other six, with vllm_mode="server" in the config:

CUDA_VISIBLE_DEVICES=0,1 trl vllm-serve --model Qwen/Qwen2.5-7B-Instruct --tensor-parallel-size 2
CUDA_VISIBLE_DEVICES=2,3,4,5,6,7 accelerate launch --num_processes 6 train_grpo.py

Full-parameter training of a 7B model with Adam needs roughly 16 bytes per parameter for weights, gradients and optimiser state, about 112 GB before activations, so the training side must shard with DeepSpeed ZeRO-3 or FSDP. With LoRA the frozen bf16 base is about 14 GB and the adapters are small, which is why a single 80 GB GPU can host both generation and training.

Sizing a step before you rent the GPUs

Size a step before renting hardware. Take 64 prompts per step with G = 8: 512 completions. If completions average 600 tokens, generation produces about 307,000 tokens per step. Measure your own vLLM throughput for the model, sequence length and GPU count; at an illustrative aggregate 10,000 generated tokens per second, generation takes about 31 seconds.

Training processes the prompt and completion tokens, say 900 per sequence, so about 461,000 tokens. A forward and backward pass costs roughly 6 FLOPs per parameter per token: 6 x 7e9 x 4.6e5 is about 1.9e16 FLOPs. Six H100-class GPUs at an assumed 40 percent of their roughly 1e15 dense bf16 FLOP/s peak deliver about 2.4e15 FLOP/s, so roughly 8 seconds, plus a forward pass for log-probabilities if importance ratios or a KL term need them. Generation dominates, which is the normal situation for GRPO: decode is memory-bandwidth-bound and long completions are expensive. The levers are therefore shorter completion budgets, fewer wasted groups through filtering, more GPUs on the rollout side, and the batching and KV-cache behaviour described in vLLM on the GPU. Weight synchronisation after each step also costs time; with LoRA only the adapters change, but the engine still needs the merged weights or adapter support, so measure it.

A run of 500 such steps is about 32,000 prompts seen and, at 40 seconds a step, under six hours of node time. Those numbers are estimates to plan with; replace every assumption with a measurement from a ten-step pilot.

What to watch during training

Log every step and read the curves together; any single metric can look healthy while the run fails.

SignalHealthyWarning sign
Mean reward per functionRises, then flattensJumps to near maximum in a few steps: suspect hacking
Groups with zero reward varianceFalls early, rises slowlyAbove about half: refilter prompts
Mean completion lengthChanges graduallyHits the cap: truncation dominates, raise cap or penalise
Clip fractionSmallLarge: learning rate or iterations per batch too high
KL to reference, if beta above 0Grows slowlyExplodes: policy collapsing to a narrow mode
Held-out pass@1Tracks training rewardFalls while training reward rises: overfitting or hacking

Have TRL log sample completions and actually read twenty of them every few hundred steps. Reward hacks are obvious to a human long before a metric shows them.

Evaluation and choosing the checkpoint

Evaluate saved checkpoints on the held-out set with fixed sampling settings, reporting pass@1 at the deployment temperature and greedy accuracy, and on a small regression suite for general ability: instruction following, a knowledge benchmark and a safety set. RL on a narrow task can degrade other skills, and LoRA reduces but does not remove that risk. Choose the checkpoint with the best held-out score that passes the regression gates, not the one with the highest training reward; late checkpoints often score higher on training prompts while generalising worse. Ship the adapter with its config, the dataset hash, the reward code version and the evaluation report, so the result can be reproduced and serving can load it as one of several task adapters on a shared base.

Report pass@k for a few values of k alongside pass@1. Several published analyses of RL on verifiable rewards found that it mainly makes the model pick its best answer more reliably, raising pass@1 sharply, while pass@k at large k, a rough measure of what the model can solve at all, moves much less and can even fall. If your use case samples many answers and verifies them, the gain may be smaller than pass@1 suggests; if it takes the first answer, pass@1 is the number that matters. Either way, compare against the base model with the same prompt template and sampling settings, because a change of template alone can move scores by several points and be mistaken for a training gain.

Failure modes

  • All-zero groups. Reward flat at zero from the first step: the model never solves anything or the parser never matches. Check reward tests and pass-rate estimates.
  • Truncation collapse. Lengths grow to the cap and reward falls; with truncated completions masked the run stalls. Lower the cap with a length penalty or raise it with more memory.
  • Reward hacking. Multiple answers, answers copied from the prompt, tests deleted in generated code. Tighten the reward and add the attack to the test suite.
  • Out-of-memory in colocate mode. vLLM's KV-cache reservation and the trainer compete; lower the vLLM memory fraction or move to server mode.
  • Stale weights. Generation keeps using an old policy because sync failed silently; compare a few log-probabilities between the engine and the trainer after sync.

Trade-offs

For outcome-reward RL, Thinking Machines' 2025 study LoRA Without Regret reported LoRA matching full fine-tuning on maths tasks even at very low rank, given a learning rate about ten times higher, so for GRPO the costs of LoRA are mostly engineering ones such as weight sync and serving, not quality. Larger G gives better advantage estimates but more generation per prompt; more prompts with smaller G gives diversity. A KL penalty limits drift at the cost of a reference model in memory and slower learning. Longer completion budgets allow longer reasoning and multiply cost. Choose by measuring on the held-out set, not by copying a recipe built for a different model.

What to do next

  1. Write the reward functions and their adversarial unit tests before touching a GPU.
  2. Estimate pass rates with vLLM at the training temperature and keep prompts strictly between 0 and 1.
  3. Run a ten-step pilot, measure generation and training time per step, and redo the sizing estimate.
  4. Start with LoRA in colocate mode on one GPU; move to server mode only when generation is the measured bottleneck.
  5. Log sample completions and read them; set alerts on zero-variance groups and length at the cap.
  6. Select the checkpoint on held-out pass@1 plus regression gates and ship it with full provenance.
  7. Compare with the human-preference pipeline in the RLHF training pipeline to see where GRPO fits among post-training methods.
Key takeaway: GRPO only learns from groups whose rewards differ, so a successful run is mostly preparation: filter prompts to those the model sometimes solves, write narrow rewards with adversarial tests, size the step knowing generation dominates, and start with LoRA on one GPU. Watch zero-variance groups, length and sample completions, and ship the checkpoint that wins on held-out data, not on training reward.