TRL (Transformer Reinforcement Learning) is Hugging Face's post-training library. It began as a PPO implementation for language models and is now the default open-source toolkit for supervised fine-tuning, preference optimisation and reinforcement learning with verifiable rewards. Its v1.0 release in April 2026 split the API into a stable tier (SFT, DPO, reward modelling, RLOO and GRPO, imported from trl) and an experimental tier under trl.experimental for methods such as KTO and ORPO.

This article is about the library rather than the algorithms: how its trainers are built, what data each expects, how it uses GPU memory, how online methods use vLLM, and how to scale a run from one GPU to many. It finishes with a worked SFT-then-DPO pipeline, failure modes and a checklist. TRL moves fast, so every hyperparameter in the code below is set explicitly; run trl env and compare against the docs for your installed version.

How TRL is put together

Datasettext / prompt-completion / pairsProcessing classtokenizer + chat templateXxxConfigTrainingArguments + method knobsXxxTrainerSFT / DPO / Reward / GRPO / RLOOtransformers Trainerloop, checkpoints, loggingPEFTLoRA / QLoRA adaptersvLLMfast generation for online RLAccelerateDDP / FSDP / DeepSpeed ZeROmap + tokenizesubclass ofpeft_configuse_vllmlaunchOffline methods (SFT, DPO, reward) read fixed data; online methods (GRPO, RLOO) generate completions every step,so generation speed, not the backward pass, usually decides their throughput.
Every method is a config plus a trainer built on the transformers Trainer, with PEFT, vLLM and Accelerate plugged in.

Configs and trainers

Every TRL method is a pair: a config dataclass and a trainer. SFTConfig, DPOConfig, RewardConfig and GRPOConfig extend transformers' TrainingArguments, so batch size, learning rate, gradient accumulation, bf16, checkpointing and report_to work exactly as in the base Trainer. Each trainer subclasses Trainer and overrides three things: dataset preparation (applying the chat template and tokenising), the collator, and compute_loss.

Two defaults differ from TrainingArguments and catch people out: gradient checkpointing is on, and bf16 is on unless you set fp16. Learning rates also differ by method (2e-5 for SFT and 1e-6 for DPO in current docs), because preference updates on an already-tuned model need far smaller steps. You can pass a model id string and TRL loads it, with model_init_kwargs forwarded to from_pretrained; set the dtype there, since TRL loads in float32 when you do not.

Data formats decide most outcomes

Most TRL bugs are data-format bugs. The trainers accept a small set of shapes, in standard (plain strings) or conversational (lists of role/content messages) form. With conversational data TRL applies the tokenizer's chat template for you.

TrainerExpected columnsLoss is computed on
SFT, language modellingtext or messagesAll tokens, or assistant turns with assistant_only_loss=True
SFT, prompt-completionprompt, completionCompletion only by default
DPOprompt, chosen, rejectedLog-prob ratio of chosen vs rejected
Rewardchosen, rejected (prompt optional)Pairwise ranking of scalar scores
GRPO / RLOOprompt plus any extra columnsGenerated completions, weighted by advantage

Assistant-only loss depends on the chat template marking generated spans with generation tags. TRL patches the template for some known families (Qwen3 among them); for others, confirm the template contains them, or the mask silently covers nothing useful. In GRPO any extra dataset column, such as a reference answer, is passed to your reward function as a keyword argument.

Planning GPU memory

Plan memory before you pick a method. Full fine-tuning with AdamW in mixed precision costs roughly 16 bytes per parameter before activations: bf16 weights and gradients (4 bytes) plus fp32 master weights and two Adam moments (12 bytes). A 7B model therefore needs around 112 GB, which means sharding across GPUs with FSDP or ZeRO-3. LoRA freezes the base, so the bf16 base is 14 GB and the optimiser state covers only the adapters. QLoRA stores the base in 4-bit, about 4 GB for 7B, and fits on a single 24 GB card with room for activations.

Method changes the bill. DPO needs reference-model log-probabilities from the policy as it was before DPO began. When you train a fresh LoRA adapter on a frozen base, that base already is the reference, so no second full copy is needed; for full fine-tuning the reference is a frozen copy unless you set precompute_ref_log_probs=True, which runs the reference once over the dataset and frees it. GRPO in current docs defaults to beta=0.0, in which case no reference model is loaded at all; any positive beta brings it back. Packing (packing=True) removes padding waste in SFT, and with the default best-fit-decreasing strategy it runs padding-free, which needs FlashAttention.

Supervised fine-tuning with QLoRA

The SFT trainer is the workhorse. The script below fine-tunes an 8B model on conversational support transcripts with QLoRA on one GPU, training only on assistant turns.

import torch
from datasets import load_dataset
from peft import LoraConfig
from transformers import BitsAndBytesConfig
from trl import SFTConfig, SFTTrainer

ds = load_dataset("json", data_files="support_chats.jsonl", split="train")  # {"messages": [...]}
split = ds.train_test_split(test_size=0.02, seed=0)

args = SFTConfig(
    output_dir="ckpt/sft",
    model_init_kwargs={"dtype": torch.bfloat16, "attn_implementation": "flash_attention_2"},
    max_length=4096,
    packing=True,                       # bfd packing, padding-free (needs FlashAttention)
    assistant_only_loss=True,           # needs generation tags in the chat template
    per_device_train_batch_size=4,
    gradient_accumulation_steps=8,
    learning_rate=1e-4,                 # adapters tolerate higher LR than full fine-tuning
    num_train_epochs=2,
    eval_strategy="steps", eval_steps=200,
    save_steps=200, save_total_limit=3,
    report_to="tensorboard",
)
trainer = SFTTrainer(
    model="Qwen/Qwen3-8B",
    args=args,
    train_dataset=split["train"],
    eval_dataset=split["test"],
    peft_config=LoraConfig(r=16, lora_alpha=32, target_modules="all-linear", task_type="CAUSAL_LM"),
    quantization_config=BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
                                           bnb_4bit_compute_dtype=torch.bfloat16),
)
trainer.train()
trainer.save_model("ckpt/sft/final")

# Merge the adapter into a bf16 copy of the base so later stages start from the SFT model.
from peft import AutoPeftModelForCausalLM
merged = AutoPeftModelForCausalLM.from_pretrained("ckpt/sft/final", dtype=torch.bfloat16).merge_and_unload()
merged.save_pretrained("ckpt/sft-merged")
trainer.processing_class.save_pretrained("ckpt/sft-merged")

Watch mean_token_accuracy and the eval loss together: accuracy that keeps climbing while eval loss turns upward is memorisation. For the algorithmic detail behind this step see supervised fine-tuning explained.

Preference optimisation with DPO

DPO starts from the merged SFT model with a fresh LoRA adapter. That ordering matters: if you instead kept training the SFT adapter and the trainer took the adapter-free base as its reference, DPO would be anchored to the pre-SFT model, a common mistake when chaining LoRA stages. With a merged starting point, the frozen weights are the SFT model whichever way the reference is built.

import torch
from datasets import load_dataset
from peft import LoraConfig
from trl import DPOConfig, DPOTrainer

pairs = load_dataset("json", data_files="support_prefs.jsonl", split="train")  # prompt/chosen/rejected

args = DPOConfig(
    output_dir="ckpt/dpo",
    beta=0.1,                    # higher = stay closer to the SFT model
    loss_type=["sigmoid"],
    learning_rate=5e-6,
    max_length=4096,
    per_device_train_batch_size=2,
    gradient_accumulation_steps=16,
    num_train_epochs=1,
    logging_steps=10,
    model_init_kwargs={"dtype": torch.bfloat16},
)
DPOTrainer(model="ckpt/sft-merged", args=args, train_dataset=pairs,
           peft_config=LoraConfig(r=16, lora_alpha=32, target_modules="all-linear",
                                  task_type="CAUSAL_LM")).train()

The logs tell you whether it works. rewards/margins should rise steadily and rewards/accuracies should climb well above 0.5. If logps/chosen falls as fast as logps/rejected, the model is lowering the probability of everything, often a sign the learning rate is too high or the pairs are too similar. The loss and its variants are covered in DPO training on GPUs; if you need an explicit scorer instead, see training reward models.

Online RL with GRPO and vLLM

Online methods generate their own training data. For each prompt GRPO samples num_generations completions, scores them with your reward functions, normalises rewards within the group to get advantages, and takes a clipped policy-gradient step. A reward function receives the prompts, the completions and any extra dataset columns, and returns one float per completion:

import re
from trl import GRPOConfig, GRPOTrainer

def exact_answer(prompts, completions, answer, **kwargs):
    """1.0 if the last \boxed{...} matches the reference column 'answer', else 0.0."""
    out = []
    for comp, ref in zip(completions, answer):
        text = comp[-1]["content"] if isinstance(comp, list) else comp
        m = re.findall(r"\\boxed\{([^}]*)\}", text)
        out.append(1.0 if m and m[-1].strip() == ref.strip() else 0.0)
    return out

args = GRPOConfig(
    output_dir="ckpt/grpo",
    num_generations=8,              # group size per prompt
    max_completion_length=1024,
    beta=0.0,                       # no KL term, no reference model
    loss_type="dapo",
    use_vllm=True, vllm_mode="colocate", vllm_gpu_memory_utilization=0.3,
    per_device_train_batch_size=8,
    gradient_accumulation_steps=4,
    learning_rate=1e-6,
)
GRPOTrainer(model="ckpt/sft-merged", args=args, reward_funcs=[exact_answer],
            train_dataset=math_prompts).train()

Generation is the bottleneck, so TRL can hand it to vLLM in two modes. In colocate mode vLLM runs inside the training processes and shares their GPUs, with vllm_gpu_memory_utilization capping its share; it is simple and suits a single node. In server mode a separate vLLM server owns dedicated GPUs and the trainer pushes updated weights to it after each step; it costs GPUs but lets generation and training overlap at larger scale. How the server is launched has changed between versions (the old trl vllm-serve command is deprecated in favour of vllm serve), so follow the vLLM integration page for your release. TRL also checks that the generation batch divides evenly into groups of num_generations. The algorithm itself is covered in GRPO on GPUs.

Scaling out with the CLI and Accelerate

TRL leans on Accelerate for distribution, and its CLI wraps both. A YAML config captures every argument, and --accelerate_config takes a built-in profile name (fsdp1, fsdp2, zero1, zero2, zero3, multi_gpu, single_gpu) or a path to your own file:

# dpo_full.yaml  --  launched with:  trl dpo --config dpo_full.yaml
model_name_or_path: ckpt/sft-merged
dataset_name: my-org/support-prefs
beta: 0.1
learning_rate: 5.0e-7
max_length: 4096
per_device_train_batch_size: 1
gradient_accumulation_steps: 16
precompute_ref_log_probs: true      # avoid holding a second full model
accelerate_config: zero3
num_processes: 8

Use LoRA plus DDP when the base fits on one GPU, and FSDP or ZeRO-3 when full fine-tuning needs to shard parameters, gradients and optimiser state. FSDP is the native PyTorch route and pairs well with recent transformers; the trade-offs are discussed in FSDP for large-model training.

Worked example: a support assistant

A team wants a support assistant on Qwen3-8B. They have 40,000 resolved support chats and 6,000 preference pairs where agents picked the better of two drafts. Step one is QLoRA SFT on a single 80 GB GPU using the script above: two epochs over about 60 million tokens take most of a working day, and assistant-only loss keeps the model from learning to imitate customers. Step two merges the adapter and runs DPO with a new adapter at beta 0.1. After one epoch rewards/accuracies sits around 0.7 on held-out pairs and a blind agent review prefers the DPO model.

They then try full-parameter DPO with ZeRO-3 on eight GPUs, using the YAML above. It wins only slightly on review and costs eight times the GPU hours, so they keep the adapter, which also lets them serve several customer-specific adapters on one base. Before release they merge the adapter for a latency test and run the same evaluation suite on both forms, because merging after 4-bit training can shift outputs slightly.

Failure modes

  • Wrong or missing chat template. The base model has no template or a different one at inference. Set chat_template_path or the EOS token explicitly and render a sample before training.
  • Truncation eats the answer. max_length defaults to 1024 tokens; longer conversations lose their final turn. Measure your token-length distribution first.
  • Version drift. Parameter names and defaults change between releases. Pin trl, transformers, peft and vLLM together in a lockfile.
  • Reward hacking in GRPO. The model finds formats that satisfy a lax regex. Inspect samples every few hundred steps and combine correctness with format rewards.
  • Colocated vLLM out of memory. Generation and training compete for the same GPU. Lower vllm_gpu_memory_utilization, shorten completions, or move to server mode.

Trade-offs

TRL trades peak efficiency for breadth and readability: specialised RL frameworks can schedule generation and training more aggressively at very large scale, but TRL covers more methods with code you can read in an afternoon and plugs into the Hugging Face ecosystem. Experimental trainers carry no stability promise, so prefer the stable tier for anything you must reproduce next year. Offline methods are cheap and predictable; online methods cost more but improve where a verifiable reward exists.

A practical order of choice: start with SFT on good demonstrations, because nothing downstream fixes bad imitation data. Add DPO when you can collect pairwise preferences cheaply, for example from reviewers choosing between two drafts. Reach for GRPO or RLOO only when correctness can be checked by code, such as maths answers, unit tests or schema validation, and budget for generation to dominate the GPU hours.

What to do next

  1. Pin trl, transformers, peft, accelerate and vLLM versions, and save trl env output with every run.
  2. Convert your data into one of the supported formats and render five samples through the chat template by hand.
  3. Measure token lengths and set max_length from the 99th percentile.
  4. Run QLoRA SFT on a small slice first and confirm the loss mask covers only the assistant turns.
  5. Add DPO only after SFT is stable, and watch rewards/margins and logps/chosen.
  6. For GRPO, write and unit-test reward functions before training, and start with colocated vLLM on one node.
  7. Move hyperparameters into a YAML config and launch with the CLI so every run is reproducible.
Key takeaway: TRL wraps every post-training method as a config plus a trainer on top of the transformers Trainer, so the skills transfer between SFT, DPO and GRPO. Get the data format and chat template right, plan memory by method, use PEFT unless full fine-tuning clearly pays, hand generation to vLLM for online RL, and pin versions because defaults move.