A chat model that declines to help with a dangerous request was not born that way. Pretrained models continue any text they are given; refusal is a behaviour added afterwards, mostly through supervised fine-tuning and reinforcement learning from human feedback (RLHF). Getting it right is harder than it looks, because there are two ways to fail. A model that refuses too little helps with things it should not. A model that refuses too much turns away a nurse asking about overdose thresholds or a developer asking how to kill a stuck process, and users learn to route around it.
This article explains how refusal behaviour is trained: the behaviour policy, the data, separate helpfulness and safety rewards, constrained optimisation, rule-based rewards, how to measure both error directions, and why the result is shallower than it looks. It assumes the basic RLHF loop; RLHF explained covers it. Examples use category placeholders rather than real harmful requests.
Start with a behaviour policy
Start with a written behaviour policy, because a reward model can only learn what labellers apply consistently. A useful policy assigns every request category one of a small number of target behaviours. OpenAI's Rule-Based Rewards paper (arXiv 2411.01111) uses three: a hard refusal for requests such as violent crime, a soft refusal for sensitive cases such as self-harm, where the model declines the harmful part with empathy and points to help, and compliance for everything else. It also specifies how a refusal should read: brief, not judgmental, not lecturing.
The policy must also say what not to refuse. Safe requests that share vocabulary with unsafe ones (killing a process, the history of a weapon, a thriller plot, drug interaction questions from a clinician) belong explicitly in the comply column. Without that, every later stage drifts towards keyword refusal, because refusing is the cheapest way to never be labelled harmful.
The training pipeline
The pipeline has five stages. Prompt collection mixes clearly harmful requests by category, borderline requests, and benign lookalikes. Supervised fine-tuning on demonstrations teaches the format of a good refusal and a good answer, which RL alone learns slowly. Preference labelling asks raters to compare two responses under a rubric that scores both helpfulness and harm. Reward modelling and RL turn those comparisons into a training signal. Evaluation measures both error directions and robustness, and its failures seed the next round. Red teaming feeds the first stage continuously.
Data: the prompt mix and the rubric
The prompt mix decides what the model learns to discriminate. If almost every prompt in the safety set is harmful, the model learns that topics are harmful, not requests. A balanced set contains, per category, harmful requests, the same topic asked for a legitimate purpose, and lookalikes that only share words. Multi-turn conversations matter too, since many real escalations are spread across turns.
Preference pairs need a rubric that separates two judgements: did the response give meaningful help with something harmful, and was it helpful for what the user legitimately needed? Anthropic's 2022 helpful-and-harmless RLHF work (Bai et al.) reported a real tension between the two objectives when they are collapsed into one preference. A rater asked "which is better?" will often prefer a refusal to a borderline but legitimate answer, and that preference, multiplied across thousands of labels, becomes over-refusal. Recording the two judgements separately keeps the tension visible and lets later stages trade them off deliberately.
Refusal style is labelled too. Two refusals can both be safe while one is preachy, accusatory or useless. Rubrics that reward a short decline plus a safe alternative (general information, a resource, a reframed version of the task) produce refusals users accept.
Turning labels into reward
There are three common ways to turn labels into reward.
One reward model trained on combined preferences is simplest and inherits the confusion above. Two reward models separate the concerns. Llama 2-Chat trained a helpfulness model and a safety model and combined them per prompt: use the safety score if the prompt is tagged as safety-relevant or the safety score is very low (below 0.15 in the paper), otherwise use the helpfulness score. Benign prompts are then judged on helpfulness, so refusing them is penalised, while risky prompts are judged on safety.
def combined_reward(prompt, response, rm_help, rm_safe, is_safety_prompt, low=0.15):
s = rm_safe(prompt, response) # probability-like safety score
h = rm_help(prompt, response)
r = s if (is_safety_prompt(prompt) or s < low) else h
return whiten(r) # normalise before PPO, as in the paperA constraint instead of a sum. Safe RLHF (Dai et al., arXiv 2310.12773, ICLR 2024) trains a reward model for helpfulness and a separate cost model for harm, then maximises reward subject to expected cost staying below a limit. The constraint is handled with a Lagrange multiplier that is learned during training: when the policy's cost is above the limit the multiplier grows and harm is penalised more; when the policy is comfortably safe it shrinks and helpfulness dominates again. That replaces a hand-tuned weight with a target.
lam = 0.0
for batch in prompts:
responses = policy.generate(batch)
r = reward_model(batch, responses) # helpfulness
c = cost_model(batch, responses) # harm, positive = worse
kl = kl_to_reference(policy, ref_policy, batch, responses)
objective = (r - lam * c) / (1 + lam) - beta * kl
ppo_step(policy, objective)
lam = max(0.0, lam + lr_lambda * (c.mean() - cost_limit)) # dual ascentRule-based rewards. OpenAI's RBR approach writes the policy as propositions about a response (for example: contains disallowed content, is judgmental, includes an apology, refers to resources), scores each with an LLM grader, and fits a small linear combination so that ideal responses rank above worse ones for each behaviour category. The result is added to the usual reward model's score during RL. The appeal is that a policy change becomes an edit to propositions and a cheap refit, not a new round of human labelling. The paper reports better balance between safety and usefulness than a human-feedback baseline on its own evaluations.
Every variant keeps the KL penalty to the reference model. Without it the policy finds whatever the reward models over-score; for refusals that is usually boilerplate disclaimers or refusing everything near a sensitive topic. How reward overoptimisation shows up generally is covered in AI Alignment, in depth. The same pairs can also train a policy directly with DPO, which drops the explicit reward model but not the need for balanced data.
Worked example: one prompt, two reward setups
Take a benign lookalike: a developer asks how to kill a process that ignores signals. Two candidate responses:
| Response | Helpfulness RM | Safety RM | Single combined RM (illustrative) |
|---|---|---|---|
| A: declines, mentions that it cannot help with harming | 0.05 | 0.98 | 0.70 |
| B: explains the signals and the forced-kill command | 0.92 | 0.95 | 0.65 |
A single reward model trained on collapsed preferences can rank A above B, because safety-flavoured words in the prompt made raters cautious and the model learned the keyword. Under the Llama 2 rule, the prompt is not tagged as safety-relevant and B's safety score is far above 0.15, so the helpfulness score decides and B wins by a wide margin. Now take a prompt in a hard-refusal category. It is tagged, so the safety score decides, and a refusal with a safe alternative beats a compliant answer even if the latter scored higher on helpfulness.
The constrained version behaves similarly but adapts. If a round of training pushes the mean cost on the safety prompt set from 0.10 to 0.30 against a limit of 0.15, the multiplier rises each step until the policy is pushed back under the limit; once cost settles at, say, 0.08, the multiplier decays and the policy is free to become more helpful on everything else. Watching the multiplier is a useful training diagnostic: steady growth means the policy cannot meet the constraint with its current data.
Measuring both error directions
Measure both directions every round, on held-out sets the training data never touched.
- Under-refusal: the rate at which the model gives meaningful help on harmful prompts, per category, judged by a calibrated classifier or rubric grader and spot-checked by humans.
- Over-refusal: the refusal rate on safe prompts that look unsafe. XSTest (Rottger et al., NAACL 2024) is a public starting point with 250 safe prompts across ten types and 200 contrasting unsafe prompts. Add lookalikes from your own domain; a medical or security product has far more legitimate sensitive questions than a general set covers.
- Refusal quality: for responses that should refuse, are they brief, non-judgmental and do they offer a safe alternative? Soft-refusal categories need this most.
- Capability regression: standard capability evaluations, since safety training that costs reasoning or coding quality will be rolled back by someone.
- Robustness: multi-turn escalation, role-play framings and paraphrases. Report how quickly refusals break, not only whether single-turn prompts are refused. LLM Jailbreak Defense Architecture covers the runtime layers that sit around a trained model.
Plot every checkpoint as a point with harmful-compliance rate on one axis and over-refusal rate on the other. Good training moves points towards the origin; most bad changes move along the trade-off curve and look like progress on whichever single metric someone happened to watch.
Why trained refusals are fragile
Three research results explain why trained refusal should be treated as one layer, not the whole defence.
Refusal is low-dimensional. Arditi et al. (NeurIPS 2024) found that in 13 open chat models up to 72B parameters, refusal is mediated by a single direction in the residual stream: removing it suppresses refusals on harmful requests, and adding it causes refusals on harmless ones. For open-weight models this means anyone with the weights can remove refusal cheaply, so release decisions cannot rely on trained refusal alone.
Alignment is shallow in token depth. Qi et al. (2024, "Safety Alignment Should Be Made More Than Just a Few Tokens Deep") showed that much of the difference between aligned and base models sits in the first few output tokens. If those tokens are forced, for example by prefilling a response, models often continue as the base model would. Their proposed mitigation is training data in which a response starts badly and then recovers into a refusal, so the decision is not concentrated at the start.
Fine-tuning erodes it. Qi et al. (2023) showed that fine-tuning aligned models on a small number of adversarial examples removes much of their safety behaviour, and that even benign fine-tuning data degrades it. Every downstream fine-tune is therefore a safety change and must re-run the safety evaluations.
Operating a refusal-trained model
- Version the policy with the model. Store the behaviour policy, rubric and proposition set alongside each checkpoint so a refusal regression can be traced to a policy change or a data change.
- Gate releases on both rates. A checkpoint that improves harmful-compliance but worsens over-refusal beyond a threshold does not ship, and vice versa.
- Measure in production. Track refusal rate by detected topic and user feedback on refusals. A rising refusal rate on a benign product surface is a real incident.
- Layer defences. Pair the trained model with input and output classifiers and tool-level permissions; trained refusal is the layer most exposed to the weaknesses above.
- Protect evaluation sets. Keep held-out safety and over-refusal sets out of every training pipeline and red-team fresh prompts each cycle.
Failure modes
- Keyword refusal. Caused by safety sets with no lookalikes and collapsed preferences. Shows up as high XSTest refusal rates.
- Disclaimer hacking. The reward model over-scores warnings, so every answer grows a paragraph of caveats. Add a proposition or rubric item that penalises unnecessary disclaimers.
- Judgmental refusals. Labellers rewarded any refusal equally; add style to the rubric.
- Multiplier runaway. In constrained training, a cost limit the data cannot meet drives the multiplier up until the policy refuses broadly. Check cost-model calibration and the limit itself.
- Inconsistent labels. Raters disagree on borderline categories, so the reward model learns noise. Measure inter-rater agreement per category and fix the policy text where it is low.
- Silent erosion. A later domain fine-tune ships without safety evaluation.
What to do next
- Write or review your behaviour policy, assigning hard refusal, soft refusal or compliance to each category, with explicit lookalikes in the comply column.
- Build two held-out sets: harmful prompts by category and safe lookalikes from your domain, and add XSTest.
- Measure your current model on both and plot it on the two-axis chart.
- If you train reward models, record helpfulness and harm judgements separately and try a two-model or constrained setup; read Reward Model Training on GPU for the training step.
- Add refusal-style criteria to the rubric and check a sample of refusals by hand.
- Add recovery examples and prefill and multi-turn tests to your robustness suite.
- Make safety evaluation a required step after every fine-tune, including benign ones.
- Pair the model with runtime defences rather than relying on trained refusal alone.