Content safety is the part of LLM engineering that decides what a system will and will not say or accept: hate and harassment, sexual content, violence, self-harm, dangerous instructions, and the categories your own product adds. It is related to security but not the same thing. Security protects the system from being subverted; content safety protects users and third parties from what the system produces, including when nobody is attacking it.

This overview gives the system view: the layers where safety is applied, the off-the-shelf classifiers, how errors compound when you stack checks, how to measure the whole pipeline including over-refusal, and how to moderate a streamed response. Deep dives on single components are linked where they matter, starting with LLM moderation architecture for taxonomy design, severity tiers and human review.

The layers of content safety

Safety is applied at several points, and each catches different failures:

  • Training. Safety fine-tuning and preference training teach the model to decline clearly harmful requests. It is the only layer that shapes reasoning itself, but you cannot update it quickly and you cannot inspect it.
  • Instructions. The system prompt states your product's policy. Cheap and immediate, but advisory: jailbreaks are exactly attacks on this layer.
  • Input checks. A classifier on the user's message blocks clearly disallowed requests before you pay for generation and routes borderline ones. See Prompt Shielding for the packing and latency details.
  • Context checks. Retrieved documents and tool output can carry harmful material the user never typed. Checking them matters most for agents that browse.
  • Output checks. A classifier on the response judges what was actually produced, in the context of the request. This is the last line and the one that catches successful jailbreaks; Output Filtering, in depth covers it end to end.
  • Review and appeals. Humans label disputed cases, which becomes both the audit trail and the next training and evaluation data.
Content safety checkpoints along one requestUser inputtext, image, fileInput checkpolicy classifierModelsafety-trainedOutput checkstreamed windowsUserRetrieved and tool textcontext checkPolicy: categories, severities, actionsone source of truth for every checkpointEval setharmful and benign, per categorySystem metricsrecall, over-blocking, latencyReview and appealslabel drift, retrainMeasure the whole chain, not each classifier: errors compound across checkpoints.
One policy feeds every checkpoint. Retrieved and tool text gets its own check, and the metrics loop at the bottom measures the chain end to end rather than each classifier alone.

Off-the-shelf classifiers

You rarely need to train a first classifier. The main families, with facts checked against their documentation:

OptionWhat it isNotes
Llama Guard 3Open-weights LLM classifier for prompts and responses14 categories S1 to S14: the 13 MLCommons hazards plus Code Interpreter Abuse; returns safe or unsafe with the violated categories
Azure AI Content SafetyHosted APIHate, sexual, violence and self-harm, scored at severity levels 0, 2, 4 and 6 (safe, low, medium, high)
Hosted moderation endpointsProvider APIs bundled with model accessCheck each provider's current category list and limits; they change
Specialised scannersToxicity, PII, prompt injection modelsNarrow and fast; each needs its own threshold calibration and a bias audit on identity terms

The choice turns on three questions. Does the taxonomy match your policy, or can you supply your own? Llama Guard accepts category definitions in its prompt, which is a strong reason teams pick it. Where may the data go? Hosted APIs mean user text leaves your boundary. What latency can you afford per checkpoint? An 8-billion-parameter classifier on every streamed window costs real GPU time. The Meta tooling is wired together in Meta Purple Llama, in depth.

Worked example: what stacking filters costs

Stacking checks feels safe, but the arithmetic deserves a look. Suppose 1,000,000 requests a day, of which 0.5 percent are genuinely harmful. Filter A catches 90 percent of harmful requests and wrongly flags 2 percent of benign ones; filter B catches 80 percent and wrongly flags 1 percent. If you block when either fires and their errors are independent:

def compose(base_rate, n, filters):
    harmful = n * base_rate
    benign = n - harmful
    miss, keep = 1.0, 1.0
    for recall, fpr in filters:
        miss *= (1 - recall)     # harmful slips past every filter
        keep *= (1 - fpr)        # benign passes every filter
    caught = harmful * (1 - miss)
    false_blocks = benign * (1 - keep)
    return dict(recall=1 - miss, fpr=1 - keep, missed=harmful - caught,
                false_blocks=false_blocks,
                precision=caught / (caught + false_blocks))

print(compose(0.005, 1_000_000, [(0.90, 0.02), (0.80, 0.01)]))

Result: combined recall 98 percent and only 100 harmful requests missed, but the false-positive rate becomes 2.98 percent, which is 29,651 benign users blocked every day, against 4,900 true catches. Precision is about 14 percent: six of every seven blocks are mistakes. Filter A alone blocked 19,900 benign users at 18 percent precision.

Two lessons follow. First, at low base rates false positives dominate the user experience, so the benign side of your evaluation matters as much as the harmful side. Second, real filters make correlated mistakes (they miss the same cleverly phrased requests), so the recall gain from stacking is usually smaller than the independent estimate while the false-positive cost is not. Measure the composed system; do not multiply vendor numbers.

Measuring the whole pipeline

An evaluation set for content safety needs two halves per category: harmful examples you must catch, and benign examples that look sensitive (a nurse asking about overdose thresholds, a history question about a massacre, a security researcher's question) that you must not block. Public sets such as XSTest exist for the second half; add cases from your own traffic. Then score the whole pipeline, not the classifier:

from collections import defaultdict

def evaluate(system, cases):
    # system(text) -> True if blocked; cases are dicts with category, text, harmful
    stats = defaultdict(lambda: {"harm": 0, "blocked_harm": 0,
                                 "benign": 0, "blocked_benign": 0})
    for case in cases:
        s = stats[case["category"]]
        blocked = system(case["text"])
        if case["harmful"]:
            s["harm"] += 1
            s["blocked_harm"] += blocked
        else:
            s["benign"] += 1
            s["blocked_benign"] += blocked
    return {cat: {"recall": s["blocked_harm"] / s["harm"] if s["harm"] else None,
                  "over_block": s["blocked_benign"] / s["benign"] if s["benign"] else None,
                  "n": s["harm"] + s["benign"]}
            for cat, s in sorted(stats.items())}

Report per category, because averages hide the failure that matters: a pipeline can look excellent overall while missing most self-harm cases written in a second language. Treat refusals by the model as blocks too, so the over-block rate includes over-cautious training, not only classifiers. Run the suite on every model, prompt or threshold change; the broader regression practice is described in LLM safety evals architecture.

Moderating a streamed response

Users expect streamed tokens, but an output classifier needs text to judge. The usual compromise holds back a window of text, checks it with some overlap so a phrase split across windows is still seen, and only then releases it:

async def moderated_stream(tokens, classify, window=40, overlap=10):
    # Hold back `window` chars; release text only after a check over it passes.
    buffer, released = "", 0
    async for tok in tokens:
        buffer += tok
        if len(buffer) - released >= window:
            if await classify(buffer[max(0, released - overlap):]):
                yield "\n[response stopped by content policy]"
                return
            yield buffer[released:]
            released = len(buffer)
    if await classify(buffer[max(0, released - overlap):]):
        yield "\n[response stopped by content policy]"
        return
    yield buffer[released:]

Tested with a toy classifier, a long response that turns bad releases its harmless opening and then stops; a short clean response is released whole at the end. Real deployments use windows of a sentence or more and classify the full response again at the end. The trade-off is unavoidable: text already released cannot be recalled, so a larger window means safer output and a slower-feeling stream. For high-severity categories, consider not streaming at all.

From verdict to action

A classifier returns a verdict; the product needs an action. Blocking is only one of them, and often the worst for users with a legitimate need. Keep the mapping from category and severity to action in configuration, versioned like code, so every checkpoint reads the same rules and an audit can say which policy version made each decision:

POLICY_VERSION = "2026-10-01"
ACTIONS = {
    # (category, severity) -> action; severity follows the classifier's scale
    ("self_harm", "low"):    "safe_complete",   # answer with care, add resources
    ("self_harm", "high"):   "support_message", # no content, show help lines
    ("violence", "low"):     "allow",           # news, history, fiction
    ("violence", "high"):    "block",
    ("weapons", "any"):      "block_and_review",
}

def decide(category: str, severity: str, stage: str) -> dict:
    action = ACTIONS.get((category, severity)) or ACTIONS.get((category, "any"), "allow")
    return {"action": action, "category": category, "severity": severity,
            "stage": stage, "policy_version": POLICY_VERSION}

"Safe completion" means the model still answers, but with a constrained prompt: for a low-severity self-harm signal it can respond to the underlying question and include support information. "Block and review" sends the case to the human queue so labels accumulate where the policy is most contested. Unknown pairs fall through to allow here; for high-risk products make the default block and enumerate what is allowed instead.

Latency belongs in the same design. An input check runs before generation, so it adds directly to time to first token; an output check on windows adds a delay per window. If each check takes 60 milliseconds and the input check is serial, users wait 60 milliseconds longer before anything appears. Running the input check in parallel with the start of generation and discarding the response on a hit hides that cost, at the price of paying for tokens you throw away.

Worked trace: one borderline request

Follow one borderline request through a pipeline built this way. A user of a health assistant writes: "What dose of paracetamol is dangerous? My teenager took several tablets an hour ago."

  1. Input check. The classifier flags self-harm at low severity: overdose language, but framed as a caregiver's question. The policy maps that pair to safe_complete, not a block.
  2. Model. The constrained prompt tells the model to put urgent guidance first. It answers that this may be an emergency, tells the user to contact emergency services or a poison control line now, and avoids giving a threshold the user might use to wait.
  3. Output check. The windowed classifier sees safety guidance, not instructions for harm, and releases each window.
  4. Logging. The decision record stores the category, severity, action, stage and policy version, with the text redacted, and samples the case into weekly review.

A naive pipeline that blocks every self-harm hit would have refused a parent in an emergency. That refusal would never show up as a safety incident, only as a benign case in your over-block metric, which is why that metric has to exist.

Operational guidance

  • Write the policy first. Classifiers implement a policy; they do not define one. Each category needs a definition, examples on both sides of the line, and an action: block, warn, route to a safer prompt, or allow.
  • Respond helpfully when blocking. A self-harm hit should lead to support resources, not a bare refusal.
  • Fail closed for high severity, open for low. If the classifier times out, decide per category whether to block or allow, and alert either way.
  • Log for audit with care. Keep decisions, scores and policy version; minimise or redact the raw text.
  • Watch drift. New slang, new languages and new product features shift both base rates and classifier accuracy. Sample blocked and allowed traffic for human review every week.
  • Layer with security. Jailbreak defences, injection scanning and tool controls belong to the same stack; LLM Defense in Depth shows how the pieces fit.

Failure modes

  • Judging output without the request. "Here are the steps" is harmless or dangerous depending on the question. Classify the pair.
  • Over-refusal creep. Every incident tightens a threshold and nobody loosens one; a year later the product refuses medical questions. Track over-blocking as a first-class metric.
  • Language and modality gaps. A classifier strong in English can be weak elsewhere, and text checks do not see images or audio.
  • Split payloads. Harmful content spread across turns or across stream windows evades per-message checks; use conversation context and overlap.
  • Single global threshold. Categories differ in harm and base rate; calibrate each separately.

What to do next

  • Write a one-page policy listing your categories, definitions and actions.
  • Build an evaluation set with harmful and borderline-benign cases per category, at least partly from your traffic.
  • Measure your current pipeline end to end with the harness above, including model refusals.
  • Pick an input and output classifier, with category definitions that match your policy, and re-measure.
  • Add windowed output moderation to streaming and decide which categories should not stream.
  • Set up weekly review of sampled blocks and passes, and feed corrections back into the evaluation set.
Key takeaway: Content safety is a pipeline, not a classifier: a written policy feeds input, context and output checks on top of a safety-trained model, and the result must be measured end to end on both harmful and borderline-benign cases. At realistic base rates false positives dominate, stacking filters multiplies them, and streamed text cannot be recalled, so measure over-blocking per category and size your output windows by severity.