An output filter is the check that runs on what a language model says before a user sees it. Input filters screen what people ask; output filters screen what the model answers, and the two catch different failures. A harmless question can produce a harmful answer, a successful jailbreak gets past every input check by definition, and retrieved documents or tool results can carry text into a response that no user typed.
This article is about the response side only: which kinds of classifier can judge a response, why the conversation around the response changes the verdict, how to set thresholds on your own traffic, what to do when the filter fires, how to stream without leaking, and how to measure what the filter misses. The wider moderation system, with its policy taxonomy, severity tiers and human review queue, is covered in LLM moderation architecture. Treating output as untrusted input to downstream code, such as HTML or SQL, is a different problem, covered in LLM output handling.
The output filter in one picture
Why the output side needs its own filter
Four routes lead to a harmful response even when the input filter is working. First, benign prompts: a question about medication doses can produce an answer that crosses into dangerous detail. Second, jailbreaks: an input filter that caught them would have blocked them, so every successful one is, by construction, an input-filter miss. Third, indirect content: a summary of a forum thread, a web page or an email can repeat slurs or threats the model was asked to process. Fourth, tool and agent output: text assembled from tool results enters the response without passing any user-input check.
Responses are also harder to classify than prompts. They are longer, often discuss harmful subjects in order to warn, quote the user, and contain code and JSON that classifiers trained on social-media text handle badly.
Two kinds of classifier
Classifiers that can judge a response fall into two families, and most production filters combine one of each.
Text-only scorers take a string and return a score per category. OpenAI's moderation endpoint is the common example: the omni-moderation-latest model is free to call, accepts text and images, and returns flagged, categories, category_scores between 0 and 1, and category_applied_input_types. Its 13 categories are harassment, harassment/threatening, hate, hate/threatening, illicit, illicit/violent, self-harm, self-harm/intent, self-harm/instructions, sexual, sexual/minors, violence and violence/graphic. Small fine-tuned encoders that you host yourself belong to the same family. They are fast and cheap, and they see only the text they are given.
Conversation-aware judges take the conversation and classify the last turn in its context. Meta's Llama Guard 4 is a 12-billion-parameter, natively multimodal model that classifies either a user prompt or a model response and generates the word safe or unsafe followed by the violated category codes. Its 14 categories, S1 to S14, are violent crimes, non-violent crimes, sex-related crimes, child sexual exploitation, defamation, specialized advice, privacy, intellectual property, indiscriminate weapons, hate, suicide and self-harm, sexual content, elections, and code interpreter abuse. Background on the Llama Guard family is in Purple Llama.
One option to avoid for new work: Jigsaw is sunsetting the Perspective API; its site says the service remains active until 31 December 2026.
Why the conversation changes the verdict
Context changes the verdict in both directions. Take the response Mixing those two products releases toxic chloramine vapours, which can injure your lungs, so never combine them. As text alone it mentions toxic vapours and injury, and a scorer may flag it as violence. With the prompt Can I clean my bathroom with bleach and an ammonia cleaner together? it is a safety warning, and blocking it harms the user.
Now reverse it. A response consisting of a numbered list of quantities and steps may look like a recipe to a text-only scorer and score near zero, while the prompt that produced it asked for something a judge would recognise as weapons help. Only a classifier that sees the request can tell the two apart. The same applies to quoting: an assistant explaining why a message it received is abusive must repeat some of it.
So score the response with the conversation whenever the classifier supports it, and treat text-only scores as a fast first signal rather than a verdict.
A two-stage filter in code
The shape that works is a cascade: a cheap scorer on every response, with low thresholds so it rarely misses, and an expensive judge only on what the scorer flags. The code below uses the verified moderation API for stage one and leaves the stage-two judge as an interface, because how you serve a 12B model, and the exact prompt template your Llama Guard version expects, belong to your stack and its model card.
from dataclasses import dataclass
from openai import OpenAI
client = OpenAI()
# Per-category thresholds, set from your own labelled outputs (see below).
THRESHOLDS = {"self-harm/instructions": 0.20, "illicit/violent": 0.30,
"sexual/minors": 0.05, "hate/threatening": 0.30}
DEFAULT_THRESHOLD = 0.60
@dataclass
class Verdict:
unsafe: bool
categories: list
stage: str
def stage1(response_text: str) -> list:
"""Text-only scorer: return the categories over their thresholds."""
result = client.moderations.create(
model="omni-moderation-latest", input=response_text).results[0]
scores = result.category_scores.model_dump(by_alias=True)
return [cat for cat, s in scores.items()
if s is not None and s >= THRESHOLDS.get(cat, DEFAULT_THRESHOLD)]
def stage2(conversation: list, response_text: str) -> Verdict:
"""Conversation-aware judge, e.g. a self-hosted Llama Guard 4.
Implement with your serving stack; parse 'safe' / 'unsafe' + codes."""
raise NotImplementedError
def check(conversation: list, response_text: str) -> Verdict:
hits = stage1(response_text)
if not hits:
return Verdict(False, [], "stage1")
return stage2(conversation, response_text)Two details matter. Thresholds are per category, because the cost of a miss differs enormously between, say, mild harassment and sexual content involving minors. And the stage-two judge overrides stage one in both directions: if it says safe, the response passes, which is how the cascade removes stage one's false positives. Libraries such as LLM Guard (now archived; its page covers running it safely) and Guardrails AI package similar scanner chains if you would rather not assemble one.
Setting thresholds on your own outputs
Scores are not probabilities, and OpenAI's own guidance is to treat them as signals for your policy and expect to recalibrate them. A score of 0.4 on your traffic may mean something quite different from 0.4 on the vendor's evaluation data, because your responses differ in length, language and subject matter.
Set thresholds from your own outputs. Sample a few thousand real responses, stratified so that each category has enough positive examples, and have them labelled against your written policy. For each category, sweep the threshold and plot recall against the rate of benign responses flagged. Choose by category: for the most severe categories, pick the threshold that reaches your recall target and accept a higher flag rate, because stage two will clean up false positives. For milder categories, favour a low flag rate. Write the chosen values, the dataset version and the date into configuration, and repeat whenever you change the generating model, the system prompt or the classifier version.
Choosing an action on a hit
A filter that only knows how to block forces a bad choice between over-blocking and under-blocking. Give it a ladder of actions and map each category and severity to a rung.
| Action | When it fits | Cost |
|---|---|---|
| Pass and log | Low-severity flags, or stage two says safe | None to the user; feeds evaluation |
| Redact a span | A quoted slur or personal detail inside an otherwise good answer | Answer may read oddly |
| Regenerate | Harm is incidental to the request; retry with a stricter system instruction | Latency; bound to one or two retries |
| Refuse with a template | The request itself seeks harmful content | User gets no answer |
| Refuse and escalate | Highest severities, or repeated attempts in a session | Human review cost |
Bound regeneration. A loop that keeps regenerating until the filter passes will eventually find wording that slips past the classifier while keeping the harmful substance, which turns your filter into an optimiser against itself. One retry, then refuse. Log every action with the scores that caused it, so you can later tell whether a refusal was a correct block or a false positive.
Streaming without leaking
Streaming makes output filtering harder, because once text has rendered it has been delivered, and retracting it afterwards does not undo the harm. The usual answer is a holdback buffer: tokens accumulate, and text is released only after the window ending at a sentence boundary has passed a check. Each check scores everything held so far, because harmful content split across fragments scores low piece by piece.
async def filtered_stream(conversation, token_stream, check_window):
held, released = "", 0
async for tok in token_stream:
held += tok
# Release only up to the last complete sentence that passed a check.
cut = max(held.rfind(". "), held.rfind("\n"))
if cut > released:
window = held[:cut + 1]
if await check_window(conversation, window):
yield held[released:cut + 1]
released = cut + 1
else:
yield "\n[Response withheld by safety filter]"
return
final = held
if await check_window(conversation, final): # full-text check at end
yield final[released:]
else:
yield "\n[Response withheld by safety filter]"The sketch is simplified: in production you would score a sliding window with a token cap rather than ever-growing text, and you would run the expensive judge only when the cheap scorer flags. The cost is latency. Sentence-level holdback adds about one sentence of delay plus one classifier call per sentence. For categories where any delivery is the harm, hold the whole response and check it once before release.
Worked example: a health-information assistant
Consider a health-information assistant answering one million questions a day. The rates below are illustrative, chosen to show the arithmetic rather than to describe any particular classifier. Suppose 0.2 percent of responses, 2,000 a day, violate policy.
With only a text-only scorer tuned for 90 percent recall, the filter catches 1,800 and misses 200. If it also flags 1 percent of the 998,000 good responses, it blocks 9,980 correct answers. Precision is 1,800 out of 11,780 flagged, about 15 percent: six of every seven blocks are wrong, and on a health assistant many of those are safety warnings, exactly the answers users most need.
Add the stage-two judge on the 11,780 flagged responses. Suppose it confirms 95 percent of true violations and wrongly confirms 10 percent of the false positives. Blocks become 1,710 true violations plus 998 false positives, so precision rises to about 63 percent and good answers blocked fall tenfold. Recall drops slightly, to 85.5 percent overall, so lower the stage-one threshold to win it back; stage two absorbs the extra flags. The judge runs on about 1.2 percent of traffic, so its cost stays small. Run this arithmetic with your own measured rates before choosing where to spend.
Measuring what the filter misses
A filter's flag rate tells you nothing about what it misses, so measure misses directly. Keep three evaluation sets. An in-distribution set, sampled from real traffic and labelled, measures everyday precision and recall. A red-team set of adversarial conversations, including the obfuscation tricks in moderation bypass, measures robustness. An over-refusal set of benign responses on sensitive subjects, such as medical warnings, history and security education, measures how much good content you block.
Report results per category and per language, never as one number, because a filter can look excellent overall while failing a category or a language entirely. In production, review a random sample of passed responses every week as well as flagged ones: passed samples are the only way to estimate the miss rate. Before switching classifiers or thresholds, run the new configuration in shadow mode beside the old one and compare decisions on the same traffic. Alert when a category's flag rate moves sharply.
Failure modes
- Language gaps. Classifiers are strongest in English; responses in other languages, or mixed scripts, are scored less reliably. Measure each language you serve.
- Obfuscation and splitting. Spaced letters, homoglyphs, encodings and harmful content spread across sentences score low. Score accumulated windows and normalise text before scoring.
- Structured output. Code, JSON and tables confuse scorers trained on prose. Extract string fields and comments and score them as text.
- Truncation. Long responses exceed the classifier's context, and the tail goes unchecked. Chunk with overlap and take the maximum score.
- Classifier outage. Decide in advance whether each category fails open or closed when the classifier times out, and alert on the timeout rate.
- Regeneration loops. Unbounded retries search for a phrasing that evades the filter. Cap retries.
- Stale thresholds. A new model version changes response style and score distributions. Recalibrate on every model change.
Trade-offs
A text-only scorer alone is cheap, but its false positives fall hardest on safety and health content. A judge on every response is accurate but adds serving cost and latency to every turn. The cascade gets most of the judge's accuracy cheaply, at the price of two systems to calibrate. Holdback protects streaming users but delays text. Write down the cost of a miss and of a false block per category, and let those numbers choose.
What to do next
- List the routes by which harmful text can reach your responses: benign prompts, jailbreaks, retrieved content and tool output.
- Put a text-only scorer on every response behind your own interface, with per-category thresholds in configuration.
- Add a conversation-aware judge on flagged responses and let its verdict override stage one.
- Label a stratified sample of your own responses and set each category's threshold from its recall and flag-rate curve.
- Define the action ladder per category, cap regeneration at one or two retries, and log scores with every action.
- For streaming, add a sentence-level holdback buffer with a final full-text check; hold whole responses for the most severe categories.
- Build in-distribution, red-team and over-refusal evaluation sets and report per category and per language.
- Review random passed responses weekly, shadow every classifier or threshold change, and recalibrate on every model change.