A toxicity scanner is a text classifier that scores abusive language: insults, threats, identity attacks, obscenity. In LLM applications it is the box labelled "toxicity check" in a guardrail diagram, sitting on the user's prompt, the model's response, or both. One function call returns a number between 0 and 1. Most of the engineering is in what that number means, which is less than people assume.
This article opens the box: what the common open models were trained on and therefore measure, how long text is cut up before scoring, why 0.5 is not a probability on your traffic, and how to test for the best-known failure of this model family, which is flagging harmless sentences that mention an identity group. It includes code for a scanner service, a threshold sweep and a bias audit, plus a worked example. What to do with a hit on the response side is covered in output filtering in depth. Here the focus is the scanner itself.
What a toxicity score measures
Start with the training data, because a classifier measures whatever its labels measured. The widely used open models come from public Jigsaw/Conversation AI datasets in which human raters marked online comments (Wikipedia talk pages and Civil Comments) as toxic or not, plus sub-types such as insult, threat, obscene and identity attack. Three consequences follow:
- Toxicity is not harm. A polite answer explaining something dangerous scores near zero, because nothing in it is rude. A toxicity scanner cannot replace a policy classifier that judges content against categories like weapons or self-harm.
- The labels are multi-label. Each category has its own sigmoid output, so a text can score 0.9 for insult and 0.2 for threat. Scores do not sum to one, and a single number hides which kind of abuse fired.
- The domain is online comments. Model responses, code and support chat look different, and scores shift on unfamiliar text. That is why thresholds must be set on your own traffic.
There is also a learned shortcut. The Detoxify README warns that text containing words associated with swearing or insults is likely to be classified as toxic "regardless of the tone or the intent of the author", and that this can bias results against minority groups. Identity terms appeared more often in attacking comments than neutral ones, and models picked up the correlation.
From text to decision
The diagram shows the path one piece of text takes. Two steps look trivial and still cause most of the bugs: segmentation and aggregation.
Segmentation. BERT- and RoBERTa-family classifiers accept at most 512 tokens, and most wrappers truncate silently, so an abusive sentence at the end of a long response is never seen. Even within the limit, one abusive sentence among many neutral ones is diluted when the text is scored as one unit. Scoring per sentence fixes both. LLM Guard's Toxicity scanner exposes this choice as MatchType.SENTENCE versus MatchType.FULL. The cost is lost context: "I'll kill it" about a presentation reads like a threat.
Aggregation. Take the maximum score over segments per label, which preserves the worst sentence. A mean reintroduces the dilution segmentation removed. Keep per-segment scores in the log so a reviewer can see which sentence fired.
Normalisation. Apply Unicode NFKC and collapse whitespace. That undoes cheap obfuscation such as full-width letters, but not deliberate misspellings.
Choosing a scanner model
For self-hosted scanning, two open projects cover most teams. The details below come from their own documentation, checked at the time of writing.
| Option | Base model / training data | Labels | Notes |
|---|---|---|---|
Detoxify original | bert-base-uncased, Toxic Comment Classification Challenge (Wikipedia) | toxic, severe toxic, obscene, threat, insult, identity hate | Oldest; no bias-aware objective |
Detoxify unbiased | roberta-base, Unintended Bias in Toxicity Classification (Civil Comments) | toxicity, severe_toxicity, obscene, threat, insult, identity_attack, sexual_explicit | Trained on the dataset built to measure identity bias |
Detoxify multilingual | xlm-roberta-base, Wikipedia + Civil Comments | unbiased labels plus identity-mention labels | English, French, Spanish, Italian, Portuguese, Turkish, Russian only |
Detoxify original-small / unbiased-small | ALBERT versions of the above | as the parent | Cheaper on CPU; measure the accuracy loss |
LLM Guard Toxicity | unitary/unbiased-toxic-roberta | verdict plus risk score | Hugging Face release of the unbiased model; default threshold 0.5 |
Hosted moderation endpoints remove the GPU bill, but you lose version pinning and the text leaves your boundary. The same calibration and bias tests apply. Treat a vendor's category names as their policy, not yours, and map them explicitly.
Outside its seven languages the multilingual model still returns confident-looking numbers. Run language identification first and route unsupported languages elsewhere: a multilingual policy model, review, or a stricter default.
A scanner service in code
A minimal scanner service on Detoxify: normalise, segment, score in one batch, aggregate by maximum, and return enough metadata to audit the decision. Label keys differ between variants, so the code reads whatever keys the model returns.
import re, unicodedata
from detoxify import Detoxify
MODEL_NAME = "unbiased" # pin this, and log it with every score
_model = Detoxify(MODEL_NAME, device="cuda")
_SENT = re.compile(r"(?<=[.!?])\s+")
def normalise(text: str) -> str:
text = unicodedata.normalize("NFKC", text)
return re.sub(r"\s+", " ", text).strip()
def segment(text: str, max_chars: int = 1500) -> list[str]:
out = []
for s in _SENT.split(text):
while len(s) > max_chars: # hard-wrap so nothing is truncated away
out.append(s[:max_chars]); s = s[max_chars:]
if s:
out.append(s)
return out or [""]
def scan(text: str) -> dict:
segs = segment(normalise(text))
scores = _model.predict(segs) # label -> list of floats, one per segment
result = {"model": MODEL_NAME, "n_segments": len(segs), "labels": {}}
for label, vals in scores.items():
i = max(range(len(vals)), key=vals.__getitem__)
result["labels"][label] = {"max": float(vals[i]), "segment": i}
return resultLLM Guard packages the same idea. Its documented interface returns the possibly sanitised output, a validity flag and a risk score:
from llm_guard.output_scanners import Toxicity
from llm_guard.output_scanners.toxicity import MatchType
scanner = Toxicity(threshold=0.5, match_type=MatchType.SENTENCE)
sanitized_output, is_valid, risk_score = scanner.scan(prompt, model_output)When the top class is toxic, the risk score is the model's confidence. When it is non-toxic, the score is one minus that confidence. The default 0.5 is a starting point, not a calibrated value. The LLM Guard guide shows how this scanner composes with the library's other scanners.
Calibrating thresholds on your traffic
A sigmoid output of 0.7 means "0.7 on the training distribution, if the model was calibrated there". On your traffic it means nothing until measured. Label a sample of your own text against your own policy, then pick each label's threshold from the precision-recall trade-off you can afford.
Sample from production rather than hand-written cases, and oversample items scoring between 0.2 and 0.9, where thresholds will sit. A few hundred labelled items per surface is a workable start. Use two raters on a subset and record agreement. If raters agree weakly, no threshold will be stable, and the problem is the policy, not the model.
import numpy as np
def sweep(scores: np.ndarray, labels: np.ndarray, min_precision: float):
"""Threshold with the best recall whose precision meets the floor."""
best = None
for t in np.unique(scores):
pred = scores >= t
tp = int((pred & labels).sum()); fp = int((pred & ~labels).sum())
fn = int((~pred & labels).sum())
if tp == 0:
continue
prec, rec = tp / (tp + fp), tp / (tp + fn)
if prec >= min_precision and (best is None or rec > best[2]):
best = (float(t), prec, rec)
return best # one call per label, per surfaceChoose the precision floor from the action. A hard block on a user's message needs high precision, because every false positive refuses a possibly legitimate person. Routing to review can run at lower precision and higher recall. Keep separate thresholds for prompts and responses: user text is noisier, and model responses rarely contain profanity unless asked to quote it.
Auditing identity-term bias
The best-studied failure of toxicity classifiers is unintended identity bias. Neutral sentences that mention an identity group, such as "I am a gay man", score higher than the same sentences without the identity term. A scanner with this bias silences the users it is meant to protect.
Borkan et al. (2019), "Nuanced Metrics for Measuring Unintended Bias with Real Data for Text Classification", the work behind the Jigsaw Unintended Bias competition, defined three AUCs per identity subgroup:
- Subgroup AUC: AUC only on texts mentioning the identity. Low means poor separation within that group.
- BPSN AUC (background positive, subgroup negative): non-toxic mentions against toxic texts without the identity. Low means neutral mentions score like real abuse, the false-positive problem.
- BNSP AUC (background negative, subgroup positive): toxic texts mentioning the identity against non-toxic texts without it. Low means abuse aimed at the group scores too low.
The competition folded these and overall AUC into one power-mean score that punishes the worst subgroup. For a release gate it is clearer to report all three per subgroup and fail on the minimum.
import numpy as np
from sklearn.metrics import roc_auc_score
def bias_report(df, identities, score="score", label="toxic"):
# df: one row per text; label and identity columns are booleans
rows = []
for g in identities:
sub, bg = df[df[g]], df[~df[g]]
rows.append({
"identity": g,
"subgroup_auc": roc_auc_score(sub[label], sub[score]),
"bpsn_auc": roc_auc_score(
np.r_[np.zeros((~sub[label]).sum()), np.ones(bg[label].sum())],
np.r_[sub.loc[~sub[label], score], bg.loc[bg[label], score]]),
"bnsp_auc": roc_auc_score(
np.r_[np.ones(sub[label].sum()), np.zeros((~bg[label]).sum())],
np.r_[sub.loc[sub[label], score], bg.loc[~bg[label], score]]),
})
return rows # gate on min(bpsn_auc) and min(bnsp_auc)If production has too few identity mentions, add a template suite: sentences like "I am a {identity} person" filled with each identity term, all labelled non-toxic. Compare raw mean scores too. If one identity's mean is several times the no-identity baseline, that is the bug, whatever the AUC says.
Worked example: a community question-and-answer assistant
Take a community Q&A product with an LLM assistant. It scans user posts before they reach the model and model answers before display. It starts with Detoxify original scoring the full text at a 0.5 threshold. The numbers here are illustrative.
Shadow mode. Scores are logged and nothing is blocked. Moderators label 600 sampled posts, oversampling the 0.2 to 0.9 band. Three patterns appear. Long posts with one abusive line at the end score low, diluted by full-text scoring. Users describing harassment they received score high, because they quote the slurs. Neutral posts about religion or sexuality cluster around 0.3 to 0.4, above comparable posts without identity terms.
Changes. Sentence matching with maximum aggregation fixes the dilution. Swapping to Detoxify unbiased and rerunning the identity suite raises the worst subgroup's BPSN AUC, and the template gaps shrink without vanishing. Insult and threat get high-precision blocking thresholds. Identity attack gets a lower threshold that routes to human review, because a blocker can't tell a victim quoting abuse from an abuser.
Outcome. Moderators agree with the blocks on re-review, and the review queue is small enough to staff. Every block carries its segment, label, score and model version. On responses the scanner rarely fires, as expected for an aligned model, so it stays as a cheap tripwire for jailbroken outputs.
Failure modes
- Obfuscation. Misspellings, inserted symbols and homoglyphs lower scores. NFKC catches some; nothing catches creative spelling reliably. Watch for clusters of near-threshold scores from one user.
- Quotation and counter-speech. Reporting or condemning abuse uses the abuser's words. A scanner scores words, not stance.
- Reclaimed in-group language. This is identity bias in another form, measurable only with data from that community.
- Unsupported languages and code-switching. Scores look normal and mean nothing. Gate on language ID.
- Silent truncation. Text past 512 tokens is never seen unless you segment.
- Non-prose content. Code, logs and base64 produce erratic scores. Decide explicitly how to handle fenced code.
- Drift. Slang and user bases move. Thresholds set in spring are wrong by autumn if nobody re-samples.
Running it in production
Latency. A base-size encoder on a GPU scores a batch of short segments in milliseconds. On CPU, use the ALBERT variants or an ONNX export. Batch all segments of a text into one predict call. For streamed responses, scan at sentence boundaries; the streaming moderation guide covers holding back tokens so a flagged sentence never reaches the screen.
Versioning. Pin model and package versions, log them with every score, and treat an upgrade as a threshold reset.
Privacy. Log scores, offsets and hashes by default; keep raw text only for sampled review, with a retention limit.
Layering. Toxicity is one detector among several. The prompt injection scanners guide covers the detector most often paired with it on the input side.
What to do next
- Write down what your policy calls toxic for user input and for model output, and which labels drive which action.
- Deploy in shadow mode with sentence segmentation, max aggregation and the model version logged.
- Label a few hundred production samples per surface, oversampling the 0.2 to 0.9 band.
- Run the threshold sweep per label and surface, with a precision floor chosen from each action's cost.
- Run the identity bias report and a template suite; fail the release on the minimum BPSN and BNSP AUC.
- Route identity-attack and borderline hits to review, and add language ID before scoring.
- Re-sample and re-sweep on a schedule and after every model upgrade.