Most LLM applications wrap the model in a moderation layer: a classifier that scores user input and model output against policy categories and blocks, flags or routes anything over a threshold. A moderation bypass is any way of getting policy-violating content through that layer, either into the model or out to the user. It is related to jailbreaking but not the same thing. A jailbreak defeats the model's own refusal training; a moderation bypass defeats the separate filter. Attackers usually need both, and a defence that relies on either alone fails when the other does.

This article is written for the people who build and run those filters. It sets out a threat model, sorts bypass techniques into six classes by the gap each one exploits, and pairs every class with a concrete defence. It then builds a moderation gateway and a safe streaming path in code, and covers how to measure robustness and what to watch in production. It deliberately contains no working bypass strings. How to design the moderation system itself (policy taxonomy, calibration, review queues) is covered in LLM moderation architecture.

Advertisement

Threat model: who bypasses moderation, and why it works

Three kinds of actor matter. Users seeking disallowed output want the model to produce something the policy forbids and phrase their input to slip past the input check. Abusers using your product as a channel want to post harassment, spam or illegal content through a feature that publishes model output or user text. Indirect attackers plant content in documents, web pages or tool results so it reaches the model without passing through the user-input filter at all. Each has different goals, but they exploit the same structural fact.

That fact is a mismatch between what the classifier evaluates and what the model understands. Moderation classifiers are usually much smaller than the model they guard, with their own tokenizer, input length and language coverage. The main model can decode encodings, follow instructions spread over many turns, read dozens of languages and infer intent from context. Every capability the model has and the classifier lacks is a place where the two see different things, and every bypass class below is an instance of that gap.

Six classes of bypass

Where moderation is bypassed: every arrow the classifier does not see, or sees in a different form from the modelUser inputtext, files, imagesInput classifierfixed window, own tokenizerLLMdecodes, translatesOutput checkper chunk?Tools / RAGretrieved textUserstreamed tokensunchecked?released early?1. Representation gapencoded or obfuscated form2. Window gapcontent past truncation3. Decomposition gapsplit over turns or fields4. Reframing gapfiction, translation, rare language5. Channel gaptool output, files, streams6. Model as decoderclassifier sees noise, model sees intentDefence principle: moderate the canonical form, the whole conversation, every channel, and the output the user actually receives.
The six gaps between what the classifier sees and what the model understands or the user receives. Each gap has its own defence; none is closed by a stronger classifier alone.
ClassWhat the attacker exploitsPrimary defence
RepresentationThe same meaning in a different surface form: character substitutions, look-alike letters, invisible characters, spacing and punctuation noiseCanonicalise before scoring; train on perturbed data
WindowThe classifier reads only the first N tokens, or a truncated summary, while the model reads everythingScore every window of the full text; take the worst
DecompositionA request split across turns, form fields, attachments or tool calls, each piece benign aloneScore the recent conversation and all fields together
ReframingFiction, role-play, hypotheticals, translation, low-resource languagesMultilingual classifiers; moderate the output, where intent becomes content
ChannelContent that enters or leaves by a path with no check: retrieved documents, tool results, images, file names, streamed chunksPut a check on every ingress and egress path
Model as decoderInput the classifier cannot read (an encoding, a cipher described in the prompt) that the model decodes and acts onOutput moderation; optionally classify the model's own restatement of the request

The table has a pattern. Input-side defences narrow the first four gaps but can never close them, because the space of ways to say something is effectively unbounded and the model can always understand more than the classifier. The output side is where intent turns into actual content, in a form the model has already decoded. Checking only input is bypassed by construction; output checks are the load-bearing control.

Advertisement

Closing the representation and window gaps

Representation attacks work because classifiers learn surface statistics. Two changes help most. First, canonicalise text before scoring. Unicode NFKC normalisation folds full-width letters, ligatures and other compatibility forms into their plain equivalents. It does not map look-alike letters from other scripts (a Cyrillic letter that looks Latin stays Cyrillic), so add a confusables mapping in the style of Unicode Technical Standard 39's skeleton. Strip invisible format characters, with care: the zero-width joiner and non-joiner are needed in emoji sequences and in scripts such as Persian. Keep the original text for the model and the logs; the canonical form is only what you classify. The Unicode smuggling defence article covers the character-level details. Second, train or fine-tune on perturbed examples, so the classifier has seen the noise it will face.

Window attacks exploit fixed input lengths. BERT-style encoders have a maximum of 512 positions, and many deployments simply truncate longer input, so anything after the cut is never scored. The fix is to split the full text into overlapping windows in the classifier's own tokens, score each and take the maximum per category. Worked example: a 3,000-token message with 512-token windows and a 64-token overlap needs windows starting every 448 tokens, which is 7 windows, so the scoring costs 7 classifier passes instead of 1. Batch them in one call to keep latency close to a single pass. Taking the maximum raises false positives slightly compared with averaging, which is the right trade: averaging lets one violating paragraph hide among six benign ones.

A moderation gateway in code

The gateway below applies those defences and adds the conversation-level check that addresses decomposition. It scores the new message alone and the recent conversation joined together, and blocks if either is over threshold. The classifier interface is abstract: tokenize must use the classifier's tokenizer, not the main model's, because window arithmetic in the wrong tokenizer re-creates the truncation gap.

import unicodedata

WINDOW, OVERLAP = 512, 64            # classifier limit in ITS tokens, and window overlap

def canonical(text: str, confusables: dict) -> str:
    # One canonical form for classification; the original is kept for the model and the logs.
    t = unicodedata.normalize("NFKC", text)                    # full-width, ligatures, compat forms
    t = "".join(ch for ch in t
                if unicodedata.category(ch) != "Cf" or ch in ("‌", "‍"))  # keep ZWNJ/ZWJ
    return "".join(confusables.get(ch, ch) for ch in t)        # TR39-style skeleton mapping

def windows(ids):
    step = WINDOW - OVERLAP
    for start in range(0, max(len(ids) - OVERLAP, 1), step):
        yield ids[start:start + WINDOW]                        # every token lands in some window

def score_text(text, clf, confusables):
    ids = clf.tokenize(canonical(text, confusables))
    scores = [clf.score_ids(w) for w in windows(ids)]          # dict: category -> probability
    return {cat: max(s[cat] for s in scores) for cat in scores[0]}   # worst window wins

def moderate_turn(conversation, new_text, clf, confusables, thresholds):
    # Score the new message alone AND the recent conversation, so split payloads are seen whole.
    recent = "\n".join(m.text for m in conversation[-6:]) + "\n" + new_text
    single, joint = score_text(new_text, clf, confusables), score_text(recent, clf, confusables)
    return {cat: max(single[cat], joint[cat]) >= thresholds[cat] for cat in thresholds}

Several details matter. The worst window decides, never the average. The conversation window is bounded (six messages here) so cost stays predictable; choose the bound from how far apart split payloads appear in your red-team data. Multi-field inputs such as forms, tool arguments and attachment text should be joined and scored together as well as separately. And thresholds are per category, because a single global threshold is miscalibrated for every category at once.

Streaming without leaking

Streaming creates its own channel gap. If the output check runs on each chunk as it arrives, a violating sentence spread over ten chunks may never appear whole in any of them. If the check runs only at the end, the user has already seen everything. The answer is a hold-back buffer: keep the most recent few hundred characters unreleased, moderate the accumulated output before releasing each older portion, and run a final check on the complete reply before releasing the tail.

async def moderated_stream(token_stream, clf, confusables, thresholds, hold_chars=400):
    # Release text only after the accumulated output (not the chunk) has passed moderation.
    produced, released = "", 0
    async for piece in token_stream:
        produced += piece
        if len(produced) - released < hold_chars:
            continue                                   # keep a hold-back buffer
        if violates(score_text(produced, clf, confusables), thresholds):
            yield {"type": "stopped", "reason": "policy"}
            return                                     # never release the buffered tail
        cut = len(produced) - hold_chars // 2          # release the older half of the buffer
        yield {"type": "text", "text": produced[released:cut]}
        released = cut
    if violates(score_text(produced, clf, confusables), thresholds):
        yield {"type": "stopped", "reason": "policy"}
        return
    yield {"type": "text", "text": produced[released:]}   # final check covers the whole reply

def violates(scores, thresholds):
    return any(scores[cat] >= thresholds[cat] for cat in thresholds)

The buffer size is a trade-off between latency and exposure. A 400-character hold-back delays text by a sentence or two, which most users do not notice, and bounds how much violating text can be released before the check sees the context around it. Scoring the whole accumulated text each time is quadratic over a long reply; in production, score only a sliding window that covers the unreleased buffer plus enough released context, and batch the calls. When the stream is stopped, replace the partial message on the client instead of leaving a truncated violating sentence on screen.

The model-as-decoder problem

The hardest class is input the classifier cannot read at all but the model can: text in an encoding, a substitution scheme explained in the same prompt, or instructions hidden in an image. Input canonicalisation cannot undo an arbitrary encoding, and a classifier that tried to decode everything would turn into a second copy of the model. Three defences combine. First, output moderation, because whatever the model decodes and acts on appears in its output in plain form (and if the output is itself encoded, flag high-entropy or encoded output for categories where that is unusual). Second, restatement classification: for high-risk surfaces, ask a model to summarise what the user is requesting, in plain language, and classify that summary. This costs a model call, so reserve it for cases where input and output scores disagree or where the input is unusually encoded. Third, rely on the model's own refusal training as one layer among several; the jailbreak defence article covers hardening that layer.

Measuring robustness

A classifier's accuracy on a clean test set says little about how easily it is bypassed. Measure bypass directly. Keep an internal, access-controlled seed set of known-violating and known-benign examples per category. Apply transformation families to each: character noise, look-alike substitution, invisible characters, padding to push content past the window, splitting across turns, translation into each supported language, and fictional framing. Then report, per category and per transformation, the bypass rate (violating items that pass) and the false positive rate on transformed benign items. A defence that halves the bypass rate on look-alike letters while tripling false positives on legitimate non-Latin text is not a win.

Track these numbers in CI for every classifier, threshold or canonicalisation change, and feed fresh techniques from red-team exercises and production incidents back into the transformation set. For background on how attackers iterate against models, see LLM jailbreaking.

Operating the defence

  • Detect probing. Bypass attempts are iterative. An account sending many near-duplicate requests that score just under a threshold, or whose input scores are low but output scores high, is probing; rate-limit it and route it to review.
  • Use account-level signals. Per-request decisions miss slow campaigns. Aggregate flags per user and per tenant over days, and escalate on accumulated risk.
  • Decide fail-closed or fail-open per surface. If the classifier times out, a public posting feature should hold the content; an internal coding assistant may reasonably continue and log.
  • Log both forms. Store the original and canonical text, the per-window scores and the decision, so reviewers can see why something passed.
  • Moderate every channel. Inventory every path into the model (user text, files, retrieved documents, tool results, images) and every path out (chat, posted content, emails, tool arguments), and confirm each has a check.

Trade-offs and failure modes

Every defence here costs something. Canonicalisation can damage legitimate text: mathematical symbols, code and some scripts lose meaning under aggressive folding, which is why the canonical form is used only for scoring. Windowing and conversation-level scoring multiply classifier calls, so batch them and cache scores for messages already seen. Taking the worst window and scoring conversations raises false positives, and those fall hardest on users writing in languages the classifier handles poorly, so measure false positive rates per language before tightening thresholds. Hold-back buffers add latency. Restatement classification adds a model call. The right mix depends on the surface: public, broadcast features justify all of them, while a private, authenticated internal tool may need only output checks and logging.

The common failure is treating moderation as a single classifier call on user input. The layered version (canonical input checks, windowed and conversation-aware scoring, per-channel checks, streamed output checks, and account-level monitoring) is more work, but each layer closes a gap the others leave open. Output handling downstream of moderation, such as rendering and tool execution, is covered in output handling.

What to do next

  1. Draw every input and output channel of your application and mark which ones currently pass through moderation; add a check to every unmarked path.
  2. Check whether your classifier truncates long input, and if it does, replace truncation with overlapping windows scored at the maximum.
  3. Add NFKC normalisation, a confusables mapping and careful format-character stripping before scoring, keeping the original text for the model.
  4. Score the recent conversation together with each new message, and join multi-field inputs before scoring.
  5. Put a hold-back buffer and a final whole-reply check on every streamed response.
  6. Build a transformation-based bypass eval per category, track bypass and false positive rates per language in CI, and alert on accounts that probe near thresholds.
Key takeaway: Moderation bypass exploits the gap between what a small classifier sees and what a capable model understands: different surface forms, truncated windows, requests split across turns, reframing, unchecked channels and encodings only the model can decode. Input checks narrow those gaps but cannot close them, so the load-bearing control is moderating the output the user actually receives. Canonicalise, score every window and the conversation, check every channel, hold back streamed text, and measure bypass rates per transformation and language.