A social app, a marketplace or a forum receives millions of pieces of user content a day, and each one needs a policy decision: leave it up, take it down, restrict it, or send it to a person. Large language models are good at this because they can read context, sarcasm and policy text, but they are also the most expensive way to make a decision that is usually easy. Most content is plainly fine. This page treats content moderation as a GPU capacity problem: how to build a cascade in which cheap models settle the easy majority, an LLM classifier settles the hard middle, and people see only what remains, and how to size the GPUs and thresholds so that the cascade is both affordable and measurably safe.
The scope is asynchronous moderation of stored user content. Guarding a chat model's own inputs and outputs in the request path is a different problem with a latency budget, covered in LLM guardrails; the policy taxonomy and review-queue design live in LLM moderation architecture.
Why a cascade
Start with cost per decision. A small embedding model with a few linear heads costs a fraction of a millisecond of GPU time per item. An 8B-parameter classifier reading a policy prompt plus the content costs tens of milliseconds. A trained reviewer costs tens of seconds of salaried time. That is roughly two orders of magnitude between each stage, so routing matters far more than any kernel optimisation: sending 10 percent of traffic to the LLM instead of 100 percent is a tenfold saving no batching trick can match.
A cascade works only if each stage can say how sure it is. A classifier that emits a bare label cannot be routed; one that emits a calibrated score can. So every stage in this design produces a number between 0 and 1 per policy category, and routing is two thresholds per category: below the low threshold the item is allowed automatically, above the high threshold it is actioned automatically, and in between it moves to the next stage. The thresholds are not guesses; they are computed from labelled data against explicit error budgets, as shown below.
Architecture
Content arrives on a queue rather than an HTTP call, which is what makes the GPU side efficient: the workers pull large batches, there is no per-request latency SLO, and the only time constraint is a freshness target such as 95 percent of items decided within five minutes of posting. Stage 1 runs on every item. Stage 2 runs on the uncertain band. Stage 3 is a review queue prioritised by severity and reach, since a post seen by a million people deserves attention before one seen by three.
The decision log is the part teams skip and regret. For each item it records the content hash, each stage's per-category score, the model and prompt versions, the policy version and the final outcome. Without it you cannot recompute thresholds, answer an appeal, or replay last month under a new policy.
Stage 1: embeddings and linear heads
Stage 1 is an encoder that turns text into a vector, followed by one logistic regression head per policy category, trained on your own labelled data. Embedding once and attaching many heads is cheap: adding a category costs a few thousand weights, not a new model. A sentence-embedding model of a few hundred million parameters processes thousands of short items per second per GPU when batched by length, so one card usually covers a large platform.
The heads are deliberately simple so that they are easy to recalibrate. Because the encoder is frozen, a policy change that adds a category needs only labels and a few minutes of training. Stage 1 also produces signals the LLM cannot: near-duplicate detection by vector distance catches the hundredth copy of a known spam post without spending any LLM time on it.
import numpy as np
from sklearn.linear_model import LogisticRegression
# X: (n, d) embeddings from the frozen encoder; Y: (n, k) 0/1 labels per category
heads = [LogisticRegression(C=1.0, max_iter=2000, class_weight="balanced").fit(X, Y[:, j])
for j in range(Y.shape[1])]
def stage1_scores(emb): # emb: (batch, d)
return np.stack([h.predict_proba(emb)[:, 1] for h in heads], axis=1)
Stage 2: an LLM classifier that writes one token
Stage 2 is an instruction-tuned safety classifier. Open options include Meta's Llama Guard 3 (1B and 8B), Llama Guard 4 (12B, which also accepts images) and Google's ShieldGemma (2B, 9B and 27B); the prompt formats are covered in Llama Guard, in depth. The trick that makes them usable in a cascade is to never let them write prose. The classifier is asked for its verdict, generation is capped at one token, and the probability of the unsafe token is read from the logprobs. That turns a generative model into a scorer, and it makes the job almost entirely prefill: one forward pass over the prompt and no decode loop.
import math
from vllm import LLM, SamplingParams
llm = LLM(model=GUARD_MODEL, enable_prefix_caching=True, max_model_len=4096)
tok = llm.get_tokenizer()
params = SamplingParams(max_tokens=1, temperature=0.0, logprobs=20)
def stage2_scores(items):
prompts = [tok.apply_chat_template([{"role": "user", "content": t}],
tokenize=False) for t in items]
outs = llm.generate(prompts, params)
scores = []
for o in outs:
top = o.outputs[0].logprobs[0] # {token_id: Logprob} for the first position
p_unsafe = p_safe = 0.0
for lp in top.values():
word = (lp.decoded_token or "").strip().lower()
if word == "unsafe":
p_unsafe += math.exp(lp.logprob)
elif word == "safe":
p_safe += math.exp(lp.logprob)
scores.append(p_unsafe / max(p_unsafe + p_safe, 1e-9))
return scoresCheck one thing against your model's tokenizer before trusting this: whether the first verdict token is the whole word or a fragment such as a leading newline, and that the rendered template plus the engine do not add the BOS token twice. Print the top logprobs for a handful of known-safe and known-unsafe items and adjust the matching. Normalising over safe plus unsafe makes the score robust to probability mass on stray tokens. Per-category scores come from the same idea applied to a per-category prompt, or from the category tokens the model emits after the verdict when you allow a few more output tokens for flagged items only.
Prefix caching matters here. When the policy text sits before the content in the prompt, every request shares that prefix, so its key-value cache is computed once and reused, and each item pays only for its own tokens. Put long, stable instructions first and the variable content last; reversing that order quietly multiplies cost.
Thresholds from labelled data
Thresholds come from a labelled validation set drawn from recent traffic, with the uncertain band oversampled so that it is measured precisely. Two budgets drive them. The miss budget says what fraction of true violations you will tolerate being auto-allowed, for example 0.5 percent for spam but 0.01 percent for child safety. The precision floor says how often an automatic takedown may be wrong, for example 99 percent precision before content is removed without a person.
import numpy as np
def pick_thresholds(scores, labels, miss_budget=0.005, precision_floor=0.99):
order = np.argsort(scores)
s, y = scores[order], labels[order]
positives = y.sum()
# low: largest cut such that violations at or below it are within the miss budget
missed = np.cumsum(y)
ok_low = np.where(missed <= miss_budget * positives)[0]
t_low = s[ok_low[-1]] if len(ok_low) else 0.0
# high: smallest cut such that everything at or above it meets the precision floor
tp_above = positives - np.concatenate([[0], missed[:-1]])
n_above = len(s) - np.arange(len(s))
ok_high = np.where(tp_above / n_above >= precision_floor)[0]
t_high = s[ok_high[0]] if len(ok_high) else 1.01
escalate = np.mean((scores > t_low) & (scores < t_high))
return t_low, t_high, escalateThe function returns the escalation rate as well, and that number is the capacity plan for the next stage. If stage 1 escalates 12 percent at a given miss budget, stage 2 must handle 12 percent of peak traffic. Tightening the miss budget widens the band and costs GPUs; that is the trade-off leadership should be shown in explicit numbers rather than discovered on an invoice. Recompute thresholds weekly from fresh labels, because content drifts and a threshold tuned in spring will be wrong by autumn.
Worked example: sizing 50 million items a day
Take a platform with 50 million new items a day, a peak-to-average ratio of 2, and items averaging 200 tokens. Stage 2 uses an 8B classifier whose prompt adds about 400 tokens of shared policy text, which prefix caching absorbs, leaving about 250 uncached tokens per item including the template suffix. A forward pass costs about 2 FLOPs per parameter per token, so each item needs 2 x 8e9 x 250 = 4e12 FLOPs. Assuming an H100-class card sustains about 400 dense BF16 TFLOP/s on large prefill batches, which is roughly 40 percent of its peak, that is 10 ms of GPU per item, or about 100 items per second per card. These are planning estimates; measure your own throughput with your real length distribution before buying anything.
| Quantity | No cascade | Cascade | How it is computed |
|---|---|---|---|
| Average rate | 579 items/s | 579 items/s | 50M / 86,400 |
| Peak rate | 1,157 items/s | 1,157 items/s | Average x 2 |
| Stage 1 GPUs | 0 | 1 | Thousands of items/s per card, one card plus a spare |
| Escalated to stage 2 | 100% | 10% | From pick_thresholds on validation data |
| Stage 2 peak load | 1,157 items/s | 116 items/s | Peak x escalation |
| Stage 2 GPUs at 70% target utilisation | 17 | 2 | load / (100 x 0.7), rounded up |
| Escalated to people (0.05%) | n/a | 25,000 items/day | Stage 2 band on validation data |
| Reviewer hours at 30 s each | n/a | 208 hours/day | About 26 eight-hour shifts |
The cascade turns 17 GPUs into 3, and it also shows where the real cost lies: reviewer time dwarfs GPU time. That changes what to optimise. A better stage 2 model that narrows its uncertain band from 0.05 percent to 0.03 percent saves ten reviewer shifts a day, which is worth far more than any throughput gain. Spend GPU to save people, not the other way round.
Policy changes and backfills
Policies change: a new category is added, a definition is tightened, a court ruling moves a line. Each change raises the question of what to do about content already live. A backfill re-scores the stored corpus under the new policy, and it is the largest GPU job moderation ever runs, so plan it as a batch job rather than letting it compete with live traffic.
Three techniques keep it affordable. First, re-run stage 1 only, with new heads, over the stored embeddings; if you kept embeddings in the decision log, this needs no encoder pass at all. Second, send to stage 2 only the items whose new stage 1 score falls in the uncertain band, and prioritise by current reach so that the most-viewed content is decided first. Third, run the backfill on a separate pool or at a lower scheduler priority so that the five-minute freshness target on new content holds. A 2-billion-item corpus at a 3 percent stage 2 rate is 60 million LLM calls, which at 100 per second per card is about 7 GPU-days: tractable on rented capacity over a weekend, ruinous if it accidentally runs against every item.
Failure modes
- Uncalibrated scores. Treating the unsafe probability from a model trained on someone else's policy as calibrated for yours. Fit thresholds on your own labels and check calibration per language.
- Language and dialect gaps. Stage 1 heads trained mostly on English auto-allow violations in other languages because their scores sit low. Set thresholds per language, or route low-resource languages straight to stage 2.
- Adversarial drift. Spammers learn the cheap stage first. Watch the share of reviewer-confirmed violations that stage 1 scored below the low threshold; evasion techniques are catalogued in moderation bypass.
- Queue starvation. A backfill or a viral event floods stage 2 and new content misses its freshness target. Use separate queues with priorities, and alert on the age of the oldest undecided item rather than on queue length.
- Silent version skew. A model or prompt update changes scores while thresholds stay put. Tie thresholds to a model and prompt hash and refuse to start a worker whose hash has no thresholds.
- Reviewer feedback bias. Labels come only from escalated items, so the training data never contains the confident mistakes. Sample a small random slice of auto-allowed and auto-actioned items into review every day.
Trade-offs
A cascade adds moving parts: two sets of thresholds per category, two models to version and a calibration job. A single LLM stage is simpler to reason about and sometimes the right call for small platforms, where a few GPUs cost less than the engineering. Larger stage 2 models such as a 27B classifier narrow the uncertain band but roughly triple the GPU cost per item; the reviewer-hour arithmetic above tells you whether that is worth it. Automatic takedowns at high precision protect users quickly but produce appeals, so budget for an appeals path. Finally, keeping embeddings and scores for every item makes backfills cheap but is itself sensitive data, so give the decision log a retention rule. For classifiers that must keep pace with live token streams instead of stored posts, see streaming moderation.
What to do next
- Measure today's volume, peak ratio and length distribution of the content you moderate.
- Label a validation set of a few thousand recent items per category, oversampling borderline cases.
- Train stage 1 heads on frozen embeddings and run the threshold picker with explicit miss and precision budgets.
- Benchmark your stage 2 classifier with one-token logprob scoring and prefix caching on real items, and record items per second per GPU.
- Size stage 2 from the escalation rate and that measured throughput, and compare it with the reviewer hours the band implies.
- Build the decision log with scores, versions and policy id before going live.
- Schedule weekly threshold recomputation and a daily random audit sample of automatic decisions.
- Write the backfill runbook: separate pool, reach-ordered, stage 1 over stored embeddings first.