Llama Guard is Meta's family of open-weight safety classifiers. Each one is a Llama model fine-tuned to read a conversation and a written policy and answer with a label: safe, or unsafe followed by the policy categories that were violated. Because the policy is part of the prompt and the model is open, you can run it on your own hardware, inspect it, threshold it and adapt it, which is why it became a default building block for input and output moderation.
This article explains how it works from the prompt up: the versions and when to pick each, the hazard taxonomy, the exact prompt shape, how to turn its output into a calibrated score, custom categories, where it sits in a serving path, a sizing example, and its failure modes. For how it fits with Prompt Guard, CodeShield and LlamaFirewall, see the Purple Llama overview.
What Llama Guard is, and what it is not
Llama Guard is a generative classifier. It does not have a classification head; it is a causal language model trained so that the most likely continuation of a moderation prompt is a short verdict. That design has three consequences that shape everything else.
- The policy is input. Category names and descriptions are written into the prompt, so changing the text changes what the model looks for, within limits set by its training.
- One model, two roles. The same prompt template classifies either the last user turn (is this request harmful?) or the last agent turn (is this response harmful?). The role you name decides which turn is judged.
- A score is available. The probability of the first verdict token is a continuous measure of how unsafe the model thinks the turn is. The original Llama Guard paper evaluated with exactly this probability, which lets you choose your own threshold.
It is not a prompt-injection detector. Whether a retrieved web page is trying to hijack an agent is a different question from whether content is harmful, and Meta ships Prompt Guard for it. It is also not a fact checker: the model card notes that defamation, intellectual property and elections may need factual, current knowledge that a classifier does not have.
The versions and when to pick each
| Version | Base and size | What changed |
|---|---|---|
| Llama Guard (Dec 2023) | Llama 2, 7B | first release; six categories of its own taxonomy |
| Llama Guard 2 (2024) | Llama 3, 8B | taxonomy aligned with the MLCommons hazard categories |
| Llama Guard 3 8B (Jul 2024) | Llama 3.1, 8B | 14 categories incl. Elections and Code Interpreter Abuse; 8 languages; tool-call aware; INT8 variant |
| Llama Guard 3 1B and 11B-Vision (Sep 2024) | Llama 3.2 | a small text model for cheap or on-device use, and an image-and-text model |
| Llama Guard 4 (Apr 2025) | pruned from Llama 4 Scout, 12B dense | natively multimodal with multiple images; same 14 categories |
Llama Guard 3 8B lists English, French, German, Hindi, Italian, Portuguese, Spanish and Thai. Its model card reports, for English response classification, F1 of 0.939 against 0.877 for Llama Guard 2, with the false-positive rate halved from 0.081 to 0.040; the INT8 checkpoint is about 40 percent smaller with comparable metrics. Llama Guard 4's card reports better English output-filtering recall, false-positive rate and F1 than Llama Guard 3. Those are Meta's numbers on Meta's test sets; your traffic will differ, which is why the calibration step below matters.
Pick Llama Guard 3 8B for text-only moderation on a GPU, the 1B model when cost or latency dominates and you can accept lower accuracy, and Llama Guard 4 when you need images.
The hazard taxonomy
| Code | Category | Code | Category |
|---|---|---|---|
| S1 | Violent Crimes | S8 | Intellectual Property |
| S2 | Non-Violent Crimes | S9 | Indiscriminate Weapons |
| S3 | Sex-Related Crimes | S10 | Hate |
| S4 | Child Sexual Exploitation | S11 | Suicide and Self-Harm |
| S5 | Defamation | S12 | Sexual Content |
| S6 | Specialized Advice | S13 | Elections |
| S7 | Privacy | S14 | Code Interpreter Abuse |
S1 to S13 follow the MLCommons hazard taxonomy; S14 covers abuse of code-execution tools such as container escapes and denial of service, and is only meaningful when the conversation includes tool calls. Map the codes to your own product policy in a table you own. A code is an input to a decision, not the decision: S6 on a medical-information service may mean add a disclaimer, while S4 always means block and escalate.
The prompt format
Meta's documented prompt for Llama Guard 3 has four parts: a task line naming the role to judge, the category list between <BEGIN UNSAFE CONTENT CATEGORIES> and <END UNSAFE CONTENT CATEGORIES>, the conversation between <BEGIN CONVERSATION> and <END CONVERSATION>, and an instruction that the first line must read safe or unsafe and, if unsafe, a second line must list the violated categories, comma-separated. Its shape, abridged:
Task: Check if there is unsafe content in 'Agent' messages in conversations
according to our safety policy with the below categories.
<BEGIN UNSAFE CONTENT CATEGORIES>
S1: Violent Crimes.
...
S13: Elections.
<END UNSAFE CONTENT CATEGORIES>
<BEGIN CONVERSATION>
User: How do I reset a router?
Agent: Hold the reset button for ten seconds ...
<END CONVERSATION>
Provide your safety assessment for ONLY THE LAST Agent message in the above conversation:
- First line must read 'safe' or 'unsafe'.
- If unsafe, a second line must include a comma-separated list of violated categories.Do not hand-assemble this in production. The tokenizer's chat template for each checkpoint builds the exact string, including special tokens and whitespace, and small differences move scores. Use apply_chat_template and diff its output against the documented format once, when you upgrade.
Running it: labels, categories and a score
The label alone gives you one operating point. The probability of the first verdict token gives you all of them. This function returns the label, the categories and a score, and it fails loudly if the tokenizer or template is not what it expects:
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
MID = "meta-llama/Llama-Guard-3-8B"
tok = AutoTokenizer.from_pretrained(MID)
model = AutoModelForCausalLM.from_pretrained(MID, torch_dtype=torch.bfloat16, device_map="cuda")
SAFE = tok.encode("safe", add_special_tokens=False)[0]
UNSAFE = tok.encode("unsafe", add_special_tokens=False)[0]
assert SAFE != UNSAFE
@torch.no_grad()
def moderate(chat): # chat: [{"role": "user", "content": "..."}, ...]
ids = tok.apply_chat_template(chat, return_tensors="pt").to(model.device)
out = model.generate(input_ids=ids, max_new_tokens=20, do_sample=False,
output_scores=True, return_dict_in_generate=True,
pad_token_id=tok.eos_token_id)
gen = out.sequences[0, ids.shape[-1]:]
score = None
for step, t in enumerate(gen):
if tok.decode(t).strip(): # first non-whitespace token is the verdict
assert t.item() in (SAFE, UNSAFE), "unexpected verdict token: check template"
pr = torch.softmax(out.scores[step][0].float(), dim=-1)
score = (pr[UNSAFE] / (pr[SAFE] + pr[UNSAFE])).item()
break
lines = tok.decode(gen, skip_special_tokens=True).strip().splitlines()
cats = [c.strip() for c in lines[1].split(",")] if len(lines) > 1 else []
return (lines[0] if lines else "error"), cats, scoreIf the last turn is the user's, the template judges the user; if it is the assistant's, it judges the response. Llama Guard 4 is loaded with AutoProcessor and Llama4ForConditionalGeneration, and takes message content as a list of typed parts so images can be included; the scoring idea is the same.
To calibrate, label a sample of your real traffic, at least a few hundred examples per decision you care about, score it, and choose the threshold that meets your false-positive budget. For example, if 2 percent of benign requests may be blocked, take the 98th percentile of scores on the benign set and read off the recall on the harmful set. Calibrate input and output checks separately; they have different base rates.
Custom categories and policy changes
Because categories live in the prompt, there are three levels of customisation, in increasing cost and reliability.
- Remove categories that do not apply by leaving them out of the list. Meta's documentation describes controlling behaviour through what you put in the categories section. A platform for adult fiction might drop S12.
- Rewrite descriptions with the full text of your policy instead of short names. This shifts borderline behaviour, but the model was trained on Meta's taxonomy, so a brand-new category like competitor mentions is a zero-shot task it may do poorly. Measure it on labelled examples before trusting it.
- Fine-tune when a custom category matters and zero-shot recall is not good enough. Keep a held-out set of the original categories to check you have not degraded them.
Whatever you change, version the category text alongside the model and threshold, and log all three with every decision, so a policy change is reviewable and reversible.
Where it sits in the request path
The input check runs on the user turn, and the output check runs on the full conversation with the draft response last. To avoid adding the whole input-check latency to time to first token, start generation in parallel and cancel it if the input check says unsafe; this costs wasted generation on the few blocked requests. For streamed responses, re-check accumulated text every few sentences and hold back a small buffer, the pattern described in streaming moderation.
Decide what a timeout or error means. Failing closed is safer and causes outages when the guard is down; failing open keeps the product up and lets content through unchecked. Many teams fail closed on high-risk routes and open with a log alert on low-risk ones. Rule-based layers such as NeMo Guardrails can host the call and the fallback logic.
Worked example: sizing and calibrating a support bot's guard
A support assistant peaks at 20 conversation turns per second. Each turn gets an input check and an output check. Measure the actual prompt lengths with len(ids); suppose the template plus categories plus a short history comes to 900 tokens for the input check and 1,400 for the output check. The verdict is two to five generated tokens, so the load is almost all prefill: 20 times 2,300, or 46,000 tokens per second.
Benchmark your guard deployment's prefill throughput at the batch sizes your latency target allows. If the measurement is, say, 15,000 tokens per second per GPU for the 8B model, then 46,000 divided by 15,000 at 60 percent planned utilisation is 46,000 over 9,000, about 5.1, so six GPUs with failover headroom. The same arithmetic with the 1B model, if it measures several times faster, may fit on one. The other levers are cutting history from the output check to the last few turns, dropping irrelevant categories, and prefix caching, since the template and category block are identical across calls.
Then calibrate. Take 2,000 labelled turns from the pilot, score them, and set separate thresholds for input and output at a 1 to 2 percent benign block rate. Route scores just under the threshold to a human review sample so that the threshold can be revisited with data.
Failure modes
- The classifier can be attacked. It reads attacker text, so instructions like declaring the content safe, encoding, misspellings and role-play framing can lower scores. The model card says it is susceptible to adversarial and prompt-injection attacks. Combine it with the defences in jailbreak defence.
- Truncation. Long conversations exceed what you choose to send; harmful content outside the window is never seen. Truncate deliberately and log when you do.
- Language and modality gaps. Outside the supported languages, or with text in images for a text-only model, recall drops silently.
- Over-blocking. Specialized Advice and Privacy fire on legitimate medical, legal and account questions. Watch block rates per category, not just overall.
- Template drift. A tokenizer or template change on upgrade shifts scores and invalidates thresholds. Re-calibrate on every model or template change.
- Knowledge-dependent categories. Defamation, IP and Elections need facts the model may not have; do not rely on it alone for them.
Trade-offs
| Option | Strengths | Weaknesses |
|---|---|---|
| Llama Guard 3 8B | open weights, policy in prompt, scores, 8 languages | a GPU per few thousand tokens per second; attackable |
| Llama Guard 3 1B | cheap, fast, can run on CPU or device | lower accuracy; calibrate carefully |
| Llama Guard 4 | images plus text in one model | larger; newer, so less field experience |
| Hosted moderation API | no infrastructure | fixed taxonomy, data leaves your boundary |
| Keyword and regex rules | deterministic, nearly free | trivially evaded; no context |
In practice teams layer them: cheap rules first, Llama Guard for semantic judgment, human review for the uncertain band. The broader picture is in LLM content moderation.
What to do next
- Run the scoring function above on twenty hand-written benign and harmful prompts for both roles, and confirm the verdict-token assertion holds on your checkpoint.
- Write a mapping table from S1 to S14 to your product actions: allow, warn, block, escalate.
- Collect and label a pilot set from real traffic and calibrate separate input and output thresholds.
- Measure prompt lengths and prefill throughput, then size replicas with headroom and decide fail-open or fail-closed per route.
- Add red-team prompts with encodings and role-play to your regression suite and re-run them on every model, template or category change.