Llama Guard is Meta's family of open-weight safety classifiers. Each one is a Llama model fine-tuned to read a conversation and a written policy and answer with a label: safe, or unsafe followed by the policy categories that were violated. Because the policy is part of the prompt and the model is open, you can run it on your own hardware, inspect it, threshold it and adapt it, which is why it became a default building block for input and output moderation.

This article explains how it works from the prompt up: the versions and when to pick each, the hazard taxonomy, the exact prompt shape, how to turn its output into a calibrated score, custom categories, where it sits in a serving path, a sizing example, and its failure modes. For how it fits with Prompt Guard, CodeShield and LlamaFirewall, see the Purple Llama overview.

What Llama Guard is, and what it is not

Llama Guard is a generative classifier. It does not have a classification head; it is a causal language model trained so that the most likely continuation of a moderation prompt is a short verdict. That design has three consequences that shape everything else.

  • The policy is input. Category names and descriptions are written into the prompt, so changing the text changes what the model looks for, within limits set by its training.
  • One model, two roles. The same prompt template classifies either the last user turn (is this request harmful?) or the last agent turn (is this response harmful?). The role you name decides which turn is judged.
  • A score is available. The probability of the first verdict token is a continuous measure of how unsafe the model thinks the turn is. The original Llama Guard paper evaluated with exactly this probability, which lets you choose your own threshold.

It is not a prompt-injection detector. Whether a retrieved web page is trying to hijack an agent is a different question from whether content is harmful, and Meta ships Prompt Guard for it. It is also not a fact checker: the model card notes that defamation, intellectual property and elections may need factual, current knowledge that a classifier does not have.

The versions and when to pick each

VersionBase and sizeWhat changed
Llama Guard (Dec 2023)Llama 2, 7Bfirst release; six categories of its own taxonomy
Llama Guard 2 (2024)Llama 3, 8Btaxonomy aligned with the MLCommons hazard categories
Llama Guard 3 8B (Jul 2024)Llama 3.1, 8B14 categories incl. Elections and Code Interpreter Abuse; 8 languages; tool-call aware; INT8 variant
Llama Guard 3 1B and 11B-Vision (Sep 2024)Llama 3.2a small text model for cheap or on-device use, and an image-and-text model
Llama Guard 4 (Apr 2025)pruned from Llama 4 Scout, 12B densenatively multimodal with multiple images; same 14 categories

Llama Guard 3 8B lists English, French, German, Hindi, Italian, Portuguese, Spanish and Thai. Its model card reports, for English response classification, F1 of 0.939 against 0.877 for Llama Guard 2, with the false-positive rate halved from 0.081 to 0.040; the INT8 checkpoint is about 40 percent smaller with comparable metrics. Llama Guard 4's card reports better English output-filtering recall, false-positive rate and F1 than Llama Guard 3. Those are Meta's numbers on Meta's test sets; your traffic will differ, which is why the calibration step below matters.

Pick Llama Guard 3 8B for text-only moderation on a GPU, the 1B model when cost or latency dominates and you can accept lower accuracy, and Llama Guard 4 when you need images.

The hazard taxonomy

CodeCategoryCodeCategory
S1Violent CrimesS8Intellectual Property
S2Non-Violent CrimesS9Indiscriminate Weapons
S3Sex-Related CrimesS10Hate
S4Child Sexual ExploitationS11Suicide and Self-Harm
S5DefamationS12Sexual Content
S6Specialized AdviceS13Elections
S7PrivacyS14Code Interpreter Abuse

S1 to S13 follow the MLCommons hazard taxonomy; S14 covers abuse of code-execution tools such as container escapes and denial of service, and is only meaningful when the conversation includes tool calls. Map the codes to your own product policy in a table you own. A code is an input to a decision, not the decision: S6 on a medical-information service may mean add a disclaimer, while S4 always means block and escalate.

The prompt format

Meta's documented prompt for Llama Guard 3 has four parts: a task line naming the role to judge, the category list between <BEGIN UNSAFE CONTENT CATEGORIES> and <END UNSAFE CONTENT CATEGORIES>, the conversation between <BEGIN CONVERSATION> and <END CONVERSATION>, and an instruction that the first line must read safe or unsafe and, if unsafe, a second line must list the violated categories, comma-separated. Its shape, abridged:

Task: Check if there is unsafe content in 'Agent' messages in conversations
according to our safety policy with the below categories.

<BEGIN UNSAFE CONTENT CATEGORIES>
S1: Violent Crimes.
...
S13: Elections.
<END UNSAFE CONTENT CATEGORIES>

<BEGIN CONVERSATION>
User: How do I reset a router?
Agent: Hold the reset button for ten seconds ...
<END CONVERSATION>

Provide your safety assessment for ONLY THE LAST Agent message in the above conversation:
 - First line must read 'safe' or 'unsafe'.
 - If unsafe, a second line must include a comma-separated list of violated categories.

Do not hand-assemble this in production. The tokenizer's chat template for each checkpoint builds the exact string, including special tokens and whitespace, and small differences move scores. Use apply_chat_template and diff its output against the documented format once, when you upgrade.

Running it: labels, categories and a score

The label alone gives you one operating point. The probability of the first verdict token gives you all of them. This function returns the label, the categories and a score, and it fails loudly if the tokenizer or template is not what it expects:

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

MID = "meta-llama/Llama-Guard-3-8B"
tok = AutoTokenizer.from_pretrained(MID)
model = AutoModelForCausalLM.from_pretrained(MID, torch_dtype=torch.bfloat16, device_map="cuda")
SAFE = tok.encode("safe", add_special_tokens=False)[0]
UNSAFE = tok.encode("unsafe", add_special_tokens=False)[0]
assert SAFE != UNSAFE

@torch.no_grad()
def moderate(chat):                      # chat: [{"role": "user", "content": "..."}, ...]
    ids = tok.apply_chat_template(chat, return_tensors="pt").to(model.device)
    out = model.generate(input_ids=ids, max_new_tokens=20, do_sample=False,
                         output_scores=True, return_dict_in_generate=True,
                         pad_token_id=tok.eos_token_id)
    gen = out.sequences[0, ids.shape[-1]:]
    score = None
    for step, t in enumerate(gen):
        if tok.decode(t).strip():        # first non-whitespace token is the verdict
            assert t.item() in (SAFE, UNSAFE), "unexpected verdict token: check template"
            pr = torch.softmax(out.scores[step][0].float(), dim=-1)
            score = (pr[UNSAFE] / (pr[SAFE] + pr[UNSAFE])).item()
            break
    lines = tok.decode(gen, skip_special_tokens=True).strip().splitlines()
    cats = [c.strip() for c in lines[1].split(",")] if len(lines) > 1 else []
    return (lines[0] if lines else "error"), cats, score

If the last turn is the user's, the template judges the user; if it is the assistant's, it judges the response. Llama Guard 4 is loaded with AutoProcessor and Llama4ForConditionalGeneration, and takes message content as a list of typed parts so images can be included; the scoring idea is the same.

To calibrate, label a sample of your real traffic, at least a few hundred examples per decision you care about, score it, and choose the threshold that meets your false-positive budget. For example, if 2 percent of benign requests may be blocked, take the 98th percentile of scores on the benign set and read off the recall on the harmful set. Calibrate input and output checks separately; they have different base rates.

Custom categories and policy changes

Because categories live in the prompt, there are three levels of customisation, in increasing cost and reliability.

  1. Remove categories that do not apply by leaving them out of the list. Meta's documentation describes controlling behaviour through what you put in the categories section. A platform for adult fiction might drop S12.
  2. Rewrite descriptions with the full text of your policy instead of short names. This shifts borderline behaviour, but the model was trained on Meta's taxonomy, so a brand-new category like competitor mentions is a zero-shot task it may do poorly. Measure it on labelled examples before trusting it.
  3. Fine-tune when a custom category matters and zero-shot recall is not good enough. Keep a held-out set of the original categories to check you have not degraded them.

Whatever you change, version the category text alongside the model and threshold, and log all three with every decision, so a policy change is reviewable and reversible.

Where it sits in the request path

Llama Guard as two checkpoints around the application modelUserpromptInput checkrole = UserApplication LLMmay start in parallelOutput checkrole = AgentDeliversafeRefuseunsafe + S codesReplaceunsafe + S codesLog and samplescores, codes, latencysafeEach check is one classifier call that generates a few tokens; prompt length drives the cost.Decide in advance whether a guard timeout fails open or closed, per route.
Input and output checks use the same model with different roles; both feed the same audit log.

The input check runs on the user turn, and the output check runs on the full conversation with the draft response last. To avoid adding the whole input-check latency to time to first token, start generation in parallel and cancel it if the input check says unsafe; this costs wasted generation on the few blocked requests. For streamed responses, re-check accumulated text every few sentences and hold back a small buffer, the pattern described in streaming moderation.

Decide what a timeout or error means. Failing closed is safer and causes outages when the guard is down; failing open keeps the product up and lets content through unchecked. Many teams fail closed on high-risk routes and open with a log alert on low-risk ones. Rule-based layers such as NeMo Guardrails can host the call and the fallback logic.

Worked example: sizing and calibrating a support bot&#x27;s guard

A support assistant peaks at 20 conversation turns per second. Each turn gets an input check and an output check. Measure the actual prompt lengths with len(ids); suppose the template plus categories plus a short history comes to 900 tokens for the input check and 1,400 for the output check. The verdict is two to five generated tokens, so the load is almost all prefill: 20 times 2,300, or 46,000 tokens per second.

Benchmark your guard deployment's prefill throughput at the batch sizes your latency target allows. If the measurement is, say, 15,000 tokens per second per GPU for the 8B model, then 46,000 divided by 15,000 at 60 percent planned utilisation is 46,000 over 9,000, about 5.1, so six GPUs with failover headroom. The same arithmetic with the 1B model, if it measures several times faster, may fit on one. The other levers are cutting history from the output check to the last few turns, dropping irrelevant categories, and prefix caching, since the template and category block are identical across calls.

Then calibrate. Take 2,000 labelled turns from the pilot, score them, and set separate thresholds for input and output at a 1 to 2 percent benign block rate. Route scores just under the threshold to a human review sample so that the threshold can be revisited with data.

Failure modes

  • The classifier can be attacked. It reads attacker text, so instructions like declaring the content safe, encoding, misspellings and role-play framing can lower scores. The model card says it is susceptible to adversarial and prompt-injection attacks. Combine it with the defences in jailbreak defence.
  • Truncation. Long conversations exceed what you choose to send; harmful content outside the window is never seen. Truncate deliberately and log when you do.
  • Language and modality gaps. Outside the supported languages, or with text in images for a text-only model, recall drops silently.
  • Over-blocking. Specialized Advice and Privacy fire on legitimate medical, legal and account questions. Watch block rates per category, not just overall.
  • Template drift. A tokenizer or template change on upgrade shifts scores and invalidates thresholds. Re-calibrate on every model or template change.
  • Knowledge-dependent categories. Defamation, IP and Elections need facts the model may not have; do not rely on it alone for them.

Trade-offs

OptionStrengthsWeaknesses
Llama Guard 3 8Bopen weights, policy in prompt, scores, 8 languagesa GPU per few thousand tokens per second; attackable
Llama Guard 3 1Bcheap, fast, can run on CPU or devicelower accuracy; calibrate carefully
Llama Guard 4images plus text in one modellarger; newer, so less field experience
Hosted moderation APIno infrastructurefixed taxonomy, data leaves your boundary
Keyword and regex rulesdeterministic, nearly freetrivially evaded; no context

In practice teams layer them: cheap rules first, Llama Guard for semantic judgment, human review for the uncertain band. The broader picture is in LLM content moderation.

What to do next

  1. Run the scoring function above on twenty hand-written benign and harmful prompts for both roles, and confirm the verdict-token assertion holds on your checkpoint.
  2. Write a mapping table from S1 to S14 to your product actions: allow, warn, block, escalate.
  3. Collect and label a pilot set from real traffic and calibrate separate input and output thresholds.
  4. Measure prompt lengths and prefill throughput, then size replicas with headroom and decide fail-open or fail-closed per route.
  5. Add red-team prompts with encodings and role-play to your regression suite and re-run them on every model, template or category change.
Key takeaway: Llama Guard is a fine-tuned Llama that reads a policy and a conversation and answers safe or unsafe with category codes. Use the official chat template, read the verdict token's probability as a score, and calibrate separate thresholds for user and agent turns on your own traffic. Size it as a prefill workload, decide how it fails, log every decision with its policy version, and remember that it can be attacked and does not detect prompt injection.