Constitutional AI began as a training method: write down a short list of principles, have the model critique and revise its own outputs against them, and use AI-generated preference labels instead of human ones to train for harmlessness. It is now also a deployment method. In 2025 Anthropic published Constitutional Classifiers, input and output classifiers trained on synthetic data generated from a constitution, as a defence against universal jailbreaks. The two uses share the core idea that a natural-language specification can be turned into training data, but they protect different things.

This article is about the safety and security side. It recaps the training method in one section, then explains why training alone is not a security boundary, how constitution-driven classifiers work, what their published results do and do not show, and how to build a scaled-down version for your own application with a worked example from banking. For the training method in depth, see Constitutional AI and the reward-model mathematics in Constitutional AI math.

The training method in one section

The original paper, Bai et al., "Constitutional AI: Harmlessness from AI Feedback" (Anthropic, December 2022), has two stages. In the supervised stage, a helpful-only model answers red-team prompts, then is asked to critique its answer against a principle drawn at random from the constitution and to revise it; the model is fine-tuned on the revisions. In the reinforcement stage, the fine-tuned model produces pairs of answers, a feedback model judges which better follows a sampled principle, those AI preferences train a preference model (mixed with human labels for helpfulness), and the policy is optimised against it with reinforcement learning. The paper reported models that were less harmful at similar helpfulness and that explained objections rather than refusing evasively.

The security-relevant point is what this changes: the model's default tendencies. A CAI-trained model usually declines a plainly harmful request. It offers no guarantee against an adversary who searches for inputs that shift the model off its trained distribution, which is what a jailbreak is.

Why training alone is not a security boundary

Jailbreaks fall into families, covered in jailbreak defense architecture: role-play, obfuscation and encoding, many-shot priming, splitting a request across turns. The most dangerous kind is a universal jailbreak, a single strategy that unlocks many restricted requests, because once found it can be shared and reused by people with no skill. Training improves the average case, but the policy is a large function with an enormous input space, and an adversary only needs one region where the training signal was thin.

Guards help because they change the attacker's problem. The input classifier sees the prompt; the output classifier sees what the model actually produced. An encoding trick that confuses the model about what is being asked still has to yield harmful output in a form the user can read, and the output guard judges that output directly. Defence in depth is not a cliche here: the attacker must now find one input that fools three different learned functions at once.

Constitutional classifiers

Constitutional Classifiers work as follows. A constitution lists categories of content that are allowed and disallowed, with the allowed side written as carefully as the disallowed side so that harmless neighbouring requests are not refused. A capable model uses the constitution to generate a large set of synthetic prompts and completions in each category, and the data is augmented, for example by translation and by rewriting in the style of known jailbreaks. Classifiers trained on that data guard the model's input and output. Because the constitution is text, updating the guard for a new threat means editing the constitution and regenerating data, not collecting a fresh labelled corpus.

The output side is built for streaming. According to the paper (arXiv 2501.18837), the output classifier predicts, at each token, the harmfulness of the full output sequence, and when the score for a token exceeds a threshold the system blocks the output and stops streaming. Users therefore keep token-by-token responses, and a harmful answer is cut off part-way rather than delivered whole.

A constitution used twice: to train the model, and to train guards around itConstitutionallowed / disallowedSynthetic data generatorprompts, completions, jailbreak stylesTrain classifiersinput guard, output guardUserpromptInput classifierblock or passModel (CAI-trained)streams tokensOutput classifierscores each chunkHalt streamreplace with refusalTraining shapes what the model wants to do; guards bound what reaches the user when training fails.
The constitution drives synthetic data, which trains an input guard and an output guard placed around a model that was itself trained with constitutional methods. The output guard can stop a response mid-stream.

What the published results show

Anthropic's published results come from three different evaluations, and it is easy to merge them into one claim that none of them made.

EvaluationSystemReported result
Automated: synthetic jailbreak promptsLater system vs. unguarded modelJailbreak success 86% without classifiers, 4.4% with them
Production-like trafficLater system0.38% absolute rise in refusals; 23.7% more inference compute
Bug bounty, two monthsEarlier prototype (higher over-refusal)183 active participants, over 3,000 hours; none answered all ten forbidden queries with one jailbreak
Public demo, February 2025Demo system339 participants, about 3,700 hours; four passed all eight levels, one with what Anthropic judged a universal jailbreak

Read correctly, the results say that constitution-trained guards raise the cost of universal jailbreaks substantially at a modest price in refusals and compute, and that a determined community still broke the demo. That is the right mental model for any guard: it buys time and shrinks the attacker pool, it does not end the contest.

Building a scaled-down version

The same pattern scales down to an application team with a narrow domain. The worked example is a retail-banking assistant. Its constitution forbids operational help with account takeover and recruiting money mules, and explicitly allows the large set of nearby legitimate requests: explaining how scams work so customers can spot them, reporting fraud, and resetting one's own credentials through the official flow.

# constitution.yaml (excerpt)
disallowed:
  - id: ato
    rule: Step-by-step help taking over someone else's account, including
          bypassing verification, SIM swapping or scripting credential stuffing.
  - id: mule
    rule: Help recruiting, scripting or managing people to move stolen funds.
allowed:
  - id: awareness
    rule: Explaining how scams and takeovers work at the level of a bank's public
          fraud-awareness page, so a customer can recognise and report them.
  - id: own_account
    rule: Helping the authenticated customer secure or recover their own account
          through official channels.

Generate data for each rule with a model you trust, and generate the allowed side with the same care, because over-refusal comes from allowed examples that were never seen. Then train a small classifier and calibrate its threshold against a false-refusal budget on real, benign traffic rather than accuracy on the synthetic set.

import numpy as np

STYLES = ["plain", "role-play as a novelist", "base64-wrapped", "split over turns",
          "translated to Spanish", "fake system message"]

def gen_examples(llm, rule, label, n=200):
    out = []
    for style in STYLES:
        prompt = (f"Write {n // len(STYLES)} diverse user messages that fall under this "
                  f"rule, phrased in the style '{style}'. One per line.\nRule: {rule}")
        out += [(t, label) for t in llm(prompt).splitlines() if t.strip()]
    return out

def calibrate(scores_benign, budget=0.005):
    # smallest threshold whose false-refusal rate on benign traffic <= budget
    return float(np.quantile(scores_benign, 1.0 - budget))

def recall_at(threshold, scores_harmful):
    return float(np.mean(np.asarray(scores_harmful) >= threshold))

The output guard must judge partial text, because waiting for the whole response means the harmful part has already streamed. Score the text so far every few dozen tokens and halt as soon as the score crosses the threshold:

def guarded_stream(model_stream, out_clf, threshold, every=32):
    buf, sent = [], 0
    for tok in model_stream:
        buf.append(tok)
        if len(buf) - sent >= every:
            if out_clf("".join(buf)) >= threshold:
                yield "\n[Response stopped: this request falls outside what I can help with.]"
                return
            yield "".join(buf[sent:]); sent = len(buf)
    if out_clf("".join(buf)) < threshold:
        yield "".join(buf[sent:])
    else:
        yield "\n[Response stopped: this request falls outside what I can help with.]"

In this worked example, the team sets a 0.5 percent false-refusal budget on a week of sampled benign chats, then measures recall on a held-out red-team set written by people, not by the generator. A typical outcome is that recall on the synthetic test split looks excellent while recall on the human set is lower, with the misses clustered on multi-turn requests that are benign one message at a time. The fix is to classify the conversation window, not the last message, which is the general lesson: the guard must see the same unit of text the attacker spreads the request across. Off-the-shelf guards such as Llama Guard make a good baseline to compare against before training your own.

Evaluating the guard

A guard is only as trustworthy as its evaluation, and a single accuracy number hides the two failures that matter. Report four numbers for every release, each on a set the classifier never trained on:

MetricMeasured onWhy it matters
Attack success rateHuman-written red-team setThe security claim; synthetic sets flatter the guard
False-refusal rateSampled real benign trafficThe cost customers pay; tied to the budget
Leak lengthRed-team prompts that eventually get haltedHow much text streams before the halt
Added latency, p95Production-shaped loadGuards on the critical path slow every reply
def release_gate(m, prev):
    checks = {
        "asr":     m["attack_success"] <= min(0.05, prev["attack_success"]),
        "refusal": m["false_refusal"] <= 0.005,
        "leak":    m["leak_tokens_p95"] <= 64,
        "latency": m["added_ms_p95"] <= 120,
    }
    return all(checks.values()), [k for k, ok in checks.items() if not ok]

The thresholds in the gate are placeholders to set from your own budgets, but the shape is the point: a new guard ships only if it is no worse than the previous one on attacks and stays inside every cost budget. Pair it with the helpfulness checks described in HHH as an engineering specification, so a safer guard that quietly makes the assistant useless is caught before release.

Failure modes

  • Generator blind spots. Classifiers learn the generator's idea of an attack. Keep a human-written red-team set that never feeds training, and measure on it.
  • Over-refusal on neighbours. A thin allowed side makes the guard block fraud education along with fraud. Track refusals on benign traffic as a first-class metric.
  • Per-message blindness. Requests split across turns pass a last-message guard. Classify a conversation window.
  • Streaming leaks. A large scoring interval lets a harmful paragraph reach the user before the halt. Size the interval from how much text you can tolerate leaking.
  • Constitution drift. Rules edited without regenerating data and recalibrating produce a guard enforcing last quarter's policy. Version the constitution, data and threshold together.
  • Guard as the only control. Account actions still need authorisation checks in the tools themselves; a classifier is not an access-control system.

Operating it and the trade-offs

Budget for the cost the published results make visible: two extra classifier passes add latency and compute, and Anthropic's figure of roughly a quarter more inference compute is a reasonable planning number for a guard of comparable size, smaller for a small classifier. Log every block with the classifier score and constitution version, sample blocks for human review weekly, and feed confirmed false refusals back into the allowed side of the data. Run a standing bounty or internal red-team rotation, since the public demo showed that a large enough pool of attackers eventually finds a way through. The trade-off is a dial, not a switch: lower thresholds stop more attacks and refuse more customers, and the false-refusal budget is the honest way to set it.

What to do next

  1. Write a constitution for your domain with as many allowed rules as disallowed ones.
  2. Generate synthetic examples per rule across several jailbreak styles, and keep a separate human-written red-team set out of training.
  3. Benchmark an existing guard such as Llama Guard on both sets before training your own.
  4. Train input and output classifiers, then calibrate thresholds against a false-refusal budget measured on real benign traffic.
  5. Put the output guard on the stream with a scoring interval you have sized deliberately.
  6. Version constitution, data, model and threshold together, and re-run the evaluation on every change.
  7. Schedule recurring red-teaming and treat each confirmed bypass as new training data.
Key takeaway: A constitution can shape a model in training and also define guards around it at run time. Training lowers the base rate of harmful output; constitution-trained input and output classifiers raise the cost of universal jailbreaks at a measurable price in compute and refusals. Write the allowed side as carefully as the disallowed side, calibrate against a false-refusal budget, guard the stream and the whole conversation, and keep red-teaming.