Constitutional AI is best known as a training method. Anthropic introduced it in 2022 to make a model harmless without having people label thousands of harmful outputs: the model critiqued and revised its own answers against a short list of written principles, and later a model compared pairs of answers against those principles to produce the preference data for reinforcement learning. The principles, the constitution, replaced most of the human harm labels.

You probably are not training a frontier model, but the core mechanism transfers directly to prompting. A written set of principles, a critique step that checks an answer against a few of them, and a revision step that fixes only what the critique found is a practical pattern for any application with rules that must hold: regulated advice, brand voice, privacy, grounding. This page shows how to build that pattern, how to use the same principles to produce preference data, where it fails and what it costs. For the training-time story in more depth, see Constitutional AI and alignment.

Advertisement

What the original method actually did

The paper, Constitutional AI: Harmlessness from AI Feedback by Bai and colleagues, has two stages. In the supervised stage, a helpful model is given prompts designed to elicit harmful answers. For each answer, a principle is sampled from the constitution, the model is asked to critique its answer against that principle and then to revise it. Repeating this produces a revised answer, and the model is fine-tuned on the revisions. In the reinforcement learning stage, the fine-tuned model produces two answers to a prompt, a model is asked which one better follows a sampled principle, and those AI judgements train a preference model that serves as the reward. The authors called this reinforcement learning from AI feedback.

Three details carry over to prompting. First, principles were sampled, not all applied at once, which keeps each critique focused. Second, a principle was written as a pair: a critique request and a matching revision request. Third, the stated aim was a model that is harmless without being evasive; the authors explicitly wanted it to engage and explain its objections rather than refuse by reflex. A constitution that only adds prohibitions drifts toward over-refusal, which is the most common failure of the pattern described below.

Writing a constitution that works

A principle is useful only if a model can check it against a specific answer and a different model call can act on the result. That rules out vague values such as be ethical. It favours narrow, observable questions with a quotable violation. Each principle gets an identifier, a scope and the two prompts. Keep the file under version control alongside the prompts that use it, because changing a principle changes behaviour as much as changing the system prompt does.

# constitution.yaml -- versioned with the prompts that use it
version: 2026-10-02
principles:
  - id: no_dosing
    applies_to: [medical]
    critique: >
      Does the response tell the user how much of a medicine to take, or change a
      dose, beyond repeating the label or pointing to a pharmacist? Quote the passage.
    revision: >
      Remove any personal dosing advice. Keep label facts, and direct the user to a
      pharmacist or prescriber for anything specific to them.
  - id: not_evasive
    applies_to: [all]
    critique: >
      Does the response refuse or hedge when a safe, useful answer was possible?
      Quote the hedge and say what a helpful answer would have contained.
    revision: >
      Answer the safe part of the question directly and specifically. Decline only
      the part that is actually unsafe, and say briefly why.
  - id: grounded
    applies_to: [all]
    critique: >
      Does the response state a fact that is not supported by the provided context?
    revision: >
      Remove or qualify unsupported claims. Do not add new facts.

Notice the second principle. Every constitution should contain at least one principle that pushes the other way, against needless refusal and hedging, or the loop will only ever make answers more cautious. Notice also the third: a critique step can enforce grounding against supplied context, which overlaps with retrieval checks but costs only a model call.

Keep principles independent. If two can conflict, for example full disclosure and privacy, say which wins inside the principle text rather than leaving the critic to decide. Ten to thirty principles is a practical range for an application; scoping by domain keeps each request's candidate pool small.

Advertisement

The loop in code

The loop has four steps. Generate a draft with your normal prompt. Sample a few applicable principles. Ask for a structured critique per principle, and revise only if a violation reaches a severity threshold. Stop when a round is clean or a round limit is reached. Structured critique output matters: a JSON verdict with a quoted piece of evidence is far easier to log, test and threshold than free text.

A constitution-guided critique and revision loop at inference timeUser requestplus contextDraftnormal promptSample principlesrelevant subsetCritiqueper principle, JSONAny violation?severity at or above barRevisefix only what was flaggedStop checkclean or max roundsFinal answerplus audit logConstitution fileversioned principles, each a critique and a revision requestyesdonenext roundOnly the critique and revision steps read the constitution; the draft prompt stays your ordinary system prompt.
The inference-time loop. The draft uses your ordinary prompt; principles are sampled from a versioned constitution and only flagged problems are passed to the revision step.

import json, random

def critique(llm, request, answer, principle):
    prompt = (
        "You are reviewing an assistant response against one principle.\n"
        f"PRINCIPLE: {principle['critique']}\n"
        f"<request>{request}</request>\n<response>{answer}</response>\n"
        'Reply with JSON only: {"violation": true|false, "severity": 0-3, '
        '"evidence": "exact quote or empty", "reason": "one sentence"}'
    )
    return json.loads(llm(prompt, temperature=0))

def revise(llm, request, answer, findings):
    notes = "\n".join(
        f"- [{f['id']}] {f['reason']} Evidence: {f['evidence']} Fix: {f['revision']}"
        for f in findings
    )
    prompt = (
        "Rewrite the response to fix ONLY the problems listed. Keep every correct, "
        "useful detail. Do not mention the review.\n"
        f"<request>{request}</request>\n<response>{answer}</response>\n"
        f"<problems>\n{notes}\n</problems>"
    )
    return llm(prompt, temperature=0.2)

def constitutional_answer(llm, request, draft, constitution, domain,
                          k=3, max_rounds=2, min_severity=2):
    pool = [p for p in constitution if domain in p["applies_to"] or "all" in p["applies_to"]]
    answer, log = draft, []
    for round_no in range(max_rounds):
        sampled = random.sample(pool, min(k, len(pool)))
        findings = []
        for p in sampled:
            verdict = critique(llm, request, answer, p)
            log.append({"round": round_no, "id": p["id"], **verdict})
            if verdict["violation"] and verdict["severity"] >= min_severity:
                findings.append({**verdict, "id": p["id"], "revision": p["revision"]})
        if not findings:
            break                        # clean under this sample of principles
        answer = revise(llm, request, answer, findings)
    return answer, log

Design choices in this code are deliberate. Critique runs at temperature zero so the same answer gets the same verdict. The revision prompt says fix only the listed problems, because an unconstrained rewrite tends to drop correct details. Severity filtering stops trivial findings from triggering a rewrite. And the log keeps every verdict, which becomes your evaluation data later.

A worked example

Take a pharmacy chain's support assistant. A user writes: I have a cold and take ibuprofen for back pain, can I also take a cold remedy, and how much ibuprofen is safe? The draft, from a general system prompt, explains that many cold remedies already contain a painkiller and then says the user can take 400 mg every four hours.

The router tags the request as medical, so the pool contains no_dosing, not_evasive and grounded. The critic flags no_dosing at severity 3, quoting the dose sentence. It passes not_evasive, because the answer engaged with the question. It flags grounded at severity 2, because the store's leaflet in the context gives a different maximum interval from the one the draft stated.

The revision keeps the useful warning about duplicated ingredients, tells the user to check the active ingredients on the cold remedy's label, removes the personal dosing statement, quotes the leaflet's label limits as label limits, and suggests asking a pharmacist, who can see what else the user takes. A second round samples the same principles and comes back clean, so the loop stops. Two critique rounds and one revision cost roughly four extra model calls on top of the draft.

How it differs from other self-correction patterns

Several patterns on this site make a model check its own work, and they are easy to confuse. Reflexion learns from task feedback across attempts: an agent fails, writes a reflection and tries again, so the signal is whether the task succeeded. Chain of verification targets factual accuracy by generating and answering verification questions about the draft's claims. A verifier scores or accepts candidate outputs, usually for correctness.

The constitutional pattern is different in what drives the check: an explicit, written, versioned set of normative rules about how an answer should behave, sampled per request. It is the right tool when the requirement is a policy, not a fact or a task outcome. In practice they combine well: a grounding check from chain of verification can be one principle among many, and a verifier can sit after the loop as a final gate.

Using the constitution for AI feedback data

The second half of the original method, AI preference labelling, is just as useful without reinforcement learning. Generate two candidate answers, ask a judge model which better satisfies a sampled principle, and you have labelled pairs for evaluation sets or for preference fine-tuning such as DPO on a smaller model. The weak point is judge bias, especially position bias, so ask twice with the order swapped and keep only consistent verdicts:

def ai_preference(llm, request, a, b, principle):
    # Ask which of two responses better satisfies one principle. Swap order and
    # ask again; keep the pair only if both orderings agree (controls position bias).
    def ask(x, y):
        out = llm(
            f"Principle: {principle['critique']}\nRequest: {request}\n"
            f"(A) {x}\n(B) {y}\nWhich response better satisfies the principle? "
            "Think step by step, then end with exactly 'ANSWER: A' or 'ANSWER: B'.",
            temperature=0,
        )
        return out.strip()[-1]
    first, second = ask(a, b), ask(b, a)
    if first == "A" and second == "B":
        return {"chosen": a, "rejected": b, "principle": principle["id"]}
    if first == "B" and second == "A":
        return {"chosen": b, "rejected": a, "principle": principle["id"]}
    return None   # inconsistent judgement: drop it rather than train on noise

Asking the judge to reason before answering improved AI labels in the original work too. Before trusting any of this data, have people label a sample of a few hundred pairs and measure agreement with the judge per principle; principles where agreement is poor need rewriting, not more data.

Failure modes

  • Over-refusal. Each round can only remove content, so answers grow timid. Include anti-evasion principles and measure refusal rate on benign prompts as a first-class metric.
  • Critic sycophancy and nit-picking. Asked to find problems, a model finds some. Severity thresholds, quoted evidence and a requirement that the evidence actually appears in the answer cut false positives.
  • Revision drift. Rewrites lose facts, change numbers or add new claims. Constrain the revision prompt, and diff key entities between draft and final answer.
  • Oscillation. Two principles push against each other and the answer flips between rounds. Cap rounds and resolve the conflict in the constitution text.
  • Injection into the critic. User text inside the response being reviewed can try to instruct the critic. Delimit the reviewed content, tell the critic to treat it as data, and keep the critic's output schema strict.
  • Silent cost growth. Every principle sampled is a call. Without budgets, adding principles multiplies spend and latency.

Running it in production

Cost scales with principles sampled times rounds, plus revisions. With three principles and up to two rounds, the worst case is nine calls per request instead of one. Three tactics keep this sane. Run the loop only where it is needed, using a cheap classifier or router to decide which requests are in a regulated domain. Use a smaller, cheaper model as the critic once its verdicts agree with a larger one on your evaluation set. And batch the critiques for one round into a single structured call when the principles are short, accepting slightly less focus for a third of the cost.

For latency, stream nothing until the loop finishes, or stream the draft only when the domain is low-risk and run the check asynchronously for monitoring. Log every verdict with the constitution version so you can see which principles fire, how often they fire on benign traffic and whether a prompt change moved those rates. Treat those logs as an evaluation suite; the prompt evaluation page covers how to build regression gates from them.

What to do next

  1. List the five to ten rules your application must never break, and write each as a critique request with a quotable violation plus a matching revision request.
  2. Add at least one anti-evasion principle and one grounding principle.
  3. Implement the loop with structured verdicts, severity thresholds and a two-round cap, and log every verdict with the constitution version.
  4. Build a test set of risky and benign prompts, and track violation rate, refusal rate and answer quality before and after the loop.
  5. Label a sample of critiques by hand and fix principles where the critic disagrees with people.
  6. Gate the loop to the domains that need it and try a smaller critic model once agreement is measured.
  7. If you fine-tune, reuse the constitution with order-swapped AI preference labels to build preference pairs.
Key takeaway: Constitutional AI replaced most human harm labels with a written list of principles that a model applies to itself. The same mechanism works at inference time: sample a few relevant principles, critique the draft with quoted evidence, revise only what was flagged, stop when clean, and log everything. It shines when requirements are policies rather than facts, it fails toward over-caution unless the constitution argues for helpfulness too, and it costs extra model calls that you should spend only where the rules matter.