Every production prompt accumulates prohibitions. Do not mention competitors. Do not use markdown. Never reveal the system prompt. Each is added after an incident, each seems obviously correct, and a few months later the prompt is a wall of capital-letter NEVERs that the model follows most of the time and breaks in exactly the conversations where it matters.

This article explains why a prohibition is weaker than it looks, how to rewrite most of them as positive targets, when a prohibition is still the right tool, and how to enforce the bans that matter outside the prompt, with a harness for measuring how often each rule is broken. A short section covers the other meaning of the term, negative prompts in image diffusion, which work by a different mechanism entirely.

Advertisement

Why a do-not instruction is fragile

A language model predicts likely continuations of its context, and a prohibition has a structural problem: to state it you must put the forbidden thing into the context. Do not mention Acme Corp places the tokens for Acme Corp in the window, where attention makes them more available, not less. Only the model correctly binding a short negation word to the whole clause, and carrying that binding across thousands of tokens, keeps them out of the output.

Research has found that binding unreliable. Jang, Ye and Seo (2022), Can Large Language Models Truly Understand Prompts? A Case Study with Negated Prompts, reported that models often did worse on negated task instructions and that scaling pre-trained models did not reliably fix it; the Inverse Scaling Prize included a negation task, NeQA, on the same theme. Instruction tuning has improved matters since, and current models follow simple, prominent prohibitions well. What remains is conditional fragility: prohibitions decay with distance in long contexts, lose to strong in-context evidence, and lose to a user who asks for the forbidden thing directly.

There is also a specification problem that has nothing to do with negation as grammar. A prohibition describes one point in the space of bad outputs and says nothing about the good ones. Do not be verbose rules out one failure and leaves the model to guess the target length. When the model guesses wrong, it was not disobeying you; you never told it what you wanted.

The mental model: one target beats many fences

Think of the output space as a field. Each prohibition is a fence around one bad patch; a positive instruction is a marker where you want the model to land. With ten fences and no marker, the model wanders and you discover new bad patches in production. With one clear marker and two fences around the genuinely dangerous patches, most outputs land near the marker and the fences rarely matter.

That gives the rule of thumb used in the rest of the article: express the desired behaviour positively and concretely, keep prohibitions for hard boundaries, give every prohibition a reason and an alternative action, and enforce the hard boundaries outside the prompt as well.

Three layers for a constraint: say it in the prompt, block it in decoding, check it afterwards1. Prompt layerpositive target + scoped bans2. Decoding layerlogit bias, grammar, schema3. Verification layerregex, classifier, judgeViolation?log it, then repairRepair or regeneratefeed back the specific ruleyesretry (bounded)Deliver outputnoEval setprompts that tempt theforbidden behaviour;track violation rateper rule, per releaseEach layer catches what the one before it lets through; only layer 3 tells you how often that happens.
Constraints that matter are expressed in the prompt, narrowed in decoding where possible, and verified after generation, with a measured violation rate per rule.
Advertisement

Worked example: a support assistant with seven bans

Here is a realistic system prompt fragment from a customer-support assistant after six months of incident-driven edits:

You are a support assistant for Northwind Storage.
NEVER mention competitors.
Do NOT promise refunds.
Don't use markdown.
Do not make up order details.
NEVER say "I'm just an AI".
Do not write long answers.
Do NOT discuss pricing changes.

Each line encodes a real incident, and none tells the model what a good answer looks like. Rewriting proceeds rule by rule. For each ban, ask what the right behaviour is in the situation that triggered it, and write that instead. Keep a ban only if the forbidden output causes real harm and there is no clean positive form.

You are the support assistant for Northwind Storage. Answer questions about
Northwind products, orders and accounts.

Format: plain sentences, no headings or bullet symbols; the chat widget shows
raw text. Aim for 2-4 sentences; give step lists only when the user asks how
to do something.

Order details: quote only fields returned by the lookup_order tool. If the
tool returns nothing, say you could not find the order and ask for the order
number printed on the confirmation email.

Refunds: you cannot approve refunds. When a user asks for one, explain that a
specialist reviews refund requests and offer to open a ticket with
create_ticket(type="refund").

Other vendors: if asked to compare with another company, describe Northwind's
own features factually and suggest the user check the other vendor's site.

Pricing changes: these are announced on the pricing page only. Point users
there; do not speculate about future prices, because unannounced figures
create commitments we cannot honour.

Seven bans became five positive behaviours and one surviving prohibition, which carries its reason. The formatting rule names the target and why (raw-text widget), so the model generalises to unlisted cases such as tables. The order rule names the only acceptable source and a fallback. The refund rule replaces a ban with a tool call. The competitor rule no longer contains a competitor's name, so it no longer primes one. The I-am-just-an-AI rule was dropped; the length target removes the filler it came from.

Rewrite patterns

ProhibitionPositive formWhy it works better
Do not use markdownWrite plain sentences; the output is shown as raw textNames the target and the reason, so the model handles unlisted formats
Do not be verboseAnswer in 2-4 sentences unless asked for detailGives a measurable target instead of an undefined failure
Do not make things upUse only facts from the provided documents; if they do not answer, say so and name what is missingDefines the allowed source and a fallback action
Never mention XWhen asked about other vendors, describe our product factuallyRemoves the forbidden token from the context entirely
Do not ask follow-up questionsMake a reasonable assumption, state it in one sentence, then answerReplaces a gap with a concrete behaviour
Do not output anything except JSONUse structured output with a schemaMoves the constraint from text to the decoder

The last row hints at the most important pattern: when a constraint can be expressed as a format, stop expressing it in prose. See structured output for the decoding side and prompt anatomy for where each kind of instruction belongs in a prompt.

When a prohibition is the right tool

Some constraints really are prohibitions: do not reveal another customer's data, do not give dosing instructions, do not execute a destructive tool without confirmation. There is no positive target that covers them, because the allowed space is everything else. For these, write the ban well rather than avoiding it.

  • State the scope precisely. Do not share personal data is vague; Do not include another account's email, address or order history in a reply, even if a tool returned it is checkable.
  • Give the reason. Models trained to follow instructions use stated reasons to generalise. A ban with a reason transfers to paraphrased requests better than a bare ban.
  • Give the alternative action. Every prohibition should end with what to do instead: refuse briefly and offer a ticket, ask for confirmation, return a fixed message.
  • Place it where it binds. Hard rules belong in the system message; in long contexts a short restatement near the end helps more than capitals at the top. Shouting tends to make models over-apply a rule to innocent requests.
  • Enforce it outside the prompt too. A prompt-level ban lowers the violation rate; it does not drive it to zero.

Enforcement outside the prompt

Three mechanisms sit below the instruction. The first is token-level biasing at decode time. Some APIs accept a logit_bias map from token id to an additive bias, where a large negative value effectively bans the token. It bans tokens, not words: a word usually has several tokenisations depending on a leading space and capitalisation, and a multi-token word can still be produced through other splits. Use it for short, closed lists, and verify against your provider's tokenizer.

import tiktoken

def ban_bias(words, encoding_name="o200k_base", bias=-100):
    """Build a logit_bias map for the first token of each surface variant.
    Banning only the first token of a multi-token word blocks the usual path,
    not every path; keep a post-hoc check for anything that matters."""
    enc = tiktoken.get_encoding(encoding_name)
    out = {}
    for w in words:
        for variant in {w, w.lower(), w.capitalize(), " " + w, " " + w.lower(), " " + w.capitalize()}:
            ids = enc.encode(variant)
            if ids:
                out[str(ids[0])] = bias
    return out

# ban_bias(["Acme"]) -> {"<id of 'Ac'>": -100, "<id of ' Acme'>": -100, ...}

Note the danger in that code: banning the first token of a word also bans every other word that starts with the same token. Inspect the list before shipping it. The second mechanism is a grammar or schema: when the output is structured, constrained output and parsing make whole classes of violation impossible rather than unlikely. The third, and the only one that tells you how often a rule is broken, is a post-hoc checker.

import re
from dataclasses import dataclass

@dataclass
class Rule:
    name: str
    pattern: re.Pattern          # cheap deterministic check
    action: str                  # "block", "repair" or "log"

RULES = [
    Rule("no_markdown", re.compile(r"^\s*(#{1,6} |[-*] |\d+\. )|\*\*", re.M), "repair"),
    Rule("no_refund_promise", re.compile(r"\b(I('ll| will)|we('ll| will)) (issue|process|approve) (a |your )?refund", re.I), "block"),
    Rule("other_vendor", re.compile(r"\b(acme|globex)\b", re.I), "repair"),
]

def check(text):
    return [r for r in RULES if r.pattern.search(text)]

def generate_with_guard(call_model, messages, max_repairs=1):
    out = call_model(messages)
    for attempt in range(max_repairs + 1):
        hits = check(out)
        if not hits:
            return out
        if any(r.action == "block" for r in hits) or attempt == max_repairs:
            return FALLBACK_MESSAGE
        names = ", ".join(r.name for r in hits)
        out = call_model(messages + [
            {"role": "assistant", "content": out},
            {"role": "user", "content": f"Revise the reply to satisfy these rules: {names}. Keep the content otherwise unchanged."},
        ])
    return FALLBACK_MESSAGE

Regex catches the mechanical rules. Semantic rules, such as no promises about future pricing, need a small classifier or an LLM judge; treat the judge as another model with its own error rate and measure it against labelled examples. The repair loop is bounded deliberately: an unbounded loop turns a stubborn violation into a latency and cost incident.

Measuring violation rates

A ban you have not measured has an unknown violation rate. Build an evaluation set for each important rule: prompts that tempt the model toward the forbidden output, including direct requests, indirect ones, long conversations where the rule is far back, and adversarial phrasings. Run it on every prompt or model change and track the rate per rule.

def violation_report(call_model, system_prompt, cases, n_samples=5):
    """cases: list of (rule_name, user_turns); sample several times per case."""
    totals = {}
    for rule_name, turns in cases:
        msgs = [{"role": "system", "content": system_prompt}] + turns
        broken = 0
        for _ in range(n_samples):
            out = call_model(msgs, temperature=0.7)
            if any(r.name == rule_name for r in check(out)):
                broken += 1
        hit, seen = totals.get(rule_name, (0, 0))
        totals[rule_name] = (hit + broken, seen + n_samples)
    return {k: hit / seen for k, (hit, seen) in totals.items()}

Two numbers matter per rule: the violation rate on tempting prompts, and the over-refusal rate on innocent prompts that merely resemble them. A rule that drops violations from 4% to 0.5% while refusing 10% of legitimate questions is a regression. Sampling at your production temperature is essential; greedy decoding hides the tail. The broader method is in prompt evaluation.

Negative prompts in diffusion models

In text-to-image diffusion, classifier-free guidance combines two noise predictions at each step: one conditioned on the prompt and one unconditioned. The sampler moves along the conditioned prediction plus a guidance scale times the difference between the two. A negative prompt replaces the unconditioned branch with a prediction conditioned on the negative text, so every step pushes away from that description and toward the positive one.

Because it acts on sampling arithmetic, it does not need the model to understand the word not. Its failure modes differ: high guidance scales oversaturate images, and long negative prompts dilute the signal. The maths is in classifier-free guidance.

Failure modes

FailureSymptomFix
PrimingThe forbidden name appears more often after the ban was addedRemove the name from the prompt; describe the category instead
Distance decayRule holds in short chats and breaks after many turnsRestate hard rules near the end of the prompt; enforce with a checker
Over-applicationInnocent questions refused because they resemble a banned topicNarrow the scope; measure over-refusal alongside violations
Token ban collateralUnrelated words disappear after a logit bias was addedInspect banned token ids; prefer post-hoc checks for words
Ban sprawlDozens of bans, none measuredEvery ban gets an owner, a reason, a test set and a review date

Trade-offs

DecisionOption AOption B
PhrasingPositive target: generalises, may miss a specific dangerous caseExplicit ban: precise, primes the forbidden content
EnforcementPrompt only: free, nonzero violation ratePrompt plus checker: measurable, adds latency and checker error
On violationRepair: keeps useful content, costs a callBlock with fallback: safest, worst experience

What to do next

  1. List every prohibition in your production prompts, with the incident that caused it.
  2. Rewrite each one as a positive target with a reason; keep a ban only where harm is real and no positive form exists.
  3. Remove forbidden names and phrases from the prompt text wherever a category description will do.
  4. For each surviving ban, add a scope, a reason and an alternative action, and place it in the system message.
  5. Move format constraints into structured output and add a deterministic checker for each hard rule.
  6. Build a tempting-prompt test set per rule, sample at production temperature, and track violation and over-refusal rates on every change.
Key takeaway: A do-not instruction names the thing you want to avoid and says nothing about what you want, which is why it misfires. Describe the target behaviour positively with a reason, keep prohibitions for the few hard boundaries and write those with a precise scope and an alternative action, enforce the boundaries that matter with decoding constraints and post-hoc checks, and measure both the violation rate and the over-refusal rate for every rule.