Most prompt engineering advice arrives as a list of tricks: give the model a role, show it examples, ask it to think step by step, make it return JSON. Each trick works somewhere and fails somewhere else, and a list gives you no way to tell which situation you are in. This article treats prompt engineering patterns the way software engineers treat design patterns: as named, reusable solutions, each one tied to the specific failure it fixes, with a known cost and a known situation where it makes things worse.

You will get a catalogue in three families (input, output and control flow), a small Python library that composes them, a worked example that adds patterns one at a time, and a method for deciding which earn their place. Deep dives on individual patterns are linked; this page is the map.

The pattern map

Observed failurefrom eval traces, not guessesClassify the failureinput, output or controlPick one patterncheapest that targets itInput patternscontract and roledelimited untrusted dataexamples, context packingOutput patternsschema-constrained outputquote, then answerexplicit abstain pathControl-flow patternschain, route, votegenerate then verifyfallback ladderAblation on a fixed eval setkeep the pattern only if the metric moves and the cost is paid forNext failureloop until the error budget is met
Start from an observed failure, classify it by family, apply the cheapest pattern that targets it, and keep it only if an ablation on a fixed eval set says it pays.

From first principles: three families

Start from what a model call is: a function from a sequence of tokens to a distribution over continuations, sampled once. Everything you call a prompt pattern changes one of three things. It changes the input sequence (what the model conditions on), it changes how the output is constrained or read (what counts as a valid answer), or it changes the program around the call (how many calls, in what order, with what checks between them). That gives three families, and the family tells you where to look when a pattern is not working.

Each family has a different cost profile. Input patterns cost prompt tokens, which are cheap per token and often cacheable, but they compete for attention with the actual task. Output patterns cost almost nothing in tokens and buy reliability of parsing, but a tight constraint can force a wrong answer into a valid shape. Control-flow patterns multiply calls, so they multiply latency and spend, but they are the only family that can catch an error after the model has made it. A sensible default order is therefore: fix the input, then constrain the output, and only then add calls. Name each pattern by the failure it fixes, so it comes with a test: if that failure rate does not fall, the pattern is not doing its job.

Input patterns

Input patterns shape what the model conditions on. The table lists the ones that pay most often, with the failure each fixes and the situation where it backfires.

PatternFailure it fixesCostDo not use when
Task contract (goal, audience, constraints, done criteria)Vague or drifting answers; the model optimises for something you did not askA few hundred tokens, cacheableThe task is already fully specified by a schema and the contract just repeats it
Role or personaWrong register or domain vocabularyTinyYou hope the role adds knowledge; it changes tone, not facts
Delimited untrusted dataInstructions inside documents or user input get obeyedTags around each sourceNever skip it for retrieved or user-supplied text
Few-shot examplesFormat or edge-case handling the model cannot infer from a descriptionHundreds to thousands of tokensExamples are unrepresentative; the model copies their surface features
Context packing (most relevant first, deduplicated, labelled)Answer ignores the relevant passage or mixes sourcesRetrieval and ranking workThe context fits easily and is all relevant

Delimiting is a security control as much as a quality one. Tags do not make prompt injection impossible, but they give the model a boundary and your evals something to test; pair them with the instruction hierarchy your provider supports.

Output patterns

Output patterns decide what counts as an answer. The most valuable is the schema-constrained output: define the fields, types and allowed values, and use the structured-output or tool-calling feature of your provider so the decoder cannot emit anything else. This removes an entire class of parse failures. Its failure mode is subtle: if the schema has no way to say "not found", the model will fill the field with something plausible. Every extraction schema therefore needs an explicit abstain path, such as a nullable field or a status enum with a value like insufficient_evidence.

The second output pattern is quote, then answer. Ask the model to copy the exact sentences that support its answer into one field before writing the answer in another. You can then verify the quotes mechanically by checking they appear verbatim in the source, and an answer whose quotes fail the check is rejected without any judgement call. This is cheap and turns a hallucination problem into a string-matching problem.

Control-flow patterns

Control-flow patterns put code around the model call. They are the most powerful and the most expensive, so each needs a clear reason.

  • Chain. Split a task into steps whose intermediate outputs you can check, such as extract, then normalise, then summarise. Fixes: one prompt doing too many jobs and failing on one of them. Cost: one call per step and latency that adds up. Avoid when the steps are not separable and the split loses context.
  • Route. Classify the request first, then send it to a specialised prompt or a smaller model. Fixes: one general prompt that is mediocre at everything, and paying large-model prices for easy requests. Cost: a cheap classification call and misroutes. Avoid when the categories are not stable.
  • Vote (self-consistency). Sample several answers at non-zero temperature and take the majority. Fixes: high-variance reasoning on tasks with a single checkable answer. Cost: N times the spend. Avoid for open-ended text, where there is nothing to vote on.
  • Generate, then verify. A second step, which can be code, a rule or another model, checks the output against the source or the schema and triggers a retry or abstention. Fixes: plausible but unsupported answers. Cost: one extra call or a cheap function. Avoid a model verifier when a deterministic check exists.
  • Fallback ladder. On failure, retry with the error message, then escalate to a stronger model, then abstain or hand off to a human. Fixes: rare hard cases that would otherwise drag everyone to the expensive model. Cost: tail latency on the failures.

Composing patterns in code

Patterns compose, and the composition is easier to reason about as code than as one ever-longer prompt. The library below is deliberately small: every combinator takes and returns a function from a dictionary of inputs to a Result that carries the parsed value and the number of model calls spent, so cost is visible at every level.

from collections import Counter
from dataclasses import dataclass
from typing import Callable, Optional

Model = Callable[[str], str]          # prompt in, text out; wrap your provider client here

@dataclass
class Result:
    value: Optional[dict]
    calls: int
    note: str = ""

def step(model: Model, template: str, parse: Callable[[str], dict]):
    """One model call: fill the template, call, parse. Parse errors propagate."""
    def run(inputs: dict) -> Result:
        return Result(parse(model(template.format(**inputs))), calls=1)
    return run

def chain(*steps):
    def run(inputs: dict) -> Result:
        calls, data = 0, dict(inputs)
        for s in steps:
            r = s(data)
            calls += r.calls
            data.update(r.value)          # each step's output feeds the next
        return Result(data, calls)
    return run

def route(classify, routes: dict, default):
    def run(inputs: dict) -> Result:
        label = classify(inputs).value["label"]
        r = routes.get(label, default)(inputs)
        return Result(r.value, r.calls + 1, note=f"route={label}")
    return run

def vote(s, n: int, key: str):
    def run(inputs: dict) -> Result:
        outs = [s(inputs) for _ in range(n)]
        top, count = Counter(o.value[key] for o in outs).most_common(1)[0]
        winner = next(o.value for o in outs if o.value[key] == top)
        return Result(winner, sum(o.calls for o in outs), note=f"agree={count}/{n}")
    return run

def verify(s, check: Callable[[dict, dict], Optional[str]], retries: int = 1):
    def run(inputs: dict) -> Result:
        calls, inputs = 0, {"feedback": "", **inputs}   # template must accept {feedback}
        for _ in range(retries + 1):
            r = s(inputs)
            calls += r.calls
            error = check(inputs, r.value)
            if error is None:
                return Result(r.value, calls)
            inputs = {**inputs, "feedback": error}
        return Result(None, calls, note=f"abstain: {error}")
    return run

def fallback(*pipelines):
    def run(inputs: dict) -> Result:
        calls = 0
        for pl in pipelines:
            r = pl(inputs)
            calls += r.calls
            if r.value is not None:
                return Result(r.value, calls, r.note)
        return Result(None, calls, "escalate to human")
    return run

Parsing and checking are plain functions, testable without a model, and the call count travels with the result so evals report cost next to accuracy. A pipeline such asfallback(verify(small_extract, quotes_ok), verify(large_extract, quotes_ok)) reads as a sentence: try the small model with a quote check, then the large model with the same check, then hand off.

Worked example: invoice extraction, one pattern at a time

Take invoice field extraction: given the text of an invoice, return the supplier, invoice number, total and due date. Suppose you have 200 labelled invoices and a baseline prompt that simply asks for the four fields as JSON. The numbers below are illustrative, to show the method; your own will differ.

StepChangeField accuracyCalls per invoiceWhat the traces showed
0Baseline prompt81%16% unparseable; dates in mixed formats; totals taken from subtotal lines
1Schema-constrained output with a status field86%1Parse failures gone; date and total errors remain
2Contract: ISO 8601 dates, total means amount payable including tax91%1Remaining errors on invoices with credit notes
3Two examples, one with a credit note93%1Some fabricated invoice numbers on scans with no number
4Quote, then answer, plus quotes_ok verify with one retry95%1.1Fabrications become abstentions
5Fallback to larger model on abstain97%1.3Residual failures are illegible scans
def quotes_ok(inputs: dict, out: dict):
    """Reject any answer whose supporting quote is not verbatim in the source."""
    if out.get("status") == "insufficient_evidence":
        return None                                   # abstaining is a valid answer
    for field, quote in out.get("evidence", {}).items():
        if quote not in inputs["document"]:
            return f"evidence for {field} is not an exact quote; copy it verbatim"
    return None

Each row targets a failure seen in the previous row's traces. Voting was not added because extracted fields rarely vary between samples, and no chain because the task is one read of one document. The average stays near 1.3 calls because the expensive path runs only on the cases the cheap path abstains on.

Choosing patterns by ablation

The method that keeps a prompt lean is ablation: run the full pipeline and the pipeline with each pattern removed on the same fixed dataset, and keep a pattern only if removing it hurts the metric by more than run-to-run noise. Score accuracy, abstention rate and calls per request together, because a pattern that raises accuracy by turning errors into abstentions may or may not be what the product wants.

def ablate(pipelines: dict, dataset, score):
    """Run every candidate pipeline on the same labelled set; report quality and cost."""
    rows = []
    for name, pl in pipelines.items():
        hits = calls = abstains = 0
        for item in dataset:
            r = pl(item["inputs"])
            calls += r.calls
            abstains += r.value is None
            hits += r.value is not None and score(r.value, item["label"])
        n = len(dataset)
        rows.append((name, hits / n, abstains / n, calls / n))
    return sorted(rows, key=lambda row: (-row[1], row[3]))

Run each configuration at least twice if any step samples, treat differences smaller than the run-to-run spread as zero, and rerun the ablation when the model version changes: patterns that compensated for an older model's weakness often become dead weight.

Failure modes

  • Pattern stacking. Every incident adds a rule and nothing is removed, until rules contradict. Fix: ablate on a schedule.
  • Examples that leak. The model copies names, lengths or phrasing from few-shot examples into answers. Fix: vary surface features across examples and test on inputs unlike them.
  • Constraint without escape. A schema with no abstain value forces invented answers. Fix: a status enum or nullable fields, and evals that include unanswerable inputs.
  • Voting on correlated errors. Five samples that share a misunderstanding agree confidently. Agreement measures variance, not correctness.
  • Verifier collusion. The same model, prompted similarly, approves its own mistakes. Prefer deterministic checks, then a different prompt or model.
  • Unbounded retries. Retry loops without a cap turn one bad input into runaway spend. Cap retries and log every abstention.

Trade-offs

Input patterns spend tokens and are the cheapest place to start, especially when the static prefix is cached. Output patterns spend flexibility, since a schema cannot express an answer it did not anticipate. Control-flow patterns spend calls and latency, and their benefit concentrates on the hard tail, which is why a fallback ladder usually beats running the expensive path for everyone. Composed functions are easier to test than one long prompt but are more things to version, so version the pipeline as a unit with its eval results attached.

Related reading

Deep dives for patterns in this catalogue: prompt chaining, routing and self-consistency voting. For the output side see structured output, for the measurement loop see evaluation harnesses and regression testing, and for how prompting changed with reasoning models see prompt engineering in 2026.

What to do next

  1. Collect 50 to 200 real inputs with labels for one task and write a scoring function before touching the prompt.
  2. Run your current prompt, read at least 30 failing traces and classify each failure as input, output or control flow.
  3. Fix the largest class with the cheapest pattern from its family, then rerun and compare.
  4. Add a schema with an explicit abstain value to every extraction or classification task.
  5. Replace any model-based check that could be a deterministic one, such as verbatim quotes, regexes or range checks.
  6. Run an ablation of the final pipeline and delete patterns that do not move the metric beyond noise.
  7. Read the deep dives linked in this article for the patterns you kept, and rerun the ablation whenever the model version changes.
Key takeaway: Treat prompt patterns as named fixes for observed failures. Input patterns such as contracts, delimiters and examples change what the model conditions on and are cheap. Output patterns such as schemas, quote-then-answer and an explicit abstain value make answers checkable. Control-flow patterns such as chains, routers, voting, verification and fallback ladders add calls and are the only way to catch an error after it happens. Compose them as small testable functions that report their own cost, add one at a time against a measured failure, and keep each only if an ablation on a fixed dataset shows it pays for itself.