Most prompt engineering advice arrives as a list of tricks: give the model a role, show it examples, ask it to think step by step, make it return JSON. Each trick works somewhere and fails somewhere else, and a list gives you no way to tell which situation you are in. This article treats prompt engineering patterns the way software engineers treat design patterns: as named, reusable solutions, each one tied to the specific failure it fixes, with a known cost and a known situation where it makes things worse.
You will get a catalogue in three families (input, output and control flow), a small Python library that composes them, a worked example that adds patterns one at a time, and a method for deciding which earn their place. Deep dives on individual patterns are linked; this page is the map.
The pattern map
From first principles: three families
Start from what a model call is: a function from a sequence of tokens to a distribution over continuations, sampled once. Everything you call a prompt pattern changes one of three things. It changes the input sequence (what the model conditions on), it changes how the output is constrained or read (what counts as a valid answer), or it changes the program around the call (how many calls, in what order, with what checks between them). That gives three families, and the family tells you where to look when a pattern is not working.
Each family has a different cost profile. Input patterns cost prompt tokens, which are cheap per token and often cacheable, but they compete for attention with the actual task. Output patterns cost almost nothing in tokens and buy reliability of parsing, but a tight constraint can force a wrong answer into a valid shape. Control-flow patterns multiply calls, so they multiply latency and spend, but they are the only family that can catch an error after the model has made it. A sensible default order is therefore: fix the input, then constrain the output, and only then add calls. Name each pattern by the failure it fixes, so it comes with a test: if that failure rate does not fall, the pattern is not doing its job.
Input patterns
Input patterns shape what the model conditions on. The table lists the ones that pay most often, with the failure each fixes and the situation where it backfires.
| Pattern | Failure it fixes | Cost | Do not use when |
|---|---|---|---|
| Task contract (goal, audience, constraints, done criteria) | Vague or drifting answers; the model optimises for something you did not ask | A few hundred tokens, cacheable | The task is already fully specified by a schema and the contract just repeats it |
| Role or persona | Wrong register or domain vocabulary | Tiny | You hope the role adds knowledge; it changes tone, not facts |
| Delimited untrusted data | Instructions inside documents or user input get obeyed | Tags around each source | Never skip it for retrieved or user-supplied text |
| Few-shot examples | Format or edge-case handling the model cannot infer from a description | Hundreds to thousands of tokens | Examples are unrepresentative; the model copies their surface features |
| Context packing (most relevant first, deduplicated, labelled) | Answer ignores the relevant passage or mixes sources | Retrieval and ranking work | The context fits easily and is all relevant |
Delimiting is a security control as much as a quality one. Tags do not make prompt injection impossible, but they give the model a boundary and your evals something to test; pair them with the instruction hierarchy your provider supports.
Output patterns
Output patterns decide what counts as an answer. The most valuable is the schema-constrained output: define the fields, types and allowed values, and use the structured-output or tool-calling feature of your provider so the decoder cannot emit anything else. This removes an entire class of parse failures. Its failure mode is subtle: if the schema has no way to say "not found", the model will fill the field with something plausible. Every extraction schema therefore needs an explicit abstain path, such as a nullable field or a status enum with a value like insufficient_evidence.
The second output pattern is quote, then answer. Ask the model to copy the exact sentences that support its answer into one field before writing the answer in another. You can then verify the quotes mechanically by checking they appear verbatim in the source, and an answer whose quotes fail the check is rejected without any judgement call. This is cheap and turns a hallucination problem into a string-matching problem.
Control-flow patterns
Control-flow patterns put code around the model call. They are the most powerful and the most expensive, so each needs a clear reason.
- Chain. Split a task into steps whose intermediate outputs you can check, such as extract, then normalise, then summarise. Fixes: one prompt doing too many jobs and failing on one of them. Cost: one call per step and latency that adds up. Avoid when the steps are not separable and the split loses context.
- Route. Classify the request first, then send it to a specialised prompt or a smaller model. Fixes: one general prompt that is mediocre at everything, and paying large-model prices for easy requests. Cost: a cheap classification call and misroutes. Avoid when the categories are not stable.
- Vote (self-consistency). Sample several answers at non-zero temperature and take the majority. Fixes: high-variance reasoning on tasks with a single checkable answer. Cost: N times the spend. Avoid for open-ended text, where there is nothing to vote on.
- Generate, then verify. A second step, which can be code, a rule or another model, checks the output against the source or the schema and triggers a retry or abstention. Fixes: plausible but unsupported answers. Cost: one extra call or a cheap function. Avoid a model verifier when a deterministic check exists.
- Fallback ladder. On failure, retry with the error message, then escalate to a stronger model, then abstain or hand off to a human. Fixes: rare hard cases that would otherwise drag everyone to the expensive model. Cost: tail latency on the failures.
Composing patterns in code
Patterns compose, and the composition is easier to reason about as code than as one ever-longer prompt. The library below is deliberately small: every combinator takes and returns a function from a dictionary of inputs to a Result that carries the parsed value and the number of model calls spent, so cost is visible at every level.
from collections import Counter
from dataclasses import dataclass
from typing import Callable, Optional
Model = Callable[[str], str] # prompt in, text out; wrap your provider client here
@dataclass
class Result:
value: Optional[dict]
calls: int
note: str = ""
def step(model: Model, template: str, parse: Callable[[str], dict]):
"""One model call: fill the template, call, parse. Parse errors propagate."""
def run(inputs: dict) -> Result:
return Result(parse(model(template.format(**inputs))), calls=1)
return run
def chain(*steps):
def run(inputs: dict) -> Result:
calls, data = 0, dict(inputs)
for s in steps:
r = s(data)
calls += r.calls
data.update(r.value) # each step's output feeds the next
return Result(data, calls)
return run
def route(classify, routes: dict, default):
def run(inputs: dict) -> Result:
label = classify(inputs).value["label"]
r = routes.get(label, default)(inputs)
return Result(r.value, r.calls + 1, note=f"route={label}")
return run
def vote(s, n: int, key: str):
def run(inputs: dict) -> Result:
outs = [s(inputs) for _ in range(n)]
top, count = Counter(o.value[key] for o in outs).most_common(1)[0]
winner = next(o.value for o in outs if o.value[key] == top)
return Result(winner, sum(o.calls for o in outs), note=f"agree={count}/{n}")
return run
def verify(s, check: Callable[[dict, dict], Optional[str]], retries: int = 1):
def run(inputs: dict) -> Result:
calls, inputs = 0, {"feedback": "", **inputs} # template must accept {feedback}
for _ in range(retries + 1):
r = s(inputs)
calls += r.calls
error = check(inputs, r.value)
if error is None:
return Result(r.value, calls)
inputs = {**inputs, "feedback": error}
return Result(None, calls, note=f"abstain: {error}")
return run
def fallback(*pipelines):
def run(inputs: dict) -> Result:
calls = 0
for pl in pipelines:
r = pl(inputs)
calls += r.calls
if r.value is not None:
return Result(r.value, calls, r.note)
return Result(None, calls, "escalate to human")
return runParsing and checking are plain functions, testable without a model, and the call count travels with the result so evals report cost next to accuracy. A pipeline such asfallback(verify(small_extract, quotes_ok), verify(large_extract, quotes_ok)) reads as a sentence: try the small model with a quote check, then the large model with the same check, then hand off.
Worked example: invoice extraction, one pattern at a time
Take invoice field extraction: given the text of an invoice, return the supplier, invoice number, total and due date. Suppose you have 200 labelled invoices and a baseline prompt that simply asks for the four fields as JSON. The numbers below are illustrative, to show the method; your own will differ.
| Step | Change | Field accuracy | Calls per invoice | What the traces showed |
|---|---|---|---|---|
| 0 | Baseline prompt | 81% | 1 | 6% unparseable; dates in mixed formats; totals taken from subtotal lines |
| 1 | Schema-constrained output with a status field | 86% | 1 | Parse failures gone; date and total errors remain |
| 2 | Contract: ISO 8601 dates, total means amount payable including tax | 91% | 1 | Remaining errors on invoices with credit notes |
| 3 | Two examples, one with a credit note | 93% | 1 | Some fabricated invoice numbers on scans with no number |
| 4 | Quote, then answer, plus quotes_ok verify with one retry | 95% | 1.1 | Fabrications become abstentions |
| 5 | Fallback to larger model on abstain | 97% | 1.3 | Residual failures are illegible scans |
def quotes_ok(inputs: dict, out: dict):
"""Reject any answer whose supporting quote is not verbatim in the source."""
if out.get("status") == "insufficient_evidence":
return None # abstaining is a valid answer
for field, quote in out.get("evidence", {}).items():
if quote not in inputs["document"]:
return f"evidence for {field} is not an exact quote; copy it verbatim"
return NoneEach row targets a failure seen in the previous row's traces. Voting was not added because extracted fields rarely vary between samples, and no chain because the task is one read of one document. The average stays near 1.3 calls because the expensive path runs only on the cases the cheap path abstains on.
Choosing patterns by ablation
The method that keeps a prompt lean is ablation: run the full pipeline and the pipeline with each pattern removed on the same fixed dataset, and keep a pattern only if removing it hurts the metric by more than run-to-run noise. Score accuracy, abstention rate and calls per request together, because a pattern that raises accuracy by turning errors into abstentions may or may not be what the product wants.
def ablate(pipelines: dict, dataset, score):
"""Run every candidate pipeline on the same labelled set; report quality and cost."""
rows = []
for name, pl in pipelines.items():
hits = calls = abstains = 0
for item in dataset:
r = pl(item["inputs"])
calls += r.calls
abstains += r.value is None
hits += r.value is not None and score(r.value, item["label"])
n = len(dataset)
rows.append((name, hits / n, abstains / n, calls / n))
return sorted(rows, key=lambda row: (-row[1], row[3]))Run each configuration at least twice if any step samples, treat differences smaller than the run-to-run spread as zero, and rerun the ablation when the model version changes: patterns that compensated for an older model's weakness often become dead weight.
Failure modes
- Pattern stacking. Every incident adds a rule and nothing is removed, until rules contradict. Fix: ablate on a schedule.
- Examples that leak. The model copies names, lengths or phrasing from few-shot examples into answers. Fix: vary surface features across examples and test on inputs unlike them.
- Constraint without escape. A schema with no abstain value forces invented answers. Fix: a status enum or nullable fields, and evals that include unanswerable inputs.
- Voting on correlated errors. Five samples that share a misunderstanding agree confidently. Agreement measures variance, not correctness.
- Verifier collusion. The same model, prompted similarly, approves its own mistakes. Prefer deterministic checks, then a different prompt or model.
- Unbounded retries. Retry loops without a cap turn one bad input into runaway spend. Cap retries and log every abstention.
Trade-offs
Input patterns spend tokens and are the cheapest place to start, especially when the static prefix is cached. Output patterns spend flexibility, since a schema cannot express an answer it did not anticipate. Control-flow patterns spend calls and latency, and their benefit concentrates on the hard tail, which is why a fallback ladder usually beats running the expensive path for everyone. Composed functions are easier to test than one long prompt but are more things to version, so version the pipeline as a unit with its eval results attached.
Related reading
Deep dives for patterns in this catalogue: prompt chaining, routing and self-consistency voting. For the output side see structured output, for the measurement loop see evaluation harnesses and regression testing, and for how prompting changed with reasoning models see prompt engineering in 2026.
What to do next
- Collect 50 to 200 real inputs with labels for one task and write a scoring function before touching the prompt.
- Run your current prompt, read at least 30 failing traces and classify each failure as input, output or control flow.
- Fix the largest class with the cheapest pattern from its family, then rerun and compare.
- Add a schema with an explicit abstain value to every extraction or classification task.
- Replace any model-based check that could be a deterministic one, such as verbatim quotes, regexes or range checks.
- Run an ablation of the final pipeline and delete patterns that do not move the metric beyond noise.
- Read the deep dives linked in this article for the patterns you kept, and rerun the ablation whenever the model version changes.