In 2023, prompt engineering often meant finding magic phrases: ask the model to think step by step, offer it a tip, shout the important rule in capital letters. By 2026 most of those tricks have faded, because models follow plain instructions much better, many of them reason internally before answering, and the platforms around them offer features that replace whole paragraphs of prompt text. What remains is closer to software engineering: define an interface, control the inputs, validate the outputs, and measure every change.
This article describes that practice. It covers what changed and what did not, how to lay out a request so that it is cheap and stable, how to prompt models that reason on their own, how to turn output formats into contracts, and how to build the evaluation loop that makes all of this safe to change. The individual techniques have their own deep dives, starting with the anatomy of a prompt; this page is the map of how they fit together today.
What changed, and what did not
| Then | Now | Consequence for prompts |
|---|---|---|
| Begging for JSON and parsing hopefully | Schema-constrained structured outputs and typed tool calls | Format rules move out of prose into schemas |
| Spelling out step-by-step reasoning | Reasoning models that think before answering | State goals and success criteria, not procedures |
| Short contexts, every token precious | Very long contexts, priced per token | The problem becomes choosing what to include |
| Each call paid in full | Prefix caching with discounted cached tokens | Stable content first, variable content last |
| Prompts edited in a playground | Prompts versioned, reviewed and evaluated in CI | Every change needs an eval run |
Some things did not change. A model still cannot use information it was not given. Ambiguous instructions still produce inconsistent results. Well-chosen examples still teach format and tone faster than descriptions do, as the few-shot deep dive shows. And untrusted text inside a prompt can still carry instructions, so separating instructions from data with clear delimiters remains a safety requirement, not a style choice.
The shape of a request
Think of a production request as layers ordered by how often they change. The system instructions, tool definitions and examples change only when you release a new prompt version. Retrieved documents and conversation history change per request. The task for this turn changes every time. Putting the stable layers first matters for two reasons. Prefix caches only reuse an identical leading sequence of tokens, so one changing byte near the top invalidates everything after it. And a stable top makes behaviour easier to reason about, because the same rules frame every request.
Contracts over pleading: structured outputs
If code consumes the answer, the answer needs a schema. Most major APIs now accept a JSON Schema and constrain generation to match it, and tool calls give the same guarantee for actions. That removes a class of failures, such as trailing prose and missing braces, and lets you delete sentences like 'respond only with valid JSON'.
A schema guarantees shape, not truth. A valid object can still put a billing ticket in the bug category. So validate in two stages: the schema check, then business rules written in code. Keep enums short and mutually exclusive, give every field a description, and include an explicit escape value such as 'other' so the model is not forced to guess. The function below works with any provider wrapper that returns JSON text.
import json, os
from pydantic import BaseModel, Field, ValidationError
from typing import Literal
class Triage(BaseModel):
category: Literal["billing", "bug", "account", "feature_request", "other"]
urgency: Literal["low", "normal", "high"]
summary: str = Field(max_length=200)
needs_human: bool
def triage(llm, ticket_text, max_attempts=2):
"""llm(messages, schema) -> str is your provider wrapper; it should request strict JSON."""
messages = [{"role": "user", "content": f"<ticket>\n{ticket_text}\n</ticket>\nTriage this ticket."}]
for attempt in range(max_attempts):
raw = llm(messages, schema=Triage.model_json_schema())
try:
result = Triage.model_validate_json(raw)
except ValidationError as err:
messages += [{"role": "assistant", "content": raw},
{"role": "user", "content": f"That output failed validation: {err}. Return corrected JSON only."}]
continue
if result.urgency == "high" and not result.needs_human:
result.needs_human = True # business rule enforced in code, not in prose
return result
raise RuntimeError("triage output invalid after retries")For reference, this is how a strict schema is requested in OpenAI's Chat Completions API. Strict modes support only a subset of JSON Schema, so check the provider's list of supported keywords and keep constraints such as length limits in your own post-validation, where Pydantic already enforces them. Other providers have their own equivalents, compared in model-specific prompt tricks.
# OpenAI Chat Completions form of a strict schema request.
# Strict mode accepts a subset of JSON Schema: every object needs additionalProperties: false
# and every property must be listed in "required".
TRIAGE_SCHEMA = Triage.model_json_schema()
TRIAGE_SCHEMA["additionalProperties"] = False
TRIAGE_SCHEMA["required"] = list(TRIAGE_SCHEMA["properties"])
response_format = {
"type": "json_schema",
"json_schema": {"name": "triage", "strict": True, "schema": TRIAGE_SCHEMA},
}
Prompting models that reason
Reasoning models spend tokens thinking before they answer, and most APIs let you set how much. Three habits from older models now hurt more than they help. Writing out the exact steps to follow constrains a model that would often find a better route. Demanding visible step-by-step reasoning in the answer duplicates work the model already did and adds tokens to parse. Stacking emphatic warnings makes the model over-apply a rule in cases where it should not.
What works instead is a clear brief: the goal, the audience, the constraints, what a good answer looks like, and how it will be checked. If a task is hard, raise the reasoning budget before rewriting the prompt. If it is easy and latency matters, lower the budget or use a non-reasoning model. Examples are still useful for output format, but keep them few and representative, since the model generalises from their content as well as their form. Vendor guidance differs in the details, so read the current documentation for the model you deploy rather than carrying advice across providers.
Layout for caching
Prefix caching lets a provider reuse the computed attention state for a leading sequence of tokens it has seen recently, which lowers cost and time to first token. The rules are mechanical: the cached part must be byte-identical and must come first. Some APIs cache automatically; others ask you to mark the breakpoint. In Anthropic's Messages API you mark it with a cache_control block, as below. Minimum cacheable lengths and cache lifetimes vary by provider and model, so check the current documentation rather than assuming.
import os, anthropic
client = anthropic.Anthropic()
SYSTEM_PROMPT = open("prompts/triage/v7/system.md").read() # long, stable
EXAMPLES = open("prompts/triage/v7/examples.md").read() # fixed per release
def ask(ticket_text):
return client.messages.create(
model=os.environ["TRIAGE_MODEL"], # pin in config, never hard-code in prompts
max_tokens=600,
system=[
{"type": "text", "text": SYSTEM_PROMPT},
{"type": "text", "text": EXAMPLES,
"cache_control": {"type": "ephemeral"}}, # everything up to here is the cached prefix
],
messages=[{"role": "user", "content": f"<ticket>\n{ticket_text}\n</ticket>"}],
)The classic mistake is a timestamp, user name or request ID in the system prompt. It changes every call, so nothing after it is ever cached. Move per-request values into the final user turn. Also keep tool definitions in a fixed order; reordering them changes the prefix. For the serving side of caching see LLM prompt caching.
Context is a budget
A large context window does not mean you should fill it. Every extra token costs money and time, and models use information unevenly across long inputs. Liu and colleagues (2023) showed that retrieval accuracy drops when the relevant passage sits in the middle of a long context, and similar position effects keep appearing in newer evaluations. Treat context as a budget with an owner.
- Include what the task needs and can use: the relevant documents, not the whole knowledge base.
- Put the most important material near the start or end, and the instruction that uses it close to the material.
- Summarise old conversation turns instead of replaying them in full; keep decisions and facts, drop chit-chat.
- Wrap every untrusted input in labelled delimiters and say in the system instructions that text inside them is data, never instructions.
- Log the token count of each layer per request so you can see which layer is growing.
Evals are the development loop
No prompt change is an improvement until a test set says so. Start with 50 to 200 real cases, each with an expected outcome, drawn from production logs and known failures. Run every version against the same cases and compare per case, not just in aggregate, because a change can fix ten cases and break eight others. With small sets the noise is large, so use a paired comparison with a confidence interval before you declare a winner.
import json, random
def run_suite(prompt_version, cases, call):
"""cases: [{"id", "input", "expected": {...}}]; call(version, input) -> parsed result or None."""
rows = []
for case in cases:
got = call(prompt_version, case["input"])
rows.append({"id": case["id"],
"valid": got is not None,
"correct": got is not None and got.category == case["expected"]["category"]})
return rows
def compare(old_rows, new_rows, iters=2000, seed=0):
"""Paired bootstrap on per-case correctness: is the new version really better?"""
diffs = [int(n["correct"]) - int(o["correct"]) for o, n in zip(old_rows, new_rows)]
rng = random.Random(seed)
means = sorted(sum(rng.choice(diffs) for _ in diffs) / len(diffs) for _ in range(iters))
return {"delta": sum(diffs) / len(diffs),
"ci95": (means[int(0.025 * iters)], means[int(0.975 * iters)]),
"regressed_ids": [o["id"] for o, n in zip(old_rows, new_rows) if o["correct"] and not n["correct"]]}The regressed_ids list matters as much as the delta. Read those cases before shipping. For how to grow this into a full system with scorers, online signals and release gates, see LLM evaluation architecture.
Prompts as versioned artifacts
- Store prompts as files in the repository, one directory per prompt with a version, next to the code that renders them.
- Pin the model ID in configuration together with the prompt version; a model upgrade is a prompt change and must pass the same evals.
- Review prompt diffs like code diffs, with the eval report attached.
- Roll out gradually: shadow traffic or a small percentage first, compare online metrics, then promote.
- Log the prompt version and model ID with every request so any output can be traced to exactly what produced it.
Worked example: a ticket triage prompt
A team classifies support tickets into five categories and three urgency levels. Version 1 is a paragraph of instructions ending in 'answer in JSON'. It parses most of the time but fails on long tickets, and the 'other' category is overused. The numbers below are illustrative of the process, not measurements from a real system.
- Collect 150 labelled tickets from logs, including every ticket a human corrected last month.
- Version 2 moves the format into a strict schema and enforces the high-urgency rule in code. Parse failures go to zero; accuracy is unchanged, as expected, because shape was the only thing fixed.
- Version 3 adds one-line definitions for each category with a boundary case for each, and four examples chosen from the confusion matrix. The paired comparison shows a gain whose interval excludes zero, and regressed_ids lists three tickets, which turn out to be mislabelled in the test set; fix the labels, not the prompt.
- Version 4 moves the ticket timestamp out of the system prompt and marks a cache breakpoint after the examples. Accuracy is identical; median latency and cost fall because the prefix is now reused.
- Ship version 4 to 10 percent of traffic, compare the human-correction rate for a week, then promote.
When a later version misbehaves on a specific ticket, follow the capture, minimise and fix routine in the prompt debugging workflow.
Failure modes and trade-offs
| Failure | Cause | Fix |
|---|---|---|
| Valid JSON, wrong answer | Schema treated as correctness | Business-rule validation and accuracy evals |
| Cost jumps after a small edit | Variable text moved into the cached prefix | Keep per-request values in the last turn; watch cache-hit metrics |
| Quality drops after a model upgrade | Prompt tuned to old model quirks | Run the suite on the new model before switching |
| Instructions followed from a document | Untrusted text not separated from instructions | Delimiters, explicit data-only rule, tool permissions limited |
| Eval says better, users say worse | Test set no longer matches traffic | Refresh cases from recent logs every release |
The main trade-off is effort against risk. A one-off internal script does not need a 200-case suite. A prompt that touches customers, money or actions does, and the cost of building the suite is repaid the first time it catches a regression before release.
What to do next
- Move every prompt you run in production into versioned files and log the version and model with each request.
- Replace prose format instructions with a strict schema and add code-level business-rule checks.
- Reorder each request so that stable layers come first, then measure your cache-hit rate.
- Build a 50-case eval set from real traffic and run it on every prompt or model change.
- Rewrite step-by-step scripts for reasoning models as goals, constraints and success criteria, and compare.
- Wrap all untrusted inputs in labelled delimiters and state that their contents are data.