In 2023, prompt engineering often meant finding magic phrases: ask the model to think step by step, offer it a tip, shout the important rule in capital letters. By 2026 most of those tricks have faded, because models follow plain instructions much better, many of them reason internally before answering, and the platforms around them offer features that replace whole paragraphs of prompt text. What remains is closer to software engineering: define an interface, control the inputs, validate the outputs, and measure every change.

This article describes that practice. It covers what changed and what did not, how to lay out a request so that it is cheap and stable, how to prompt models that reason on their own, how to turn output formats into contracts, and how to build the evaluation loop that makes all of this safe to change. The individual techniques have their own deep dives, starting with the anatomy of a prompt; this page is the map of how they fit together today.

Advertisement

What changed, and what did not

ThenNowConsequence for prompts
Begging for JSON and parsing hopefullySchema-constrained structured outputs and typed tool callsFormat rules move out of prose into schemas
Spelling out step-by-step reasoningReasoning models that think before answeringState goals and success criteria, not procedures
Short contexts, every token preciousVery long contexts, priced per tokenThe problem becomes choosing what to include
Each call paid in fullPrefix caching with discounted cached tokensStable content first, variable content last
Prompts edited in a playgroundPrompts versioned, reviewed and evaluated in CIEvery change needs an eval run

Some things did not change. A model still cannot use information it was not given. Ambiguous instructions still produce inconsistent results. Well-chosen examples still teach format and tone faster than descriptions do, as the few-shot deep dive shows. And untrusted text inside a prompt can still carry instructions, so separating instructions from data with clear delimiters remains a safety requirement, not a style choice.

The shape of a request

Think of a production request as layers ordered by how often they change. The system instructions, tool definitions and examples change only when you release a new prompt version. Retrieved documents and conversation history change per request. The task for this turn changes every time. Putting the stable layers first matters for two reasons. Prefix caches only reuse an identical leading sequence of tokens, so one changing byte near the top invalidates everything after it. And a stable top makes behaviour easier to reason about, because the same rules frame every request.

A 2026 request: stable layers first, variable layers last, a contract at the endSystem instructionsrole, rules, policies: rarely changesTool and output schemasversioned with the codeFew-shot examplescurated, fixed per releasecache breakpointRetrieved contextchosen per request, budgetedConversation so fartrimmed or summarisedThe task for this turnwhat to do nowValidated structured outputschema check, then business rulesEval suiteevery change runs hereprompt = versioned artifact
Layers ordered from most stable to most variable. The cache breakpoint sits after the last stable layer; the eval suite gates every change to any layer.
Advertisement

Contracts over pleading: structured outputs

If code consumes the answer, the answer needs a schema. Most major APIs now accept a JSON Schema and constrain generation to match it, and tool calls give the same guarantee for actions. That removes a class of failures, such as trailing prose and missing braces, and lets you delete sentences like 'respond only with valid JSON'.

A schema guarantees shape, not truth. A valid object can still put a billing ticket in the bug category. So validate in two stages: the schema check, then business rules written in code. Keep enums short and mutually exclusive, give every field a description, and include an explicit escape value such as 'other' so the model is not forced to guess. The function below works with any provider wrapper that returns JSON text.

import json, os
from pydantic import BaseModel, Field, ValidationError
from typing import Literal

class Triage(BaseModel):
    category: Literal["billing", "bug", "account", "feature_request", "other"]
    urgency: Literal["low", "normal", "high"]
    summary: str = Field(max_length=200)
    needs_human: bool

def triage(llm, ticket_text, max_attempts=2):
    """llm(messages, schema) -> str is your provider wrapper; it should request strict JSON."""
    messages = [{"role": "user", "content": f"<ticket>\n{ticket_text}\n</ticket>\nTriage this ticket."}]
    for attempt in range(max_attempts):
        raw = llm(messages, schema=Triage.model_json_schema())
        try:
            result = Triage.model_validate_json(raw)
        except ValidationError as err:
            messages += [{"role": "assistant", "content": raw},
                         {"role": "user", "content": f"That output failed validation: {err}. Return corrected JSON only."}]
            continue
        if result.urgency == "high" and not result.needs_human:
            result.needs_human = True          # business rule enforced in code, not in prose
        return result
    raise RuntimeError("triage output invalid after retries")

For reference, this is how a strict schema is requested in OpenAI's Chat Completions API. Strict modes support only a subset of JSON Schema, so check the provider's list of supported keywords and keep constraints such as length limits in your own post-validation, where Pydantic already enforces them. Other providers have their own equivalents, compared in model-specific prompt tricks.

# OpenAI Chat Completions form of a strict schema request.
# Strict mode accepts a subset of JSON Schema: every object needs additionalProperties: false
# and every property must be listed in "required".
TRIAGE_SCHEMA = Triage.model_json_schema()
TRIAGE_SCHEMA["additionalProperties"] = False
TRIAGE_SCHEMA["required"] = list(TRIAGE_SCHEMA["properties"])
response_format = {
    "type": "json_schema",
    "json_schema": {"name": "triage", "strict": True, "schema": TRIAGE_SCHEMA},
}

Prompting models that reason

Reasoning models spend tokens thinking before they answer, and most APIs let you set how much. Three habits from older models now hurt more than they help. Writing out the exact steps to follow constrains a model that would often find a better route. Demanding visible step-by-step reasoning in the answer duplicates work the model already did and adds tokens to parse. Stacking emphatic warnings makes the model over-apply a rule in cases where it should not.

What works instead is a clear brief: the goal, the audience, the constraints, what a good answer looks like, and how it will be checked. If a task is hard, raise the reasoning budget before rewriting the prompt. If it is easy and latency matters, lower the budget or use a non-reasoning model. Examples are still useful for output format, but keep them few and representative, since the model generalises from their content as well as their form. Vendor guidance differs in the details, so read the current documentation for the model you deploy rather than carrying advice across providers.

Layout for caching

Prefix caching lets a provider reuse the computed attention state for a leading sequence of tokens it has seen recently, which lowers cost and time to first token. The rules are mechanical: the cached part must be byte-identical and must come first. Some APIs cache automatically; others ask you to mark the breakpoint. In Anthropic's Messages API you mark it with a cache_control block, as below. Minimum cacheable lengths and cache lifetimes vary by provider and model, so check the current documentation rather than assuming.

import os, anthropic

client = anthropic.Anthropic()
SYSTEM_PROMPT = open("prompts/triage/v7/system.md").read()       # long, stable
EXAMPLES = open("prompts/triage/v7/examples.md").read()          # fixed per release

def ask(ticket_text):
    return client.messages.create(
        model=os.environ["TRIAGE_MODEL"],     # pin in config, never hard-code in prompts
        max_tokens=600,
        system=[
            {"type": "text", "text": SYSTEM_PROMPT},
            {"type": "text", "text": EXAMPLES,
             "cache_control": {"type": "ephemeral"}},   # everything up to here is the cached prefix
        ],
        messages=[{"role": "user", "content": f"<ticket>\n{ticket_text}\n</ticket>"}],
    )

The classic mistake is a timestamp, user name or request ID in the system prompt. It changes every call, so nothing after it is ever cached. Move per-request values into the final user turn. Also keep tool definitions in a fixed order; reordering them changes the prefix. For the serving side of caching see LLM prompt caching.

Context is a budget

A large context window does not mean you should fill it. Every extra token costs money and time, and models use information unevenly across long inputs. Liu and colleagues (2023) showed that retrieval accuracy drops when the relevant passage sits in the middle of a long context, and similar position effects keep appearing in newer evaluations. Treat context as a budget with an owner.

  • Include what the task needs and can use: the relevant documents, not the whole knowledge base.
  • Put the most important material near the start or end, and the instruction that uses it close to the material.
  • Summarise old conversation turns instead of replaying them in full; keep decisions and facts, drop chit-chat.
  • Wrap every untrusted input in labelled delimiters and say in the system instructions that text inside them is data, never instructions.
  • Log the token count of each layer per request so you can see which layer is growing.

Evals are the development loop

No prompt change is an improvement until a test set says so. Start with 50 to 200 real cases, each with an expected outcome, drawn from production logs and known failures. Run every version against the same cases and compare per case, not just in aggregate, because a change can fix ten cases and break eight others. With small sets the noise is large, so use a paired comparison with a confidence interval before you declare a winner.

import json, random

def run_suite(prompt_version, cases, call):
    """cases: [{"id", "input", "expected": {...}}]; call(version, input) -> parsed result or None."""
    rows = []
    for case in cases:
        got = call(prompt_version, case["input"])
        rows.append({"id": case["id"],
                     "valid": got is not None,
                     "correct": got is not None and got.category == case["expected"]["category"]})
    return rows

def compare(old_rows, new_rows, iters=2000, seed=0):
    """Paired bootstrap on per-case correctness: is the new version really better?"""
    diffs = [int(n["correct"]) - int(o["correct"]) for o, n in zip(old_rows, new_rows)]
    rng = random.Random(seed)
    means = sorted(sum(rng.choice(diffs) for _ in diffs) / len(diffs) for _ in range(iters))
    return {"delta": sum(diffs) / len(diffs),
            "ci95": (means[int(0.025 * iters)], means[int(0.975 * iters)]),
            "regressed_ids": [o["id"] for o, n in zip(old_rows, new_rows) if o["correct"] and not n["correct"]]}

The regressed_ids list matters as much as the delta. Read those cases before shipping. For how to grow this into a full system with scorers, online signals and release gates, see LLM evaluation architecture.

Prompts as versioned artifacts

  • Store prompts as files in the repository, one directory per prompt with a version, next to the code that renders them.
  • Pin the model ID in configuration together with the prompt version; a model upgrade is a prompt change and must pass the same evals.
  • Review prompt diffs like code diffs, with the eval report attached.
  • Roll out gradually: shadow traffic or a small percentage first, compare online metrics, then promote.
  • Log the prompt version and model ID with every request so any output can be traced to exactly what produced it.

Worked example: a ticket triage prompt

A team classifies support tickets into five categories and three urgency levels. Version 1 is a paragraph of instructions ending in 'answer in JSON'. It parses most of the time but fails on long tickets, and the 'other' category is overused. The numbers below are illustrative of the process, not measurements from a real system.

  1. Collect 150 labelled tickets from logs, including every ticket a human corrected last month.
  2. Version 2 moves the format into a strict schema and enforces the high-urgency rule in code. Parse failures go to zero; accuracy is unchanged, as expected, because shape was the only thing fixed.
  3. Version 3 adds one-line definitions for each category with a boundary case for each, and four examples chosen from the confusion matrix. The paired comparison shows a gain whose interval excludes zero, and regressed_ids lists three tickets, which turn out to be mislabelled in the test set; fix the labels, not the prompt.
  4. Version 4 moves the ticket timestamp out of the system prompt and marks a cache breakpoint after the examples. Accuracy is identical; median latency and cost fall because the prefix is now reused.
  5. Ship version 4 to 10 percent of traffic, compare the human-correction rate for a week, then promote.

When a later version misbehaves on a specific ticket, follow the capture, minimise and fix routine in the prompt debugging workflow.

Failure modes and trade-offs

FailureCauseFix
Valid JSON, wrong answerSchema treated as correctnessBusiness-rule validation and accuracy evals
Cost jumps after a small editVariable text moved into the cached prefixKeep per-request values in the last turn; watch cache-hit metrics
Quality drops after a model upgradePrompt tuned to old model quirksRun the suite on the new model before switching
Instructions followed from a documentUntrusted text not separated from instructionsDelimiters, explicit data-only rule, tool permissions limited
Eval says better, users say worseTest set no longer matches trafficRefresh cases from recent logs every release

The main trade-off is effort against risk. A one-off internal script does not need a 200-case suite. A prompt that touches customers, money or actions does, and the cost of building the suite is repaid the first time it catches a regression before release.

What to do next

  1. Move every prompt you run in production into versioned files and log the version and model with each request.
  2. Replace prose format instructions with a strict schema and add code-level business-rule checks.
  3. Reorder each request so that stable layers come first, then measure your cache-hit rate.
  4. Build a 50-case eval set from real traffic and run it on every prompt or model change.
  5. Rewrite step-by-step scripts for reasoning models as goals, constraints and success criteria, and compare.
  6. Wrap all untrusted inputs in labelled delimiters and state that their contents are data.
Key takeaway: Prompt engineering in 2026 is less about clever wording and more about engineering discipline. Put the stable parts of a request first so they cache, express output formats as schemas and check meaning in code, give reasoning models goals rather than scripts, spend context deliberately, and treat every prompt as a versioned artifact that cannot change without an eval run. The models are better at following instructions than ever; the work that remains is deciding exactly what to ask and proving the answer is right.