Every team that ships on more than one large language model collects folklore: wrap things in XML for Claude, tell GPT to think step by step, set temperature to zero for determinism. Some of it was true for one model generation. Some of it is now harmful, and a few tricks no longer even pass API validation. Prompt advice is specific to a model family and a date, and it ages faster than most code.

This article separates what is genuinely model-specific from what is good prompting everywhere, using the vendors' own published guidance as checked on 1 October 2026: Anthropic's prompting best-practices page, OpenAI's reasoning best-practices guide, and Google's Gemini prompting strategies. It then shows how to keep that knowledge in one place in code, how to migrate a prompt when a trick stops working, and how to measure whether a trick still earns its place.

Advertisement

Three layers of model specificity

"Model-specific" covers three different things, and they age at different speeds.

  1. API surface. Message roles, parameters and features: whether there is a system or developer role, whether you can prefill the assistant turn, how reasoning depth is controlled, how a JSON schema is enforced. These are hard facts that either work or return an error.
  2. Trained conventions. Formats a model family was trained to read well: XML-tagged sections, where long documents go, whether examples help. These are soft: a different format still works, just less reliably.
  3. Behavioural calibration. How strongly a model reacts to emphasis, how verbose it is by default, how eagerly it calls tools. These shift between versions of the same family, which is why vendors publish per-model notes.

Everything else is common ground and should be in every prompt regardless of model: state the task plainly, say what the output is for and who reads it, give the information the model needs, define the output format, and test on real cases. The model-specific parts are adjustments on top, not a substitute.

Claude: what Anthropic's guidance says

Anthropic's best-practices page recommends XML tags to separate the parts of a prompt that mixes instructions, context, examples and inputs, for example <instructions>, <context> and <input>, with multiple documents wrapped in <document> tags holding <document_content> and <source>. For long inputs of 20,000 tokens or more it says to put the documents at the top and the query at the end, noting that queries at the end improved response quality by up to 30 percent in their tests on complex, multi-document inputs.

Three points change prompts written for older Claude models. First, prefilling the last assistant turn is no longer supported starting with the Claude 4.6 models: such requests return a 400 error. The guide suggests structured outputs or plain instructions for format control, direct instructions to skip preambles, and moving injected reminders into the user turn. Second, explicit thinking budgets are going away: budget_tokens is deprecated on the 4.6 models and returns a 400 error from Claude 4.7 onward, replaced by adaptive thinking (thinking: {type: "adaptive"}) with depth guided by an effort setting such as "high" in output_config. Third, the guide notes that Claude Opus 4.5 and 4.6 follow the system prompt more closely than earlier models, so emphatic language written to overcome under-triggering, such as "CRITICAL: You MUST use this tool", can now cause over-triggering; the guide recommends ordinary wording such as "Use this tool when...".

The page also recommends giving the reason behind an instruction, because the model generalises from the explanation, and telling the model what to do rather than what not to do, for example asking for flowing prose paragraphs instead of saying "do not use markdown".

Advertisement

OpenAI reasoning models: what OpenAI&amp;amp;amp;#x27;s guidance says

OpenAI's reasoning best-practices guide applies to its reasoning models, not to every model it serves, and that scope matters. It says these models take developer messages rather than system messages, a change made to fit OpenAI's chain-of-command rules. It says that because the models reason internally, asking them to think step by step or explain their reasoning is unnecessary. It recommends delimiters such as markdown, XML tags and section titles to mark the parts of the input, and it recommends trying prompts without examples first, because reasoning models often do not need few-shot examples.

One detail catches teams out: reasoning models in the API avoid markdown formatting by default, and the guide says to put the string Formatting re-enabled on the first line of the developer message when markdown output is wanted. For schema-constrained output, Structured Outputs takes a JSON schema through response_format with strict set to true, which constrains decoding to the schema rather than merely asking for it.

Gemini: what Google&amp;amp;amp;#x27;s guidance says

Google's prompting strategies page takes the opposite default on examples: it recommends always including few-shot examples, because they shape format, phrasing and patterns more reliably than instructions alone. Its Gemini 3 template uses tagged sections such as <role>, <constraints> and <output_format> in the system instruction. For long context it says to supply all the context first and put the specific question at the very end, bridged with a phrase such as "Based on the information above".

The sampling advice reverses an old habit. For Gemini 3.x models Google strongly recommends leaving temperature and related parameters at their defaults, warning that setting temperature below 1.0 can cause looping or degraded performance. A team that set temperature to 0.2 everywhere "for determinism" is following advice that this model family explicitly contradicts. Gemini models also think internally, and complex JSON is best requested through structured output.

The differences side by side

DimensionClaude (Anthropic)OpenAI reasoning modelsGemini 3.x (Google)
Instruction slotSystem promptDeveloper messageSystem instruction
StructureXML tags recommendedMarkdown, XML or section titlesTagged sections in the template
ExamplesUseful; wrap in tagsTry zero-shot firstAlways include few-shot
ReasoningAdaptive thinking guided by effortInternal; no step-by-step promptInternal; no explicit steps needed
Format controlInstructions, structured outputs or tools; no prefill on 4.6+Structured Outputs with a strict schemaStructured output with a response schema
Long contextDocuments first, query lastNot covered by the reasoning guideContext first, question last
SamplingNot covered hereNot covered hereKeep defaults; below 1.0 can loop

Two of these rows conflict across vendors, and those are the ones worth designing around: examples (always for Gemini, zero-shot first for OpenAI reasoning models) and anything that touches API validation (prefill and thinking budgets on Claude). Most other rows agree: separate the parts of the prompt clearly, put long material before the question, and do not script the model's reasoning.

Keep the tricks in renderers, not in prompts

The maintainable pattern is to write each prompt once as a neutral specification and render it per provider. The specification holds what you want: instructions with their reasons, documents, optional examples, the question and an optional schema. Each renderer applies only its vendor's documented advice. When guidance changes, you edit one function and rerun the evals.

One prompt spec, three renderers, one eval gatePromptSpecinstructions, documents,examples, question, schemarender_claudeXML tags, docs first, no prefillrender_openai_reasoningdeveloper role, zero-shot firstrender_geminifew-shot, default temperatureEval matrixsame cases, every modelscore per rendererModel-specific advice lives only in the renderers. The spec stays neutral; the eval decides which advice still helps.
A neutral PromptSpec is rendered per provider; vendor advice is confined to the renderers, and one eval matrix decides which advice still helps.
from dataclasses import dataclass, field

@dataclass
class PromptSpec:
    instructions: str                     # what to do, and why
    documents: list[str] = field(default_factory=list)
    examples: list[tuple[str, str]] = field(default_factory=list)
    question: str = ""
    json_schema: dict | None = None


def _docs_xml(docs):
    return "\n".join(
        f'<document index="{i}"><document_content>{d}</document_content></document>'
        for i, d in enumerate(docs, 1))


def render_claude(spec, model):
    # Anthropic: XML-tagged sections, long material first, question last, no prefill.
    user = f"<documents>\n{_docs_xml(spec.documents)}\n</documents>\n\n"
    if spec.examples:
        shots = "\n".join(f"<example><input>{q}</input><output>{a}</output></example>"
                          for q, a in spec.examples)
        user += f"<examples>\n{shots}\n</examples>\n\n"
    user += f"<question>{spec.question}</question>"
    return {"model": model, "max_tokens": 4000,
            "system": spec.instructions,
            "messages": [{"role": "user", "content": user}]}


def render_openai_reasoning(spec, model):
    # OpenAI reasoning models: developer message, no step-by-step instruction,
    # zero-shot first (examples are added only if the eval shows they help).
    user = _docs_xml(spec.documents) + "\n\n" + spec.question
    req = {"model": model,
           "messages": [{"role": "developer", "content": spec.instructions},
                        {"role": "user", "content": user}]}
    if spec.json_schema:
        req["response_format"] = {"type": "json_schema", "json_schema": {
            "name": "answer", "schema": spec.json_schema, "strict": True}}
    return req


def render_gemini(spec):
    # Google: always include few-shot examples, context first, question last,
    # sampling parameters left at their defaults for Gemini 3.x.
    from google.genai import types
    shots = "\n\n".join(f"Input: {q}\nOutput: {a}" for q, a in spec.examples)
    contents = ("\n\n".join(spec.documents) + "\n\n" + shots
                + "\n\nBased on the information above, " + spec.question)
    config = types.GenerateContentConfig(
        system_instruction=spec.instructions,
        response_mime_type="application/json" if spec.json_schema else None,
        response_schema=spec.json_schema)
    return contents, config

The renderers return request shapes rather than calling SDKs, so they can be unit-tested and diffed. Note what each one leaves out as much as what it adds: the Claude renderer never adds an assistant message; the OpenAI renderer drops the examples and never says "think step by step"; the Gemini renderer sets no temperature. The delimiters guide covers tag design, and structured output covers schema enforcement in depth.

Worked example: migrating a classifier

A support-ticket classifier was written two years ago against an earlier Claude model. It used three tricks: an assistant prefill of {"category": " to force JSON, a system prompt reading "CRITICAL: you MUST output only JSON", and few-shot examples. Moving it to a current Claude model fails immediately with a 400 error because of the prefill. The migration:

# Before: steering the format with an assistant prefill.
# Worked on earlier Claude models; returns HTTP 400 on Claude 4.6 and later.
messages = [
    {"role": "user", "content": ticket_prompt},
    {"role": "assistant", "content": '{"category": "'},
]

# After: the format lives in the instructions, with the reason, and is checked in code.
system = (
    "Classify the support ticket. Reply with only a JSON object: "
    '{"category": "billing" | "outage" | "account" | "other", "reason": "<one sentence>"}. '
    "A program parses your reply, so do not write anything before or after the object."
)
messages = [{"role": "user", "content": ticket_prompt}]

def parse_or_retry(call, request, retries=1):
    for attempt in range(retries + 1):
        text = call(request)
        try:
            return json.loads(text)
        except json.JSONDecodeError:
            if attempt == retries:
                raise

The prefill is replaced by an instruction that states the format and the reason for it, and a parser that retries once and then fails loudly. The capitalised warning becomes a plain sentence. The examples stay, wrapped in <example> tags. On an eval set of 300 labelled tickets, the team compares the old and new versions before switching traffic.

The same team then adds an OpenAI reasoning model as a fallback and a Gemini 3 model for a cheaper tier. The neutral spec does not change. The OpenAI renderer starts zero-shot and adds a Structured Outputs schema with the four categories as an enum; the Gemini renderer keeps the examples and removes the inherited temperature=0.2.

Measure every trick with an eval matrix

A trick is a hypothesis about one model. The only way to know whether it still helps is to run the same cases through each renderer, with and without the trick, and compare. Keep the harness small enough that people run it on every prompt change:

from dataclasses import replace

def run_matrix(cases, renderers, call, score):
    """cases: [(PromptSpec, expected)]; renderers: {name: spec -> request}, e.g.
    functools.partial(render_claude, model=MODEL); call(name, request) -> text;
    score(text, expected) -> 0 or 1."""
    table = {}
    for name, render in renderers.items():
        passed = sum(score(call(name, render(spec)), expected) for spec, expected in cases)
        table[name] = passed / len(cases)
    return table

# Ablate one trick at a time: same cases, renderer with and without the trick.
baseline = run_matrix(cases, {"gemini": render_gemini}, call, exact_label)
no_shots = run_matrix([(replace(s, examples=[]), e) for s, e in cases],
                      {"gemini": render_gemini}, call, exact_label)
print(baseline, no_shots)   # keep the examples only if they earn their tokens

Score with exact checks where possible: label match, schema validity, presence of a required citation. Track cost beside quality: examples cost tokens on every call, which prompt caching reduces for a stable prefix. Rerun the matrix whenever a provider ships a new model version, not only when you change the prompt. The evals guide covers building the case set.

Failure modes

  • Folklore that now fails validation. Prefill on Claude 4.6 and later, budget_tokens on Claude 4.7 and later. These fail loudly, which is the good case; search your code base for them before upgrading.
  • Emphasis that now overshoots. Capitalised MUST and CRITICAL written for less attentive models make newer ones over-apply a rule, such as calling a tool on every turn.
  • Scripted reasoning on reasoning models. "Think step by step" is, per OpenAI's guide, unnecessary, and it spends tokens for no measured gain.
  • Old sampling settings. Low temperature carried over to Gemini 3.x can cause looping.
  • Missing formatting. Markdown silently disappears from OpenAI reasoning-model output unless the developer message re-enables it.
  • Single-model evals. A change validated on one model is shipped to all. Run the matrix across every renderer in production.

Trade-offs

Renderers add a layer and a test burden; for a single-model product with one prompt, a well-written prompt plus an eval set may be enough. Provider-specific features such as strict schemas or adaptive thinking are worth using even though they reduce portability, because they move guarantees from wording into the API; isolate them in the renderer so switching providers stays a local change. Prefer advice that holds across vendors over one model's quirks; the general advice survives upgrades.

What to do next

  1. Grep your code for assistant-turn prefills, budget_tokens, "step by step" and hard-coded temperatures, and list which model each call targets.
  2. Re-read each vendor's current prompting guide for the models you use and note the date you read it.
  3. Refactor your most important prompt into a neutral spec with one renderer per provider.
  4. Replace emphatic capitals with plain instructions that include the reason.
  5. Build a 100 to 300 case eval set from real traffic and run it through every renderer.
  6. Ablate one trick at a time, such as examples, tags or document order, and keep only those that measurably help.
  7. Rerun the matrix on every provider model release, and read role prompting next for the same evidence-first treatment of personas.
Key takeaway: Model-specific prompting is three things: API facts that either work or return errors, trained conventions that improve reliability, and behaviour that shifts between versions. As of October 2026 the vendor guides agree on clear sections, long context before the question and not scripting reasoning, and disagree on examples: Gemini says always include them, OpenAI's reasoning guide says try zero-shot first. Claude prompts must drop prefill and thinking budgets on newer models, and Gemini 3.x should keep default temperature. Write each prompt once as a neutral spec, put vendor advice in renderers, and let an eval matrix decide which tricks stay.