Every team that ships on more than one large language model collects folklore: wrap things in XML for Claude, tell GPT to think step by step, set temperature to zero for determinism. Some of it was true for one model generation. Some of it is now harmful, and a few tricks no longer even pass API validation. Prompt advice is specific to a model family and a date, and it ages faster than most code.
This article separates what is genuinely model-specific from what is good prompting everywhere, using the vendors' own published guidance as checked on 1 October 2026: Anthropic's prompting best-practices page, OpenAI's reasoning best-practices guide, and Google's Gemini prompting strategies. It then shows how to keep that knowledge in one place in code, how to migrate a prompt when a trick stops working, and how to measure whether a trick still earns its place.
Three layers of model specificity
"Model-specific" covers three different things, and they age at different speeds.
- API surface. Message roles, parameters and features: whether there is a system or developer role, whether you can prefill the assistant turn, how reasoning depth is controlled, how a JSON schema is enforced. These are hard facts that either work or return an error.
- Trained conventions. Formats a model family was trained to read well: XML-tagged sections, where long documents go, whether examples help. These are soft: a different format still works, just less reliably.
- Behavioural calibration. How strongly a model reacts to emphasis, how verbose it is by default, how eagerly it calls tools. These shift between versions of the same family, which is why vendors publish per-model notes.
Everything else is common ground and should be in every prompt regardless of model: state the task plainly, say what the output is for and who reads it, give the information the model needs, define the output format, and test on real cases. The model-specific parts are adjustments on top, not a substitute.
Claude: what Anthropic's guidance says
Anthropic's best-practices page recommends XML tags to separate the parts of a prompt that mixes instructions, context, examples and inputs, for example <instructions>, <context> and <input>, with multiple documents wrapped in <document> tags holding <document_content> and <source>. For long inputs of 20,000 tokens or more it says to put the documents at the top and the query at the end, noting that queries at the end improved response quality by up to 30 percent in their tests on complex, multi-document inputs.
Three points change prompts written for older Claude models. First, prefilling the last assistant turn is no longer supported starting with the Claude 4.6 models: such requests return a 400 error. The guide suggests structured outputs or plain instructions for format control, direct instructions to skip preambles, and moving injected reminders into the user turn. Second, explicit thinking budgets are going away: budget_tokens is deprecated on the 4.6 models and returns a 400 error from Claude 4.7 onward, replaced by adaptive thinking (thinking: {type: "adaptive"}) with depth guided by an effort setting such as "high" in output_config. Third, the guide notes that Claude Opus 4.5 and 4.6 follow the system prompt more closely than earlier models, so emphatic language written to overcome under-triggering, such as "CRITICAL: You MUST use this tool", can now cause over-triggering; the guide recommends ordinary wording such as "Use this tool when...".
The page also recommends giving the reason behind an instruction, because the model generalises from the explanation, and telling the model what to do rather than what not to do, for example asking for flowing prose paragraphs instead of saying "do not use markdown".
OpenAI reasoning models: what OpenAI&amp;amp;#x27;s guidance says
OpenAI's reasoning best-practices guide applies to its reasoning models, not to every model it serves, and that scope matters. It says these models take developer messages rather than system messages, a change made to fit OpenAI's chain-of-command rules. It says that because the models reason internally, asking them to think step by step or explain their reasoning is unnecessary. It recommends delimiters such as markdown, XML tags and section titles to mark the parts of the input, and it recommends trying prompts without examples first, because reasoning models often do not need few-shot examples.
One detail catches teams out: reasoning models in the API avoid markdown formatting by default, and the guide says to put the string Formatting re-enabled on the first line of the developer message when markdown output is wanted. For schema-constrained output, Structured Outputs takes a JSON schema through response_format with strict set to true, which constrains decoding to the schema rather than merely asking for it.
Gemini: what Google&amp;amp;#x27;s guidance says
Google's prompting strategies page takes the opposite default on examples: it recommends always including few-shot examples, because they shape format, phrasing and patterns more reliably than instructions alone. Its Gemini 3 template uses tagged sections such as <role>, <constraints> and <output_format> in the system instruction. For long context it says to supply all the context first and put the specific question at the very end, bridged with a phrase such as "Based on the information above".
The sampling advice reverses an old habit. For Gemini 3.x models Google strongly recommends leaving temperature and related parameters at their defaults, warning that setting temperature below 1.0 can cause looping or degraded performance. A team that set temperature to 0.2 everywhere "for determinism" is following advice that this model family explicitly contradicts. Gemini models also think internally, and complex JSON is best requested through structured output.
The differences side by side
| Dimension | Claude (Anthropic) | OpenAI reasoning models | Gemini 3.x (Google) |
|---|---|---|---|
| Instruction slot | System prompt | Developer message | System instruction |
| Structure | XML tags recommended | Markdown, XML or section titles | Tagged sections in the template |
| Examples | Useful; wrap in tags | Try zero-shot first | Always include few-shot |
| Reasoning | Adaptive thinking guided by effort | Internal; no step-by-step prompt | Internal; no explicit steps needed |
| Format control | Instructions, structured outputs or tools; no prefill on 4.6+ | Structured Outputs with a strict schema | Structured output with a response schema |
| Long context | Documents first, query last | Not covered by the reasoning guide | Context first, question last |
| Sampling | Not covered here | Not covered here | Keep defaults; below 1.0 can loop |
Two of these rows conflict across vendors, and those are the ones worth designing around: examples (always for Gemini, zero-shot first for OpenAI reasoning models) and anything that touches API validation (prefill and thinking budgets on Claude). Most other rows agree: separate the parts of the prompt clearly, put long material before the question, and do not script the model's reasoning.
Keep the tricks in renderers, not in prompts
The maintainable pattern is to write each prompt once as a neutral specification and render it per provider. The specification holds what you want: instructions with their reasons, documents, optional examples, the question and an optional schema. Each renderer applies only its vendor's documented advice. When guidance changes, you edit one function and rerun the evals.
from dataclasses import dataclass, field
@dataclass
class PromptSpec:
instructions: str # what to do, and why
documents: list[str] = field(default_factory=list)
examples: list[tuple[str, str]] = field(default_factory=list)
question: str = ""
json_schema: dict | None = None
def _docs_xml(docs):
return "\n".join(
f'<document index="{i}"><document_content>{d}</document_content></document>'
for i, d in enumerate(docs, 1))
def render_claude(spec, model):
# Anthropic: XML-tagged sections, long material first, question last, no prefill.
user = f"<documents>\n{_docs_xml(spec.documents)}\n</documents>\n\n"
if spec.examples:
shots = "\n".join(f"<example><input>{q}</input><output>{a}</output></example>"
for q, a in spec.examples)
user += f"<examples>\n{shots}\n</examples>\n\n"
user += f"<question>{spec.question}</question>"
return {"model": model, "max_tokens": 4000,
"system": spec.instructions,
"messages": [{"role": "user", "content": user}]}
def render_openai_reasoning(spec, model):
# OpenAI reasoning models: developer message, no step-by-step instruction,
# zero-shot first (examples are added only if the eval shows they help).
user = _docs_xml(spec.documents) + "\n\n" + spec.question
req = {"model": model,
"messages": [{"role": "developer", "content": spec.instructions},
{"role": "user", "content": user}]}
if spec.json_schema:
req["response_format"] = {"type": "json_schema", "json_schema": {
"name": "answer", "schema": spec.json_schema, "strict": True}}
return req
def render_gemini(spec):
# Google: always include few-shot examples, context first, question last,
# sampling parameters left at their defaults for Gemini 3.x.
from google.genai import types
shots = "\n\n".join(f"Input: {q}\nOutput: {a}" for q, a in spec.examples)
contents = ("\n\n".join(spec.documents) + "\n\n" + shots
+ "\n\nBased on the information above, " + spec.question)
config = types.GenerateContentConfig(
system_instruction=spec.instructions,
response_mime_type="application/json" if spec.json_schema else None,
response_schema=spec.json_schema)
return contents, configThe renderers return request shapes rather than calling SDKs, so they can be unit-tested and diffed. Note what each one leaves out as much as what it adds: the Claude renderer never adds an assistant message; the OpenAI renderer drops the examples and never says "think step by step"; the Gemini renderer sets no temperature. The delimiters guide covers tag design, and structured output covers schema enforcement in depth.
Worked example: migrating a classifier
A support-ticket classifier was written two years ago against an earlier Claude model. It used three tricks: an assistant prefill of {"category": " to force JSON, a system prompt reading "CRITICAL: you MUST output only JSON", and few-shot examples. Moving it to a current Claude model fails immediately with a 400 error because of the prefill. The migration:
# Before: steering the format with an assistant prefill.
# Worked on earlier Claude models; returns HTTP 400 on Claude 4.6 and later.
messages = [
{"role": "user", "content": ticket_prompt},
{"role": "assistant", "content": '{"category": "'},
]
# After: the format lives in the instructions, with the reason, and is checked in code.
system = (
"Classify the support ticket. Reply with only a JSON object: "
'{"category": "billing" | "outage" | "account" | "other", "reason": "<one sentence>"}. '
"A program parses your reply, so do not write anything before or after the object."
)
messages = [{"role": "user", "content": ticket_prompt}]
def parse_or_retry(call, request, retries=1):
for attempt in range(retries + 1):
text = call(request)
try:
return json.loads(text)
except json.JSONDecodeError:
if attempt == retries:
raiseThe prefill is replaced by an instruction that states the format and the reason for it, and a parser that retries once and then fails loudly. The capitalised warning becomes a plain sentence. The examples stay, wrapped in <example> tags. On an eval set of 300 labelled tickets, the team compares the old and new versions before switching traffic.
The same team then adds an OpenAI reasoning model as a fallback and a Gemini 3 model for a cheaper tier. The neutral spec does not change. The OpenAI renderer starts zero-shot and adds a Structured Outputs schema with the four categories as an enum; the Gemini renderer keeps the examples and removes the inherited temperature=0.2.
Measure every trick with an eval matrix
A trick is a hypothesis about one model. The only way to know whether it still helps is to run the same cases through each renderer, with and without the trick, and compare. Keep the harness small enough that people run it on every prompt change:
from dataclasses import replace
def run_matrix(cases, renderers, call, score):
"""cases: [(PromptSpec, expected)]; renderers: {name: spec -> request}, e.g.
functools.partial(render_claude, model=MODEL); call(name, request) -> text;
score(text, expected) -> 0 or 1."""
table = {}
for name, render in renderers.items():
passed = sum(score(call(name, render(spec)), expected) for spec, expected in cases)
table[name] = passed / len(cases)
return table
# Ablate one trick at a time: same cases, renderer with and without the trick.
baseline = run_matrix(cases, {"gemini": render_gemini}, call, exact_label)
no_shots = run_matrix([(replace(s, examples=[]), e) for s, e in cases],
{"gemini": render_gemini}, call, exact_label)
print(baseline, no_shots) # keep the examples only if they earn their tokensScore with exact checks where possible: label match, schema validity, presence of a required citation. Track cost beside quality: examples cost tokens on every call, which prompt caching reduces for a stable prefix. Rerun the matrix whenever a provider ships a new model version, not only when you change the prompt. The evals guide covers building the case set.
Failure modes
- Folklore that now fails validation. Prefill on Claude 4.6 and later,
budget_tokenson Claude 4.7 and later. These fail loudly, which is the good case; search your code base for them before upgrading. - Emphasis that now overshoots. Capitalised MUST and CRITICAL written for less attentive models make newer ones over-apply a rule, such as calling a tool on every turn.
- Scripted reasoning on reasoning models. "Think step by step" is, per OpenAI's guide, unnecessary, and it spends tokens for no measured gain.
- Old sampling settings. Low temperature carried over to Gemini 3.x can cause looping.
- Missing formatting. Markdown silently disappears from OpenAI reasoning-model output unless the developer message re-enables it.
- Single-model evals. A change validated on one model is shipped to all. Run the matrix across every renderer in production.
Trade-offs
Renderers add a layer and a test burden; for a single-model product with one prompt, a well-written prompt plus an eval set may be enough. Provider-specific features such as strict schemas or adaptive thinking are worth using even though they reduce portability, because they move guarantees from wording into the API; isolate them in the renderer so switching providers stays a local change. Prefer advice that holds across vendors over one model's quirks; the general advice survives upgrades.
What to do next
- Grep your code for assistant-turn prefills,
budget_tokens, "step by step" and hard-coded temperatures, and list which model each call targets. - Re-read each vendor's current prompting guide for the models you use and note the date you read it.
- Refactor your most important prompt into a neutral spec with one renderer per provider.
- Replace emphatic capitals with plain instructions that include the reason.
- Build a 100 to 300 case eval set from real traffic and run it through every renderer.
- Ablate one trick at a time, such as examples, tags or document order, and keep only those that measurably help.
- Rerun the matrix on every provider model release, and read role prompting next for the same evidence-first treatment of personas.