A prompt in production is code that is interpreted by a model you do not control. Change one sentence and accuracy can move by several points; keep the sentence and let the provider update the model behind an alias, and it can move again. Prompt versioning is the discipline that lets you answer three questions at any time: exactly what produced this output, what changed between the version that worked and the one that does not, and how fast can we go back.

This article covers the version itself: what it must contain, how to identify it, how changes are reviewed and gated, and how versions are resolved, logged, rolled out and rolled back. The service that stores and serves versions is described in Prompt registry architecture, and designing the evaluations that gate changes is covered in Prompt Evaluation Architecture.

Advertisement

The prompt is the wrong unit

Teams usually start by versioning the template text, and then discover that outputs changed when nobody touched it. Model behaviour is a function of everything sent and everything that configures the call. A version must therefore capture the whole bundle: the system and user templates; the few-shot examples, which are often loaded from a file or a database and change independently; the tool definitions and output schema, because a renamed field or reworded tool description changes behaviour as much as the template does; the model identifier; and the sampling parameters such as temperature and the token limit.

Two further inputs are often relevant and frequently forgotten: retrieval configuration for prompts fed by search, such as the index, the number of chunks and the ranking model, and any preprocessing applied to user input before it is substituted. If an input can change the output and is not in the bundle, it is an unversioned dependency, and an incident investigation will eventually hit it.

A prompt version is a bundle, identified by its content hashTemplate textsystem + user partsFew-shot examplespinned setTool + output schemasJSON SchemaModel snapshotdated id, not aliasSampling paramstemperature, max tokensCanonical JSONsha256 = version idEval reportscores bound to the hashAliasesprod, canary point at a hashTracesevery call logs the hashevaluated asreleased asresolved at runtimeRollback = move the alias back to an older hashno redeploy, no edit of the old versionChange any input and you get a new id; nothing is ever edited in place.
Every input that can change model behaviour goes into one bundle; its canonical hash is the version identity that evals, aliases and traces refer to.

Identity: content hashes underneath, labels on top

A version needs an identity that cannot lie. Hand-maintained numbers drift: someone edits a file and forgets to bump the number, and two different bundles share one label. The robust approach is content addressing. Serialize the bundle canonically, with sorted keys, fixed separators and referenced files inlined, then hash it. The hash is the version identifier; any change to any input yields a new one, and identical bundles get identical ids wherever they are built.

Humans still need names, so layer two kinds of label on top. A semantic version such as 1.5.0 communicates intent in review. Aliases such as prod and canary are mutable pointers to immutable versions, and releasing or rolling back means moving a pointer. Never edit a released version in place; if it is wrong, release a new one.

# prompts/ticket_classifier/1.5.0.yaml  -- immutable once released
name: ticket_classifier
version: 1.5.0
model: vendor-model-2026-06-15        # a dated snapshot id, never a moving alias
params: {temperature: 0, max_tokens: 200}
template:
  system: |
    You classify customer support tickets into exactly one category.
    Categories: billing, outage, account_access, feature_request, other.
  user: |
    Ticket:
    {ticket_text}
few_shot: examples/ticket_classifier/v3.jsonl
output_schema: schemas/ticket_category.v2.json
import hashlib, json, pathlib, yaml

ROOT = pathlib.Path("prompts")

def load_bundle(name: str, version: str) -> dict:
    spec = yaml.safe_load((ROOT / name / f"{version}.yaml").read_text())
    # Inline every referenced file so the hash covers its contents, not its path.
    spec["few_shot"] = (ROOT.parent / spec["few_shot"]).read_text()
    spec["output_schema"] = json.loads((ROOT.parent / spec["output_schema"]).read_text())
    return spec

def version_id(bundle: dict) -> str:
    identity = {k: v for k, v in bundle.items() if k != "version"}   # labels are not identity
    canonical = json.dumps(identity, sort_keys=True, separators=(",", ":"), ensure_ascii=False)
    return "pv_" + hashlib.sha256(canonical.encode("utf-8")).hexdigest()[:16]

def resolve(name: str, alias: str, aliases: dict) -> tuple[dict, str]:
    version = aliases[name][alias]            # e.g. {"ticket_classifier": {"prod": "1.4.2"}}
    bundle = load_bundle(name, version)
    return bundle, version_id(bundle)

Note that the hash excludes the version label itself, so relabelling does not create a new identity, and that it inlines the few-shot file and schema so editing an example changes the hash even though the manifest did not change. The templates are kept as separate fields rather than concatenated, which keeps diffs readable; reusable template patterns are covered in Prompt templates.

Advertisement

Change classes and what each requires

Semantic versioning maps onto prompts imperfectly, because even a typo fix can shift behaviour. It is still worth defining change classes by their effect on consumers. A major change alters the output contract: a new category, a renamed field, a different format. Downstream parsers and stored data must change with it, so it ships with a schema version bump and a migration plan. A minor change alters behaviour within the contract: new instructions, new examples, a different model snapshot. A patch is intended to be behaviour-neutral: wording, formatting, comments.

The rule that matters is that every class runs the evaluation suite, because intent is not evidence. What differs is the bar and the rollout. Patches need no regression beyond the noise floor; minor changes need an improvement or a justification; major changes need consumer sign-off and a coordinated release, because a parser expecting the old schema will fail on the new output immediately. Structured output contracts are discussed in Structured output.

The git and CI workflow

Keep prompt bundles in the same repository as the code that calls them, in plain text files, so a pull request shows the diff of the template, the examples and the schema together with the code that consumes them. Reviewers should be able to read the change as prose. On every pull request that touches a bundle, CI computes the new version id, runs the evaluation set against both the production version and the candidate, and fails the build if quality drops beyond agreed limits.

# ci/check_prompt_change.py -- runs on every pull request that touches prompts/
import json, sys

base = json.load(open("eval/base_report.json"))       # scores for the version in prod
cand = json.load(open("eval/candidate_report.json"))  # scores for the changed bundle

assert cand["version_id"] != base["version_id"], "bundle unchanged but files differ?"
failures = []
for metric, floor in {"accuracy": -0.01, "schema_valid": 0.0}.items():
    delta = cand["overall"][metric] - base["overall"][metric]
    if delta < floor:
        failures.append(f"{metric} moved {delta:+.3f} (allowed {floor:+.3f})")
for slice_name, s in cand["slices"].items():           # regressions hide in slices
    if s["accuracy"] < base["slices"][slice_name]["accuracy"] - 0.03:
        failures.append(f"slice {slice_name} regressed")
if failures:
    sys.exit("prompt change blocked:\n" + "\n".join(failures))
print("prompt change ok:", cand["version_id"])

Gate on slices, not only on the overall score. A change that improves average accuracy by two points while dropping a 5% slice, say non-English tickets, by ten points will pass an average-only gate and fail real users. Store each evaluation report keyed by version id, so the evidence for every release is retrievable later. When the suite is expensive, run a fast subset on every commit and the full set before merge.

Some teams keep prompts in a registry service with a web editor so non-engineers can edit them. That is a valid choice, but the same properties must hold: immutable versions, content identity, mandatory evaluation before an alias moves, and an audit log of who moved which alias. A registry without those is a shared text box.

Runtime: resolve, call, log the hash

At runtime the application asks for a prompt by name and alias, resolves it to a bundle and version id, and calls the model. Resolution can happen at deploy time, where the bundle is baked into the build, or at request time, where the service reads the alias from a registry with a short cache. Deploy-time resolution is simpler and reproducible; request-time resolution allows rollback without a deploy. Many teams resolve at request time with a cache of a minute or so, and fall back to the last known good bundle if the registry is unavailable.

Whatever the mechanism, log the version id with every model call, alongside the model identifier, the alias, latency, token counts and whether the output parsed. Without it, a complaint about a bad answer cannot be traced to the bundle that produced it, and A/B comparisons between versions are impossible. Store the rendered inputs for a sample of calls as well, subject to your data-retention rules, so failures can be replayed against a candidate version. The debugging loop that uses these traces is covered in Prompt debugging workflow.

Pinning the model

Providers typically offer two kinds of model identifier: a moving alias that is updated to newer models over time, and a dated or versioned snapshot that is not supposed to change. A bundle that names an alias is not a version, because its behaviour can change without any change on your side. Pin the snapshot in the bundle, and treat moving to a newer model as a minor change that goes through the full evaluation gate like any other.

Snapshots are retired on a published schedule, so model migration is routine work rather than an emergency. Track the deprecation dates of every snapshot your bundles pin, start the migration well before them, and expect to retune: prompts tuned for one model often need adjustment for its successor, particularly few-shot examples and formatting instructions. Running the evaluation suite against the new snapshot with the old prompt first tells you how much retuning is needed.

Rollout and rollback

Moving an alias directly from one version to another exposes every user at once. A canary alias receives a small, deterministic share of traffic, chosen by hashing a stable key such as the user id so a user does not flip between versions mid-conversation. Compare the canary's online metrics with production: parse failure rate, refusals, latency, cost per call, and any downstream signal such as escalation rate. Widen the share in steps, and promote by pointing prod at the canary's version.

import hashlib

def pick_alias(user_id: str, canary_percent: int) -> str:
    # Deterministic: the same user always sees the same version during a rollout.
    bucket = int(hashlib.sha256(user_id.encode()).hexdigest(), 16) % 100
    return "canary" if bucket < canary_percent else "prod"

alias = pick_alias(request.user_id, canary_percent=5)
bundle, vid = resolve("ticket_classifier", alias, ALIASES)
result = call_model(bundle, ticket_text=request.text)
log.info("llm_call", prompt="ticket_classifier", prompt_version=vid, alias=alias,
         model=bundle["model"], latency_ms=result.latency_ms, parsed_ok=result.parsed_ok)

Rollback is moving the alias back, which takes effect within the resolution cache's lifetime. Two things can make it harder than that. Outputs stored under a new schema may not be readable by consumers of the old one, so major changes need consumers that accept both formats during the rollout. And caches keyed only by input will serve outputs from the bad version after rollback; include the version id in any cache key for model outputs.

Worked example: a ticket classifier release

A support team's classifier at version 1.4.2 routes tickets into five categories. Agents report that login problems are often classified as billing. An engineer adds two instructions and four new examples, producing 1.5.0. The figures that follow are illustrative. On the 1,200-ticket evaluation set, overall accuracy rises from 0.91 to 0.93 and the account-access slice from 0.82 to 0.90, but the Spanish-language slice falls from 0.88 to 0.81, because all the new examples are English and the model now over-weights their phrasing. The slice gate blocks the merge.

The engineer adds two Spanish examples, producing a new bundle with a new hash. All slices are now within tolerance and the change merges. In canary at 5% for a day, parse failures and latency match production, and agent reclassification of canary tickets falls. The alias moves to 25% and then 100%. A week later the provider announces a retirement date for the pinned snapshot; the team runs 1.5.1, identical except for the newer snapshot, through the same gate, finds formatting drift in the output, adjusts one instruction in 1.5.2 and ships it through the same canary process. Every trace from the period identifies which of the four versions handled the ticket.

Failure modes

SymptomCauseFix
Outputs changed, nobody edited the promptmodel alias, examples or schema outside the bundlecontent-hash the whole bundle; pin snapshots
Cannot say which prompt produced a bad answerversion id not loggedlog the hash on every call
Average improved, complaints roseslice regressiongate on slices
Rollback did not fix itoutput cache keyed without versioninclude the version id in cache keys
Parser errors after releaseoutput contract changed as a minorclassify contract changes as major; dual-read
Two environments disagreelabel reused for different contentidentity by hash, labels advisory

Trade-offs

Git-based versioning gives you review, history and atomic changes with code for free, but makes every prompt change a deploy unless aliases are resolved at runtime. A registry service lets product owners iterate without engineers and allows instant rollback, but needs its own access control, audit and availability. Content hashing adds a little machinery and guarantees honest identity. Strict evaluation gates slow iteration and catch the regressions that cost the most. For most teams the right shape is bundles in git, hashes as identity, aliases resolved at runtime with a short cache, and evaluation as a required check.

What to do next

  1. List every input to each production prompt call and move anything outside the bundle, especially examples and schemas, into it.
  2. Compute a canonical content hash per bundle and log it, with the model id, on every model call.
  3. Replace moving model aliases in bundles with dated snapshots, and record each snapshot's retirement date.
  4. Add a CI gate that compares candidate and production evaluation reports overall and per slice.
  5. Introduce prod and canary aliases with deterministic user bucketing, and rehearse a rollback.
  6. Add the version id to every cache key for model outputs.
Key takeaway: Version the bundle, not the template: template text, examples, tool and output schemas, a pinned model snapshot and sampling parameters together determine behaviour. Identify each bundle by a canonical content hash, put human labels and mutable aliases on top, and never edit a released version. Gate every change, including typo fixes, on evaluation overall and per slice, log the version id with every call, roll out through a deterministic canary, and roll back by moving an alias, with version ids in cache keys so rollback really takes effect.