Rolling back a web service is a solved problem in principle: keep the previous binary, route traffic back to it, and deal carefully with schema migrations. An AI application complicates this because the thing you ship is not one binary. A production LLM feature is a combination of model weights, often an adapter on top of a base model, a set of prompt templates and tool schemas, a retrieval index built with a particular embedding model, guardrail thresholds, sampling parameters, and the application code that glues them together. Any of these can change independently, several of them change on schedules owned by different teams, and some of them are only valid in combination with specific versions of the others.
This article is about reversing those AI-specific artifacts. It assumes the general mechanics of rolling back application code and database migrations are already in place, and focuses on what is different: how to describe a release so that it can be reversed as a unit, which artifacts are coupled so tightly that they must move together, how to keep the previous state warm enough that rollback takes seconds rather than an hour of weight loading or index rebuilding, and what cannot be undone at all, such as actions an agent already took or documents a user asked you to delete.
A release is a manifest, not a model
The first design decision is what the unit of rollback is. Teams that version each artifact separately, with a model registry here, a prompt repository there and an index naming convention somewhere else, discover during an incident that nobody can say which combination was live at 14:05 yesterday. The fix is a release manifest: a small immutable document that pins every artifact by content digest, gets its own identifier, and is the only thing the serving layer reads to decide what to run.
release: r-0412
created: 2026-09-28T10:14:00Z
parent: r-0409
model:
base: registry://llm-base@sha256:7c1e...
adapter: registry://support-lora@sha256:a93f...
tokenizer: registry://tok@sha256:11b0...
prompts:
bundle: prompts://support@sha256:4d2a... # templates + tool JSON schemas
retrieval:
index: idx-0912
embedder: registry://e5-v3@sha256:9e77...
tombstone_log_offset: 88231
params:
temperature: 0.2
max_output_tokens: 1024
guardrails:
toxicity_block: 0.82
pii_redaction: v5Two properties matter. Every reference is to something immutable: a digest, not a tag like latest, because a tag can move underneath you and a rollback to a moving target is not a rollback. And every request log line carries the release identifier, so any metric, complaint or eval result can be attributed to exactly one combination. With that in place, rollback means pointing production at the previous manifest, and the question of which prompt went with which index answers itself.
Pointers, not copies
Serving should resolve an alias such as prod to a manifest identifier, and deploys and rollbacks should both be a single atomic update of that alias. Nothing is copied or rebuilt during a rollback; the old artifacts are already in storage because they are immutable and retained. Use a compare-and-set so two operators, or an operator and an automated rollback controller, cannot race each other into an unintended state.
def flip_alias(store, alias: str, expected: str, target: str, actor: str, reason: str):
"""Atomically move `alias` from `expected` to `target`, or fail loudly."""
manifest = store.get_manifest(target)
verify_compatibility(manifest) # see next section; refuse bad combos
ok = store.compare_and_set(f"alias/{alias}", expected, target)
if not ok:
current = store.get(f"alias/{alias}")
raise RuntimeError(f"{alias} is {current}, not {expected}; re-read before flipping")
store.append_audit(alias=alias, frm=expected, to=target, actor=actor, reason=reason)Serving replicas watch the alias and switch on change. The rollback latency you actually get is the propagation time of that watch plus any per-artifact warm-up, so measure it end to end rather than assuming the flip is instantaneous. A replica that caches the alias with a five-minute refresh gives you a five-minute rollback no matter how fast the flip is.
Compatibility edges between artifacts
Several artifacts are only meaningful in combination, and a partial rollback across one of these edges produces a system that is worse than either release. The most important is the embedding model and the index built with it. Vectors from two different embedding models live in different spaces; querying an index built with model A using a query vector from model B returns confident nonsense with no error. If a release upgraded the embedder and rebuilt the index, rolling back the embedder without the index, or the reverse, silently destroys retrieval.
Prompts are coupled to models more tightly than they look. A prompt tuned for one model's chat template, tool-calling format or instruction-following quirks can degrade sharply on another. Output parsers are coupled to prompts: if the new prompt asks for a JSON field the old parser does not know, or the old prompt omits a field the new parser requires, failures appear downstream as validation errors. Adapters are coupled to the exact base weights they were trained on, and a tokenizer change invalidates any cached token counts or prefix caches.
Encode these edges as machine-checked constraints rather than tribal knowledge. The manifest validator should refuse a combination where the index metadata names a different embedder digest than the manifest, where an adapter's recorded base digest differs from the base, or where the prompt bundle declares a minimum parser version the application does not meet. Because rollbacks move whole manifests, these checks mostly protect against hand-assembled emergency releases, which is exactly when people make mistakes.
Rolling back a model
The practical obstacle is load time. A few-billion-parameter model reloads in seconds, but a large model across multiple GPUs can take many minutes to pull weights, load and warm up, and during an incident those minutes are the outage. The standard answer is to keep the previous release warm: retain enough capacity to serve the previous manifest, either as a small always-on pool that can scale out or as the old deployment left running at reduced replicas for a soak period after every release. This costs GPU hours, so size the soak window to how long regressions typically take to surface; for many teams a day or two covers the bulk of detected regressions.
Watch for state that is keyed implicitly to a model. Prefix and KV caches must be keyed by model and tokenizer digest. Semantic response caches that store answers by query similarity will happily serve answers from the bad release after rollback unless they are keyed by release identifier or flushed. Long conversations are another trap: a session that started on the new model may carry a history full of formats or tool calls the old model handles poorly, so either pin in-flight sessions to the release they started on and drain them, or accept and monitor a short period of degraded multi-turn quality.
Rolling back a prompt
Prompts change more often than anything else, and they are usually treated as configuration fetched at runtime so they can change without a code deploy. That is good for rollback speed and bad for discipline. The loader should always resolve prompts through the release manifest, cache with a short time to live, and support an emergency pin that overrides the manifest for one prompt while an incident is investigated.
class PromptLoader:
def __init__(self, store, ttl_s=30):
self.store, self.ttl_s, self._cache = store, ttl_s, {}
def get(self, release_id: str, name: str) -> Prompt:
pin = self.store.get(f"pin/prompt/{name}") # incident override
key = pin or f"{release_id}/{name}"
hit = self._cache.get(key)
if hit and time.monotonic() - hit.at < self.ttl_s:
return hit.prompt
prompt = self.store.load_prompt(key) # immutable by digest
self._cache[key] = Cached(prompt, time.monotonic())
return promptA pin is a deliberate break in the manifest-as-unit rule, so it must be visible: log it on every request, alert while any pin is active, and require that the next release either absorbs or removes it. Also remember that a prompt's effects outlive it. Conversation memory, summaries and agent scratchpads written under the new prompt remain in storage after rollback. If the new prompt changed the format of stored memories, the old prompt may misread them, which is another reason to version persisted formats and not only the prompts that produce them.
Rolling back a retrieval index
Indexes should be built blue-green: a new index is built alongside the live one, evaluated, and made live by moving an alias. Keep the previous index for the soak window, so rollback is an alias flip rather than a multi-hour rebuild. The subtlety is that the live index kept changing after the new one was built. Documents were added and edited, and some were deleted, possibly because a user or a legal process required it.
Rolling back to an index that is two days old must not resurrect those deletions. The robust design keeps an ordered log of document changes, with deletes as tombstones, and records in the manifest the log offset each index has applied. Rolling back means flipping the alias and then replaying the log from the old index's offset, applying at least the tombstones before the index serves traffic, and the upserts shortly after. If replay cannot finish quickly, apply deletes as a query-time filter against a deny list until it does. Treat an index that could serve deleted content as a compliance incident, not a quality blip.
What cannot be rolled back
Some effects are external and permanent. Answers already shown to users, emails an agent sent, tickets it closed and API calls it made cannot be recalled by flipping an alias. For these, rollback is containment plus remediation: stop the bad release, then use request logs keyed by release identifier to enumerate affected users and actions, and run compensating operations where they exist. This is also the argument for staging risky capabilities behind a separate flag with its own kill switch, so that disabling agent write actions does not require reverting the whole release.
Feedback and training data are a quieter version of the same problem. Ratings, corrections and conversations collected during a bad release will flow into evaluation sets and fine-tuning data unless they are tagged with the release identifier and can be filtered. A regression that is rolled back in production but then learned by the next training run has not really been rolled back.
Deciding when to roll back
Rollback should be cheap enough to be the default response to a suspected regression: restore the previous state first, then diagnose. Decide the triggers before the release, not during the incident. Good automated triggers are guardrail metrics with fast signal: structured-output parse failure rate, tool-call error rate, refusal rate, safety-filter block rate, retrieval empty-result rate, p95 latency and cost per request, each compared against the previous release on concurrent traffic during a canary. Slower quality signals such as user ratings or online eval scores feed a human decision rather than an automatic flip, because they are noisy over short windows.
def should_roll_back(canary, baseline, min_requests=2000):
if canary.requests < min_requests:
return None # not enough evidence yet
checks = {
"parse_fail": canary.parse_fail_rate > baseline.parse_fail_rate * 1.5 + 0.002,
"tool_errors": canary.tool_error_rate > baseline.tool_error_rate * 1.5 + 0.002,
"p95_latency": canary.p95_ms > baseline.p95_ms * 1.25,
"cost": canary.cost_per_req > baseline.cost_per_req * 1.30,
"empty_ctx": canary.empty_retrieval > baseline.empty_retrieval + 0.01,
}
tripped = [k for k, bad in checks.items() if bad]
return tripped or FalseThe additive floors keep tiny baselines from producing hair-trigger ratios, and the minimum request count avoids acting on noise. Roll forward instead only when the fix is trivial and certain, or when rollback is impossible because the old release depends on something that no longer exists, such as a deprecated hosted model. That last case is worth planning for: if you depend on a provider's model versions, record their deprecation dates next to your manifests, because a release whose rollback target is scheduled to disappear has no rollback.
Drills and verification
A rollback path that has never been exercised is a hope, not a capability. Run rollbacks deliberately in staging on every release, and periodically in production during low traffic, measuring the time from decision to the previous release serving all traffic. After any rollback, run verification probes: a small set of golden queries with known good answers, a retrieval probe that confirms recently deleted documents do not appear, a check that no replica still reports the rolled-back release identifier, and a cache check that responses are not being served from the bad release's entries. Track rollback time as an objective alongside deploy frequency; if it creeps up because warm capacity was trimmed to save money, that is a risk decision someone should make explicitly.
Failure modes
- Mutable references. A manifest that points at latest tags cannot be restored because the tags have moved.
- Split coupling. Rolling back the embedder but not the index, or the prompt but not the parser, silently breaks the system.
- Cold standby. The previous model has to be loaded from scratch and the rollback takes longer than the incident budget.
- Stale caches. Semantic or response caches keep serving the bad release's answers after the flip.
- Resurrected documents. An old index returns content deleted after it was built because tombstones were not replayed.
- Poisoned feedback. Data collected during the bad release flows into the next training or eval set.
- Forgotten pins. An emergency prompt pin outlives the incident and silently overrides future releases.