Rolling back a web service usually takes seconds: point the load balancer at the previous container image and the old code is running before anyone notices. Rolling back an LLM serving release on GPUs is different. The artifact is tens or hundreds of gigabytes of weights that must cross a network or a disk into GPU memory, the serving engine must profile memory and capture CUDA graphs before it can take traffic, the prefix cache that made the fleet fast is invalid the moment the weights change, and the accelerators you would roll back onto are probably already running the release you are trying to escape.
So an LLM rollback has to be designed before the release, not improvised during the incident. This article explains where rollback time goes on GPUs, what exactly must be rolled back together, how to trigger it automatically, how to handle in-flight streams, why adapter-based releases roll back almost instantly, and how to budget the standing GPU cost of a fast rollback. The release patterns themselves, canary and blue-green, are covered in LLM canary deployment and LLM blue-green deployment; the artifact bookkeeping is in rollback strategies for models, prompts and indexes.
Where rollback time goes
Time to roll back is the sum of four terms, and only the first is fast:
- Routing change. Setting the router weights back to the old version takes seconds, but only helps if old-version replicas are still running.
- Weight transfer. A 70-billion-parameter model in bf16 is about 140 GB. That must reach every replica, from object storage, a local NVMe cache or host memory. This term dominates when nothing is cached.
- Engine startup. Engines such as vLLM load weights onto each GPU in the tensor-parallel group, run a profiling pass to size the KV cache, and capture CUDA graphs for common batch sizes. Expect tens of seconds to a few minutes; measure yours.
- Cache warmup. Prefix-cache entries are KV tensors computed by specific weights, so they are useless after a weight change. Time to first token stays elevated until shared system prompts are recomputed and cached, as described in paged KV cache.
The architecture follows from that list. The fastest rollback never has to load anything, which means old-version replicas are still warm somewhere. Everything else is about making the slow path shorter and keeping it from overloading the fleet while it runs.
Roll back the whole serving tuple
The most common rollback that does not work is one that restores the weights and nothing else. An LLM release is a tuple, and every member is coupled to the others:
| Component | Why it must match | Failure if it does not |
|---|---|---|
| Weights | the model itself | - |
| Tokenizer | token IDs are model-specific | garbled output, broken logit bias |
| Chat template | trained turn and tool-call format | malformed tool calls, role leakage |
| Engine image | kernels, schedulers, defaults | new OOMs, changed sampling |
| Quantization config | scales are computed per checkpoint | silent quality loss |
| Speculative draft model | must share the tokenizer and track the target | acceptance rate collapses |
| LoRA adapters | trained against one base model | nonsense or load failure |
Pin each member by immutable digest, not by tag, and verify the whole tuple when the previous release is retained: hash the weight shards on the cache disk, check the tokenizer and template hashes, and confirm the engine image digest is still pullable. A rollback target that was garbage-collected from the registry last week is not a rollback target.
Rollbacks of the engine image and of the weights are different events. If a new engine version caused the problem, keep the new weights and roll back the image, and vice versa. Changing both in one release doubles the search space during the incident.
Automatic triggers and draining
Human-triggered rollbacks are slow because someone has to notice, believe the graphs and decide. Write the decision down in advance as guard metrics compared between the new cohort and the old one over the same window, and let a controller act on it:
import time
GUARDS = { # metric -> (kind, threshold), new cohort vs old cohort
"error_rate": ("ratio", 1.5),
"ttft_p95_ms": ("ratio", 1.3),
"tpot_p95_ms": ("ratio", 1.3),
"refusal_rate": ("delta", 0.02),
"judge_win_rate": ("floor", 0.45), # pairwise LLM-judge win rate of new over old
}
def breached(metric, new, base):
kind, t = GUARDS[metric]
if kind == "ratio":
return new > base * t
if kind == "delta":
return new - base > t
return new < t
def rollback_controller(router, metrics, release, min_samples=2000, strikes_needed=3):
strikes = 0
while release.active():
new = metrics.window(release.new_cohort, minutes=5)
old = metrics.window(release.old_cohort, minutes=5)
if new.samples >= min_samples:
bad = [m for m in GUARDS if breached(m, new[m], old[m])]
strikes = strikes + 1 if bad else 0
if strikes >= strikes_needed:
router.set_weights({release.old: 100, release.new: 0}) # stop new traffic
release.mark("rolled_back", reason=bad)
router.drain(release.new, max_wait_s=120) # let streams end
return bad
time.sleep(30)Three details matter. Compare cohorts, not the new version against yesterday, so that a traffic surge does not look like a regression. Require a minimum sample and several consecutive breaches, so one slow batch does not flap the release. And include at least one quality signal, because the worst LLM regressions are fast, healthy-looking and wrong; a mirrored-traffic comparison like the one in LLM shadow deployment is the best source of it before the release takes real traffic.
Streaming responses complicate the cut. Drain rather than kill: stop sending new requests to the bad version and give open streams a bounded time to finish. Kill them only if the regression is harmful output rather than slowness. Conversations need no migration, because chat history is text and the old model can continue it, unless the new release changed tool schemas or stored tokenized state.
Adapter releases roll back instantly
When a release is only a fine-tuned adapter on an unchanged base model, rollback can be nearly free. With vLLM, start the server with --enable-lora and register both the current and the previous adapter with --lora-modules; clients select an adapter by its name in the model field, so rolling back is a gateway change from one name to the other, with no weight transfer and no restart.
vllm serve /models/base-8b --enable-lora --max-loras 4 \
--lora-modules support-v6=/models/adapters/support-v6 \
support-v7=/models/adapters/support-v7
# gateway config: rollback = point the alias back at the previous adapter
# support-bot -> support-v7 (release)
# support-bot -> support-v6 (rollback)vLLM can also load and unload adapters at runtime through /v1/load_lora_adapter and /v1/unload_lora_adapter when VLLM_ALLOW_RUNTIME_LORA_UPDATING is set, but its documentation warns that this feature carries security risks and should not be used in production outside an isolated, fully trusted environment. Preloading both versions avoids the question. The same idea extends to whole models: keep the previous version resident on a fraction of the fleet so the rollback is a routing change.
Worked example: a 70B fleet
Consider a 70B dense model in bf16 served with tensor parallelism 4, so each replica uses 4 GPUs and holds about 140 GB of weights (35 GB per GPU). The fleet is 12 replicas, 48 GPUs, and v2 has just reached 100% of traffic. The figures below use assumed rates, not measured specifications: 2 GB/s per node from object storage, 6 GB/s from local NVMe, and 60 seconds of engine startup. Replace them with your own measurements.
| Strategy | First relief | Full v1 capacity | Standing cost |
|---|---|---|---|
| Cold: pull v1 from object storage, 3 replicas at a time | about 130 s | about 9 min, at 75% capacity throughout | none |
| NVMe cache: v1 kept on local disk, 3 at a time | about 83 s | about 6 min | disk space only |
| Partial warm: 3 v1 replicas kept running | seconds, for 25% of traffic | about 4 min via NVMe for the rest | 12 GPUs (25%) |
| Blue-green: full v1 fleet kept running | seconds, for all traffic | seconds | 48 GPUs (100%) |
The cold row is 140 GB at 2 GB/s, 70 seconds, plus 60 seconds of startup, repeated in four waves so that three quarters of the fleet keeps serving. Twelve replicas pulling at once would need 24 GB/s from storage, and many object stores throttle long before that. The NVMe row cuts transfer to about 23 seconds per wave. The partial warm pool absorbs a quarter of traffic immediately, while the remainder is either served by the bad version a few minutes longer or queued, which is a choice to make in advance per regression type.
The usual compromise is to keep the warm pool for the first day after a release, when most regressions surface, then release those GPUs and fall back to the NVMe path. A quarter of the fleet for one day per release is often far cheaper than one hour of a fleet-wide incident.
Validate the table with a drill rather than arithmetic. In staging, or on a small production slice at a quiet hour, trigger the controller deliberately and time each phase: router flip, first v1 replica healthy, last replica healthy, and time-to-first-token back within its objective. The gaps between those timestamps tell you which term to attack next. If transfer dominates, stage weights closer to the GPUs; if startup dominates, look at graph capture and profiling settings; if cache warmup dominates, replay the most common system prompts before reopening traffic.
Failure modes
- No capacity to roll back onto. Every GPU runs v2, so restoring v1 means removing serving capacity at peak. Plan the surge or warm pool before the release.
- Thundering herd on storage. All replicas pull 140 GB at once and the object store throttles, making a 2-minute rollback a 20-minute one. Stage weights on local disk.
- Partial tuple. Weights restored, template not: tool calls break in a new way and the incident gets worse.
- TTFT spike after the flip. The old pool's prefix cache is cold and its batch is suddenly four times larger. Pre-warm shared system prompts with synthetic requests.
- Flapping. Automatic rollback with no hysteresis bounces between versions. Require consecutive breaches and lock the release once rolled back until a human clears it.
- Irreversible side effects. Rollback stops new harm but does not undo messages already sent or tool actions already taken by the bad version. Log them for follow-up.
Trade-offs
Rollback speed is bought with idle GPUs or with disk and engineering effort. Blue-green buys instant rollback at double cost during the window. A partial warm pool buys instant relief for part of the traffic. Local weight caches buy minutes instead of tens of minutes for the price of disk. Adapter-based releases make rollback almost free but constrain what a release can change. Choose per model by asking how much an hour of a bad release costs, and spend up to that on standing readiness. The broader reliability picture, from health checks to multi-region failover, is covered in GPU serving reliability.
What to do next
- Measure your real times: weight pull per node, engine startup and time for TTFT to recover after a restart.
- Define the serving tuple for each model and pin every member by digest.
- Keep the previous tuple on local NVMe of every serving node until the next release is proven.
- Write guard metrics and thresholds, and run the controller in shadow mode on the next release.
- Decide the warm-pool size and duration per model from the cost of an hour of regression.
- Rehearse a rollback in staging every release, timing each phase.
- For adapter releases, preload both adapters and make rollback a gateway alias change.