Most incidents in LLM serving are not caused by the model. They come from a change somebody believed was harmless: a new engine image with a different default for prefix caching, a chat template that dropped a newline, a driver upgrade rolled out with the same ticket as a quantized checkpoint, a system-prompt edit that nobody tied to a release. The outputs shift, latency moves, memory headroom shrinks, and nobody can say which of four simultaneous changes did it.
This article treats change management as an engineering system rather than paperwork. It defines what counts as a change in a GPU serving stack, puts all of it into one content-addressed release manifest, classifies risk automatically from the manifest diff, attaches required gates to each class, and makes rollback a capacity plan rather than a hope. The live traffic split itself is covered in LLM Canary Deployment; this page is everything before and around it.
What counts as a change
Start with an inventory of everything that can change model behaviour or capacity. The table is deliberately broader than the weights, because the weights are the layer teams already control well.
| Layer | Examples | What it can break |
|---|---|---|
| Weights | new checkpoint, merged adapter | quality, safety, output format |
| Numerics | bf16 to FP8 weights, KV cache dtype | quality at the tails, long-context recall, memory headroom |
| Interface | tokenizer, chat template, special tokens | tool calling, stop sequences, parsing downstream |
| Runtime | engine image, scheduler flags, speculative decoding | latency, throughput, small output drift |
| Platform | driver, GPU SKU, tensor-parallel degree | kernels selected, memory per GPU, failure rate |
| Behaviour | system prompt, sampling defaults | tone, refusals, task success |
Two observations follow. Several layers change outputs without changing a single weight: kernel selection depends on driver, SKU, batch shape and parallel layout, and a template change alters every token the model sees. And several layers change capacity without changing outputs: FP8 weights free memory for KV cache, while a higher tensor-parallel degree changes per-GPU memory and interconnect traffic. Change management has to see all of them, which is why the record of a release must be more than a model name and a tag.
The release manifest
The release manifest is one document that pins every layer by content hash. Tags such as latest or v2 are banned; images are referenced by digest, weight files by the hash of their contents, templates and prompts by the hash of the text. The hash of the canonicalised manifest is the release id, so two releases with identical content have the same id by construction, and any difference, however small, produces a new one.
release: chat-assistant
weights:
repo: models/assistant-70b
digest: sha256:9f1c...e2 # content hash of the weight files, not a tag
quantization: bf16
tokenizer:
digest: sha256:41aa...07
chat_template:
digest: sha256:c3d0...9b
engine:
image: registry.local/serving@sha256:77be...10
flags: {max_num_seqs: 256, prefix_caching: true, kv_cache_dtype: auto}
platform:
gpu_sku: H100-80GB
driver: "<pinned driver version>"
tensor_parallel: 4
behavior:
system_prompt_digest: sha256:0b9e...44
sampling_defaults: {temperature: 0.7, top_p: 0.9}Two practices make the manifest trustworthy. First, every serving replica reports its effective manifest at startup, computed from what it actually loaded, not from configuration it was given. A drift detector compares reported ids with the intended release and pages when they differ, which catches nodes that pulled a stale image or a hand-edited flag. Second, the manifest lives in version control next to the evaluation configs, so the review of a change is a diff, and the diff is what the risk classifier reads.
The pipeline in one picture
Classifying risk from the diff
The classifier maps each changed manifest key to a risk class, then takes the union of the required gates. Unmapped keys are treated as model changes, the strictest class, so adding a field to the manifest without classifying it costs more gates rather than fewer.
RISK = {
"weights.digest": "model",
"weights.quantization": "numerics",
"engine.flags.kv_cache_dtype": "numerics",
"tokenizer.digest": "interface",
"chat_template.digest": "interface",
"engine.image": "runtime",
"engine.flags": "runtime", # any other engine flag
"platform.driver": "platform",
"platform.gpu_sku": "platform",
"platform.tensor_parallel": "platform",
"behavior.system_prompt_digest": "behavior",
"behavior.sampling_defaults": "behavior",
}
GATES = {
"model": ["task_evals", "safety_evals", "golden_set", "perf_bench", "kv_capacity", "canary_full"],
"numerics": ["task_evals", "golden_set", "perf_bench", "kv_capacity", "compat_matrix", "canary_full"],
"interface": ["template_render_tests", "task_evals", "golden_set", "canary_full"],
"runtime": ["golden_set", "perf_bench", "compat_matrix", "canary_short"],
"platform": ["compat_matrix", "node_burn_in", "perf_bench", "golden_set", "canary_short"],
"behavior": ["task_evals", "safety_evals", "canary_short"],
}
def flatten(d, prefix=""):
out = {}
for k, v in d.items():
key = f"{prefix}{k}"
if isinstance(v, dict):
out.update(flatten(v, key + "."))
else:
out[key] = v
return out
def risk_of(key):
# Longest matching prefix wins; anything unmapped is treated as a model change.
parts = key.split(".")
for n in range(len(parts), 0, -1):
if ".".join(parts[:n]) in RISK:
return RISK[".".join(parts[:n])]
return "model"
def classify(old, new):
a, b = flatten(old), flatten(new)
changed = sorted(k for k in a.keys() | b.keys() if a.get(k) != b.get(k))
classes = sorted({risk_of(k) for k in changed})
gates = sorted({g for c in classes for g in GATES[c]})
return {"changed": changed, "classes": classes, "gates": gates, "split": len(classes) > 1}Run against a change request that switches the weights from bf16 to FP8 and bumps the engine image in the same release, the classifier reports two changed keys, engine.image and weights.quantization, classes numerics and runtime, seven gates (canary_full, canary_short, compat_matrix, golden_set, kv_capacity, perf_bench, task_evals), and split set to true. The split flag is the most valuable output. A release that mixes classes cannot be diagnosed when it regresses, so the pipeline asks for two releases: engine first, because it should not change quality and has the shorter canary, then quantization on top of a known-good engine.
Gates and what each one proves
Each gate answers a specific question, and each has a pass rule decided before the change is built, not after the numbers arrive.
- Task evals and safety evals measure quality on held-out suites with a minimum detectable difference you have computed in advance. Regression statistics and noise floors are covered in LLM Performance Regression Analysis.
- Golden set is a few hundred fixed prompts decoded greedily, compared with the previous release. It is a tripwire for gross breakage such as template or tokenizer errors, not a quality measure.
- Template render tests render every conversation shape your product sends (system prompts, tool calls, multi-turn) and compare token ids against expected fixtures.
- Perf bench replays a recorded traffic shape and checks time to first token, inter-token latency and throughput at the target concurrency.
- KV capacity computes how many tokens of cache fit after weights, and fails if the new release cannot hold the concurrency the fleet is sized for.
- Compatibility matrix checks that the container's CUDA runtime is supported by the node driver, that the engine version supports the model architecture and quantization format, and that the GPU SKU supports the required data types.
- Node burn-in runs a stress workload on upgraded nodes before they rejoin the serving pool, so a bad driver or firmware combination fails on a drained node rather than under traffic.
Why same output is the wrong test
A tempting golden-set rule is bitwise equality: same prompts, greedy decoding, identical outputs. It fails on changes that should pass. Floating-point reductions are not associative, so a kernel that sums in a different order produces slightly different logits, and when two candidate tokens are nearly tied the argmax flips and the continuation diverges. A September 2025 post from Thinking Machines traced most run-to-run variation in served LLMs to kernels whose results depend on batch size, which itself depends on load. The practical consequence is that even an unchanged release is not bitwise reproducible under a production server unless it uses batch-invariant kernels.
So golden-set checks use tolerant rules. Compare the position of first divergence rather than whole strings, and require most prompts to agree for a reasonable prefix. Compare per-token log-probabilities of the reference continuation under the new release, which is far more stable than comparing sampled text. Run the old release through the same harness twice to measure its own variation, and set thresholds above that floor. A template or tokenizer error produces divergence at the first token on most prompts; a benign kernel change produces late, scattered divergence. The two look nothing alike.
Capacity: the gate people forget
Numerics and platform changes move memory, and memory is concurrency. For a model with the Llama-3-70B shape (80 layers, 8 key-value heads, head dimension 128), the KV cache costs 2 * 80 * 8 * 128 * 2 = 327,680 bytes per token in bf16. On four 80 GB GPUs with 90 percent of memory available to the engine, bf16 weights of about 70.6 billion parameters leave room for roughly 448,000 cached tokens; FP8 weights leave roughly 663,000, ignoring activation workspace. FP8 KV cache would halve the bytes per token and double each figure. Activation memory and fragmentation reduce these, so the gate uses figures measured from the engine's own startup log, not the arithmetic; the arithmetic tells you which direction to expect. Sizing details are in KV cache sizing.
The same check catches the reverse problem. A change from tensor-parallel degree four to two halves the GPUs per replica and may leave too little cache for the long-context traffic the fleet serves, even though the model loads and the golden set passes.
Rollback is a capacity plan
Rollback is only real if the previous release can serve traffic within your recovery objective. For a 70B model, bf16 weights are about 141 GB. At a line rate of 10 gigabits per second, pulling them takes about 113 seconds; at 25 gigabits, about 45 seconds; and real registries and object stores rarely deliver line rate to many nodes at once. Add engine startup, weight loading into GPU memory and graph capture, and a cold rollback is easily ten minutes or more. The fix is to keep the N-1 weights staged on local NVMe on every node, and for critical services to keep a small pool of N-1 replicas warm until the new release has baked.
Some changes are one-way doors and need explicit sign-off. A chat template change that alters stored conversation formats, adapters trained against a specific base digest, fine-tuning data collected from the new release's output format, and client code that parses a new tool-call schema all make rollback incomplete. The classifier can flag interface-class changes for this review, but humans must decide whether rollback is still possible.
Process rules and the audit trail
Process rules that pay for themselves: one risk class per release; platform changes on their own schedule, rolled node pool by node pool with drains, never bundled with model changes; a change freeze around known traffic peaks; and a release calendar so two teams do not ship a template change and a driver upgrade in the same hour. Every gate result, approval and promotion writes to an audit trail keyed by release id, so an incident responder can ask what changed in the last 24 hours and get a list of manifest diffs with evidence, not a chat transcript. Topology choices that change how releases roll out are discussed in LLM deployment patterns.
Failure modes
- Tags instead of digests. A re-pushed tag changes production without a release. Pin everything by content hash.
- Behaviour changes outside the pipeline. System prompts edited in an admin console bypass every gate. Put them in the manifest.
- Bitwise golden tests. They fail benign changes, teams learn to override them, and then they miss real breakage. Use tolerant, calibrated rules.
- Bundled releases. Two classes in one release make a regression undiagnosable. Split them.
- Rollback never rehearsed. Weights garbage-collected from nodes or a registry that throttles under fleet-wide pulls turn rollback into an outage. Rehearse it on a schedule.
What to do next
- Write the release manifest for your current production state and pin every field by digest.
- Make replicas report their effective manifest id and alert on drift.
- Encode the risk table and gates as code; treat unmapped keys as model changes.
- Measure your golden set's own variation by running the current release twice, and set thresholds above it.
- Add a KV-capacity gate fed by the engine's startup log.
- Stage N-1 weights on local disk and rehearse a rollback, timing every step.
- Move platform upgrades onto a separate, node-pool-by-node-pool schedule.
- Hand passing releases to the canary process and record the outcome against the release id.