Ask a team running an LLM in production what exactly is serving traffic and you usually get a model name and an image tag. Neither answers the questions that come up in an incident: which tokenizer revision and chat template are loaded, which LoRA adapters, which quantisation kernels, which CUDA, cuDNN and NCCL builds, which driver and GPU firmware. Each of those can change outputs, latency or security exposure, and each can change without the model name changing.
An LLM bill of materials (LLM-BoM) is the answer written down: a machine-readable inventory of every component a served model depends on, from the system prompt down to the GPU, with versions and content hashes. This article covers what belongs in it, how to capture it from the running process rather than from build files, how to emit it as CycloneDX, how to diff two deployments, and how to query a fleet. Formats and admission checks for third-party models are covered in AIBOM, in depth; training-side run manifests in LLM training reproducibility.
Why a package list is not enough
A conventional software bill of materials lists packages. For an LLM service that misses most of what determines behaviour. Model weights are data, not packages. The chat template is a string inside a tokenizer config. The attention and matrix-multiply kernels that actually run depend on the GPU's compute capability and on which optional libraries were importable at start-up. The driver comes from the host, not the container. A requirements file records intent; the BoM must record what was loaded.
That gives the LLM-BoM three jobs. It explains change: when quality or latency moves, diff today's BoM with last week's. It answers exposure questions: when a driver, library or model advisory lands, list the deployments affected in seconds. And it supports accountability: a release can be traced to exact artefacts and owners.
Owners matter as much as versions. The five layers below are usually changed by different teams: product owns the system prompt, an ML team owns weights and adapters, a platform team owns the engine and container, and an infrastructure team owns drivers and node images. Record an owning team for every component, because the most useful output of a diff is not the changed version but the name of the team that can explain it. The same field routes advisories: a driver bulletin goes to infrastructure, a retracted adapter to the ML team. Without owners, a fleet query returns a list of affected deployments and nobody who is responsible for fixing them, and the list ages in a ticket queue while the exposure stays open.
The five layers of a served model
Think of a served model as five layers. Every layer can change output tokens, not just speed, which is why all five belong in one document.
| Layer | Components | How it changes behaviour | How to identify it |
|---|---|---|---|
| Prompt | system prompt, guard or classifier models, stop sequences | directly changes answers and refusals | content hash, guard model hash |
| Model artefacts | weights, tokenizer, chat template, generation defaults, adapters | different tokens, formats, sampling defaults | SHA-256 of each file, repo revision |
| Engine and kernels | inference engine, attention backend, GEMM and quantisation kernels, KV-cache dtype | numerics, batching effects, precision | package versions, resolved engine config |
| CUDA user space | CUDA runtime, cuBLAS, cuDNN, NCCL | kernel selection and reduction order | versions reported by the framework |
| Driver and hardware | driver, GPU model, VBIOS, compute capability, count | available kernels, rare numeric and stability differences | NVML queries |
Two details are easy to forget. The chat template should be hashed on its own, because it is the component most often changed by accident, and it can live in two places: a field in the tokenizer config or a separate chat_template.jinja file, which takes precedence when both exist. Hash the template the loaded tokenizer or engine actually uses, not either file. And the engine's resolved configuration (the effective data type, maximum context, KV-cache precision, quantisation method) matters more than its command line, because engines fill in defaults that differ between versions.
Capture, store, query
The collector runs inside the serving process during warm-up, after the model is loaded and before the node reports ready. It reads versions from the loaded modules, queries the driver through NVML, hashes the artefacts it actually opened, emits a CycloneDX document and stores it keyed by deployment id. Everything downstream, the diff and the fleet queries, works from the store.
Capturing the stack from a running node
The collector below uses only introspection that exists in the standard tools: PyTorch's version attributes, importlib.metadata for installed packages, NVML through the nvidia-ml-py bindings for the driver and each GPU, and SHA-256 over the files the server loaded. Pass it the tokenizer object and the resolved engine configuration from your server rather than re-deriving it. Hashing multi-gigabyte weights takes time, so cache digests keyed by path, size and modification time, or take them from a signed manifest produced when the model was published, as described in model signing and provenance.
import hashlib, importlib.metadata as md, json, platform
from pathlib import Path
import pynvml, torch
def sha256(path, chunk=1 << 24):
h = hashlib.sha256()
with open(path, "rb") as f:
while block := f.read(chunk):
h.update(block)
return h.hexdigest()
def pkg(name):
try:
return md.version(name)
except md.PackageNotFoundError:
return None
def collect(model_dir, tokenizer, adapters, system_prompt, engine_config):
"""tokenizer: the object the server actually loaded (template may be str or dict)."""
model_dir = Path(model_dir)
template = json.dumps(tokenizer.chat_template, sort_keys=True)
pynvml.nvmlInit()
gpus = []
for i in range(pynvml.nvmlDeviceGetCount()):
h = pynvml.nvmlDeviceGetHandleByIndex(i)
gpus.append({"name": pynvml.nvmlDeviceGetName(h),
"vbios": pynvml.nvmlDeviceGetVbiosVersion(h),
"cc": "%d.%d" % pynvml.nvmlDeviceGetCudaComputeCapability(h)})
return {
"prompt": {"system_prompt_sha256": hashlib.sha256(system_prompt.encode()).hexdigest()},
"model": {
"weights": {p.name: sha256(p) for p in sorted(model_dir.glob("*.safetensors"))},
"tokenizer_json": sha256(model_dir / "tokenizer.json"),
"chat_template_sha256": hashlib.sha256(template.encode()).hexdigest(),
"adapters": {a: sha256(Path(a) / "adapter_model.safetensors") for a in adapters},
},
"engine": {"config": engine_config,
"packages": {n: pkg(n) for n in
("torch", "transformers", "vllm", "flash-attn", "triton")}},
"cuda": {"torch_cuda": torch.version.cuda,
"cudnn": torch.backends.cudnn.version(),
"nccl": ".".join(map(str, torch.cuda.nccl.version()))},
"host": {"driver": pynvml.nvmlSystemGetDriverVersion(),
"driver_cuda": pynvml.nvmlSystemGetCudaDriverVersion(),
"kernel": platform.release(), "gpus": gpus},
}
Emitting CycloneDX
Emit a standard format so security and procurement tools can read it. CycloneDX 1.6 has component types that fit each layer: machine-learning-model for weights and adapters, data for the chat template and system prompt, library for packages and CUDA libraries, device-driver for the driver and device for GPUs. Anything without a standard field, such as the VBIOS version or compute capability, goes into name and value properties. A production version would add model cards and dependency relationships; this shape is enough for diffs and queries.
import uuid, datetime
def to_cyclonedx(bom, deploy_id):
comps = [{"type": "machine-learning-model", "name": f"weights:{shard}",
"hashes": [{"alg": "SHA-256", "content": d}]} # one component per shard
for shard, d in bom["model"]["weights"].items()]
comps += [{"type": "data", "name": "chat-template",
"hashes": [{"alg": "SHA-256", "content": bom["model"]["chat_template_sha256"]}]},
{"type": "data", "name": "system-prompt",
"hashes": [{"alg": "SHA-256", "content": bom["prompt"]["system_prompt_sha256"]}]}]
comps += [{"type": "machine-learning-model", "name": f"adapter:{a}",
"hashes": [{"alg": "SHA-256", "content": d}]}
for a, d in bom["model"]["adapters"].items()]
comps += [{"type": "library", "name": n, "version": v,
"purl": f"pkg:pypi/{n}@{v}"}
for n, v in bom["engine"]["packages"].items() if v]
comps += [{"type": "library", "name": k, "version": str(v)} for k, v in bom["cuda"].items()]
comps.append({"type": "device-driver", "name": "nvidia-driver",
"version": bom["host"]["driver"]})
comps += [{"type": "device", "name": g["name"],
"properties": [{"name": "vbios", "value": g["vbios"]},
{"name": "compute_capability", "value": g["cc"]}]}
for g in bom["host"]["gpus"]]
return {"bomFormat": "CycloneDX", "specVersion": "1.6",
"serialNumber": f"urn:uuid:{uuid.uuid4()}",
"metadata": {"timestamp": datetime.datetime.now(datetime.timezone.utc).isoformat(),
"properties": [{"name": "deploy_id", "value": deploy_id}]},
"components": comps}
Diffing two deployments
The most frequent use is answering “what changed?” Index each document by component type and name, compare identities, and print the differences. The flag on each row is deliberately conservative: nearly every component type can change output tokens, so the diff tells you where to look, and an evaluation run against a fixed prompt set tells you whether it mattered. For why identical code on different kernels gives different numbers, see ML reproducibility on GPUs.
NUMERIC = {"machine-learning-model", "data", "library", "device-driver", "device"}
def index(cdx):
out = {}
for c in cdx["components"]:
ident = c.get("version") or ",".join(h["content"][:12] for h in c.get("hashes", []))
props = ",".join(f'{p["name"]}={p["value"]}' for p in c.get("properties", []))
out[(c["type"], c["name"])] = ident + (f" [{props}]" if props else "")
return out
def diff(old, new):
a, b = index(old), index(new)
rows = []
for key in sorted(set(a) | set(b)):
if a.get(key) != b.get(key):
rows.append({"component": f"{key[0]}:{key[1]}", "old": a.get(key),
"new": b.get(key), "may_change_outputs": key[0] in NUMERIC})
return rows
Fleet queries
Once every deployment writes its BoM to a store, fleet questions become queries. Typical ones: which deployments run a driver branch named in a new vendor security bulletin; which serve a given base model hash, so a licence or safety notice can be acted on; which still run an adapter that was retracted; which mix GPU compute capabilities within one replica set, so their numerics differ; which use a chat template that no longer matches the model publisher's. Load the CycloneDX documents into any store that can filter JSON, such as a document database or a warehouse table with one row per component, and keep at least the history of every deployment that served traffic.
Two policies make the data trustworthy. A node that fails to produce a BoM does not become ready. And the deployment record links to the BoM, so the BoM cannot quietly drift from what the scheduler thinks is running. Container images give you part of this for free, since the image digest pins user-space libraries, but the driver, GPU and mounted weights live outside the image; see GPU containers for that boundary.
Worked example: the Tuesday quality drop
On a Tuesday an internal assistant's evaluation score drops four points and users report answers that ignore the system prompt's formatting rules. The model name, image tag and weights are unchanged. The on-call engineer diffs the BoM of the current deployment with Monday's. Three rows differ: the node pool's driver moved to a newer release within the same branch, the engine's resolved configuration now has a different maximum context, and the chat-template hash changed.
The template is the obvious suspect. The tokenizer files had been re-downloaded from the model repository by a cache-refresh job, and the publisher had updated the template to handle tool calls, which changed how the system message was placed. Pinning the previous repository revision restores the score on the fixed prompt set; the driver change, tested separately on a canary, makes no measurable difference. Without the BoM the team would have started by rolling back the driver, the change with the most operational risk and, here, no effect. The incident also produced two lasting fixes: weights and tokenizer are now loaded only by pinned revision, and a template hash change blocks a deploy until the evaluation set passes. The details are illustrative; template drift of this kind is a common real cause.
Failure modes
- BoM from build files. Generated from requirements and Dockerfiles, it misses the host driver, mounted weights and whichever optional kernels actually imported.
- Names without hashes. “llama-style-8b” and “latest” identify nothing. Every artefact needs a content hash or a pinned revision.
- Template not tracked. The chat template changes inside an unchanged tokenizer repository, and nobody sees it.
- Stale BoMs. Captured at build, not at warm-up, so hot-swapped adapters and changed config never appear.
- Slow capture blocking start-up. Hashing hundreds of gigabytes on every boot delays readiness; cache digests or use signed manifests.
- Secrets in the BoM. Raw system prompts or engine configs with credentials get copied into a widely readable store. Store hashes, not contents.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Runtime capture vs build-time | records what really loaded | code in the serving path; start-up time |
| Full file hashes vs revision ids | detects silent artefact changes | hashing time for large weights |
| CycloneDX vs custom JSON | tool interoperability | some fields only fit as properties |
| Per-deployment vs per-node BoM | smaller store | misses heterogeneous nodes in one pool |
| Hashing prompts vs storing them | no secrets in the store | diff shows that, not how, a prompt changed |
What to do next
- List the five layers for one service and note where each version currently lives.
- Add a collector that runs at warm-up and blocks readiness if it fails.
- Hash weights, tokenizer, chat template, adapters and system prompt; cache the digests.
- Emit CycloneDX 1.6 and store it keyed by deployment id, keeping history.
- Add the BoM diff to your incident runbook as the first step for quality or latency changes.
- Write the fleet queries you will need for driver, library and model advisories, and test them.
- Pin model and tokenizer revisions and gate deploys on template-hash changes.