Ask a team running an LLM in production what exactly is serving traffic and you usually get a model name and an image tag. Neither answers the questions that come up in an incident: which tokenizer revision and chat template are loaded, which LoRA adapters, which quantisation kernels, which CUDA, cuDNN and NCCL builds, which driver and GPU firmware. Each of those can change outputs, latency or security exposure, and each can change without the model name changing.

An LLM bill of materials (LLM-BoM) is the answer written down: a machine-readable inventory of every component a served model depends on, from the system prompt down to the GPU, with versions and content hashes. This article covers what belongs in it, how to capture it from the running process rather than from build files, how to emit it as CycloneDX, how to diff two deployments, and how to query a fleet. Formats and admission checks for third-party models are covered in AIBOM, in depth; training-side run manifests in LLM training reproducibility.

Why a package list is not enough

A conventional software bill of materials lists packages. For an LLM service that misses most of what determines behaviour. Model weights are data, not packages. The chat template is a string inside a tokenizer config. The attention and matrix-multiply kernels that actually run depend on the GPU's compute capability and on which optional libraries were importable at start-up. The driver comes from the host, not the container. A requirements file records intent; the BoM must record what was loaded.

That gives the LLM-BoM three jobs. It explains change: when quality or latency moves, diff today's BoM with last week's. It answers exposure questions: when a driver, library or model advisory lands, list the deployments affected in seconds. And it supports accountability: a release can be traced to exact artefacts and owners.

Owners matter as much as versions. The five layers below are usually changed by different teams: product owns the system prompt, an ML team owns weights and adapters, a platform team owns the engine and container, and an infrastructure team owns drivers and node images. Record an owning team for every component, because the most useful output of a diff is not the changed version but the name of the team that can explain it. The same field routes advisories: a driver bulletin goes to infrastructure, a retracted adapter to the ML team. Without owners, a fleet query returns a list of affected deployments and nobody who is responsible for fixing them, and the list ages in a ticket queue while the exposure stays open.

The five layers of a served model

Think of a served model as five layers. Every layer can change output tokens, not just speed, which is why all five belong in one document.

LayerComponentsHow it changes behaviourHow to identify it
Promptsystem prompt, guard or classifier models, stop sequencesdirectly changes answers and refusalscontent hash, guard model hash
Model artefactsweights, tokenizer, chat template, generation defaults, adaptersdifferent tokens, formats, sampling defaultsSHA-256 of each file, repo revision
Engine and kernelsinference engine, attention backend, GEMM and quantisation kernels, KV-cache dtypenumerics, batching effects, precisionpackage versions, resolved engine config
CUDA user spaceCUDA runtime, cuBLAS, cuDNN, NCCLkernel selection and reduction orderversions reported by the framework
Driver and hardwaredriver, GPU model, VBIOS, compute capability, countavailable kernels, rare numeric and stability differencesNVML queries

Two details are easy to forget. The chat template should be hashed on its own, because it is the component most often changed by accident, and it can live in two places: a field in the tokenizer config or a separate chat_template.jinja file, which takes precedence when both exist. Hash the template the loaded tokenizer or engine actually uses, not either file. And the engine's resolved configuration (the effective data type, maximum context, KV-cache precision, quantisation method) matters more than its command line, because engines fill in defaults that differ between versions.

Capture, store, query

The collector runs inside the serving process during warm-up, after the model is loaded and before the node reports ready. It reads versions from the loaded modules, queries the driver through NVML, hashes the artefacts it actually opened, emits a CycloneDX document and stores it keyed by deployment id. Everything downstream, the diff and the fleet queries, works from the store.

Inference nodePrompt layersystem prompt, guard modelsModel artefactsweights, tokenizer, adaptersEngine + kernelsengine, attention, GEMM libsCUDA user spaceruntime, cuDNN, NCCLDriver + GPUdriver, VBIOS, SKUBoM collectorruns at warm-upintrospectCycloneDX JSONper deploymentBoM storekeyed by deploy idDiffwhy did it change?Fleet querywho runs X?The BoM describes what is loaded and running, read from the process, not what a build file asked for.
Capture the BoM from the running process at warm-up, store it per deployment, then diff and query.

Capturing the stack from a running node

The collector below uses only introspection that exists in the standard tools: PyTorch's version attributes, importlib.metadata for installed packages, NVML through the nvidia-ml-py bindings for the driver and each GPU, and SHA-256 over the files the server loaded. Pass it the tokenizer object and the resolved engine configuration from your server rather than re-deriving it. Hashing multi-gigabyte weights takes time, so cache digests keyed by path, size and modification time, or take them from a signed manifest produced when the model was published, as described in model signing and provenance.

import hashlib, importlib.metadata as md, json, platform
from pathlib import Path
import pynvml, torch

def sha256(path, chunk=1 << 24):
    h = hashlib.sha256()
    with open(path, "rb") as f:
        while block := f.read(chunk):
            h.update(block)
    return h.hexdigest()

def pkg(name):
    try:
        return md.version(name)
    except md.PackageNotFoundError:
        return None

def collect(model_dir, tokenizer, adapters, system_prompt, engine_config):
    """tokenizer: the object the server actually loaded (template may be str or dict)."""
    model_dir = Path(model_dir)
    template = json.dumps(tokenizer.chat_template, sort_keys=True)
    pynvml.nvmlInit()
    gpus = []
    for i in range(pynvml.nvmlDeviceGetCount()):
        h = pynvml.nvmlDeviceGetHandleByIndex(i)
        gpus.append({"name": pynvml.nvmlDeviceGetName(h),
                     "vbios": pynvml.nvmlDeviceGetVbiosVersion(h),
                     "cc": "%d.%d" % pynvml.nvmlDeviceGetCudaComputeCapability(h)})
    return {
        "prompt": {"system_prompt_sha256": hashlib.sha256(system_prompt.encode()).hexdigest()},
        "model": {
            "weights": {p.name: sha256(p) for p in sorted(model_dir.glob("*.safetensors"))},
            "tokenizer_json": sha256(model_dir / "tokenizer.json"),
            "chat_template_sha256": hashlib.sha256(template.encode()).hexdigest(),
            "adapters": {a: sha256(Path(a) / "adapter_model.safetensors") for a in adapters},
        },
        "engine": {"config": engine_config,
                   "packages": {n: pkg(n) for n in
                                ("torch", "transformers", "vllm", "flash-attn", "triton")}},
        "cuda": {"torch_cuda": torch.version.cuda,
                 "cudnn": torch.backends.cudnn.version(),
                 "nccl": ".".join(map(str, torch.cuda.nccl.version()))},
        "host": {"driver": pynvml.nvmlSystemGetDriverVersion(),
                 "driver_cuda": pynvml.nvmlSystemGetCudaDriverVersion(),
                 "kernel": platform.release(), "gpus": gpus},
    }

Emitting CycloneDX

Emit a standard format so security and procurement tools can read it. CycloneDX 1.6 has component types that fit each layer: machine-learning-model for weights and adapters, data for the chat template and system prompt, library for packages and CUDA libraries, device-driver for the driver and device for GPUs. Anything without a standard field, such as the VBIOS version or compute capability, goes into name and value properties. A production version would add model cards and dependency relationships; this shape is enough for diffs and queries.

import uuid, datetime

def to_cyclonedx(bom, deploy_id):
    comps = [{"type": "machine-learning-model", "name": f"weights:{shard}",
              "hashes": [{"alg": "SHA-256", "content": d}]}      # one component per shard
             for shard, d in bom["model"]["weights"].items()]
    comps += [{"type": "data", "name": "chat-template",
              "hashes": [{"alg": "SHA-256", "content": bom["model"]["chat_template_sha256"]}]},
             {"type": "data", "name": "system-prompt",
              "hashes": [{"alg": "SHA-256", "content": bom["prompt"]["system_prompt_sha256"]}]}]
    comps += [{"type": "machine-learning-model", "name": f"adapter:{a}",
               "hashes": [{"alg": "SHA-256", "content": d}]}
              for a, d in bom["model"]["adapters"].items()]
    comps += [{"type": "library", "name": n, "version": v,
               "purl": f"pkg:pypi/{n}@{v}"}
              for n, v in bom["engine"]["packages"].items() if v]
    comps += [{"type": "library", "name": k, "version": str(v)} for k, v in bom["cuda"].items()]
    comps.append({"type": "device-driver", "name": "nvidia-driver",
                  "version": bom["host"]["driver"]})
    comps += [{"type": "device", "name": g["name"],
               "properties": [{"name": "vbios", "value": g["vbios"]},
                              {"name": "compute_capability", "value": g["cc"]}]}
              for g in bom["host"]["gpus"]]
    return {"bomFormat": "CycloneDX", "specVersion": "1.6",
            "serialNumber": f"urn:uuid:{uuid.uuid4()}",
            "metadata": {"timestamp": datetime.datetime.now(datetime.timezone.utc).isoformat(),
                         "properties": [{"name": "deploy_id", "value": deploy_id}]},
            "components": comps}

Diffing two deployments

The most frequent use is answering “what changed?” Index each document by component type and name, compare identities, and print the differences. The flag on each row is deliberately conservative: nearly every component type can change output tokens, so the diff tells you where to look, and an evaluation run against a fixed prompt set tells you whether it mattered. For why identical code on different kernels gives different numbers, see ML reproducibility on GPUs.

NUMERIC = {"machine-learning-model", "data", "library", "device-driver", "device"}

def index(cdx):
    out = {}
    for c in cdx["components"]:
        ident = c.get("version") or ",".join(h["content"][:12] for h in c.get("hashes", []))
        props = ",".join(f'{p["name"]}={p["value"]}' for p in c.get("properties", []))
        out[(c["type"], c["name"])] = ident + (f" [{props}]" if props else "")
    return out

def diff(old, new):
    a, b = index(old), index(new)
    rows = []
    for key in sorted(set(a) | set(b)):
        if a.get(key) != b.get(key):
            rows.append({"component": f"{key[0]}:{key[1]}", "old": a.get(key),
                         "new": b.get(key), "may_change_outputs": key[0] in NUMERIC})
    return rows

Fleet queries

Once every deployment writes its BoM to a store, fleet questions become queries. Typical ones: which deployments run a driver branch named in a new vendor security bulletin; which serve a given base model hash, so a licence or safety notice can be acted on; which still run an adapter that was retracted; which mix GPU compute capabilities within one replica set, so their numerics differ; which use a chat template that no longer matches the model publisher's. Load the CycloneDX documents into any store that can filter JSON, such as a document database or a warehouse table with one row per component, and keep at least the history of every deployment that served traffic.

Two policies make the data trustworthy. A node that fails to produce a BoM does not become ready. And the deployment record links to the BoM, so the BoM cannot quietly drift from what the scheduler thinks is running. Container images give you part of this for free, since the image digest pins user-space libraries, but the driver, GPU and mounted weights live outside the image; see GPU containers for that boundary.

Worked example: the Tuesday quality drop

On a Tuesday an internal assistant's evaluation score drops four points and users report answers that ignore the system prompt's formatting rules. The model name, image tag and weights are unchanged. The on-call engineer diffs the BoM of the current deployment with Monday's. Three rows differ: the node pool's driver moved to a newer release within the same branch, the engine's resolved configuration now has a different maximum context, and the chat-template hash changed.

The template is the obvious suspect. The tokenizer files had been re-downloaded from the model repository by a cache-refresh job, and the publisher had updated the template to handle tool calls, which changed how the system message was placed. Pinning the previous repository revision restores the score on the fixed prompt set; the driver change, tested separately on a canary, makes no measurable difference. Without the BoM the team would have started by rolling back the driver, the change with the most operational risk and, here, no effect. The incident also produced two lasting fixes: weights and tokenizer are now loaded only by pinned revision, and a template hash change blocks a deploy until the evaluation set passes. The details are illustrative; template drift of this kind is a common real cause.

Failure modes

  • BoM from build files. Generated from requirements and Dockerfiles, it misses the host driver, mounted weights and whichever optional kernels actually imported.
  • Names without hashes. “llama-style-8b” and “latest” identify nothing. Every artefact needs a content hash or a pinned revision.
  • Template not tracked. The chat template changes inside an unchanged tokenizer repository, and nobody sees it.
  • Stale BoMs. Captured at build, not at warm-up, so hot-swapped adapters and changed config never appear.
  • Slow capture blocking start-up. Hashing hundreds of gigabytes on every boot delays readiness; cache digests or use signed manifests.
  • Secrets in the BoM. Raw system prompts or engine configs with credentials get copied into a widely readable store. Store hashes, not contents.

Trade-offs

ChoiceGainCost
Runtime capture vs build-timerecords what really loadedcode in the serving path; start-up time
Full file hashes vs revision idsdetects silent artefact changeshashing time for large weights
CycloneDX vs custom JSONtool interoperabilitysome fields only fit as properties
Per-deployment vs per-node BoMsmaller storemisses heterogeneous nodes in one pool
Hashing prompts vs storing themno secrets in the storediff shows that, not how, a prompt changed

What to do next

  1. List the five layers for one service and note where each version currently lives.
  2. Add a collector that runs at warm-up and blocks readiness if it fails.
  3. Hash weights, tokenizer, chat template, adapters and system prompt; cache the digests.
  4. Emit CycloneDX 1.6 and store it keyed by deployment id, keeping history.
  5. Add the BoM diff to your incident runbook as the first step for quality or latency changes.
  6. Write the fleet queries you will need for driver, library and model advisories, and test them.
  7. Pin model and tokenizer revisions and gate deploys on template-hash changes.
Key takeaway: A served LLM is five layers deep, from system prompt to GPU firmware, and each layer can change its answers. Capture the bill of materials from the running process at warm-up with content hashes, emit it as CycloneDX, store it per deployment, and make the BoM diff the first step whenever quality, latency or exposure changes.