A container that serves a large language model is a supply chain of its own. It stacks a Linux base, the CUDA runtime and math libraries, a deep tree of Python wheels, an inference server, your application code and, somewhere, tens of gigabytes of model weights. Each layer comes from a different publisher, changes on a different schedule and can carry its own attack. A compromised wheel runs with GPU access. A weights file in the wrong format can execute code when it loads. A server started with remote code enabled will run Python fetched from a model repository.
The general practices of software supply chain security still apply: pin inputs, build in CI, generate an SBOM, sign, verify at admission. They are covered in software supply chain security architecture and SBOMs and supply chain security. This article is about what is different for LLM serving: the size and shape of the images, weights as a separate artifact that needs its own signature, the formats that turn data into code, and the GPU driver boundary that no image scan can see. It ends with a worked rollout and a checklist.
Why LLM serving images are different
Four properties make LLM serving images unusual.
- Size. The official vLLM OpenAI-compatible image is over 10 GB before any weights, mostly PyTorch and CUDA libraries built for several GPU architectures. Large images are slow to scan and tempting to rebuild rarely, so patches lag.
- Native code everywhere. CUDA kernels, NCCL, cuDNN and compiled extensions in wheels are binaries that most scanners identify only by package metadata. If a wheel vendors a library without declaring it, the SBOM misses it.
- Weights are data that can act like code. PyTorch's older checkpoint format is a Python pickle, and unpickling can run arbitrary code. Hugging Face's
trust_remote_codeoption loads and runs Python modelling files shipped alongside the weights. - The host driver is inside the trust boundary. The NVIDIA Container Toolkit mounts the host's driver libraries into the container at start. They are not in your image, not in your SBOM and not covered by your signature, yet every kernel runs through them.
Anatomy of a serving image
Look at the image as layers with different owners and change rates. The table guides where to spend effort.
| Layer | Typical source | Size | Changes | Main risk |
|---|---|---|---|---|
| OS base | Distro or minimal base | Tens of MB | Weekly patches | Known CVEs |
| CUDA runtime, cuDNN, NCCL | Vendor base image | Several GB | Per CUDA release | Opaque binaries, slow patching |
| PyTorch and wheels | Package index | Several GB | Per release | Typosquats, compromised releases, undeclared vendored code |
| Inference server | vLLM, TGI, Triton | Hundreds of MB | Every few weeks | Unsafe flags, network listeners |
| Application code | Your repository | Small | Daily | Ordinary bugs, secrets |
| Model weights | Model hub or internal | 1 to 400+ GB | Per model release | Tampering, pickle execution, licence |
| Host GPU driver | Node image | Outside the image | Per node update | Invisible to image scanning |
Two conclusions follow. Weights should not live in the same artifact as code, because they differ in size, cadence and verification. And the node image, including the driver version, needs the same inventory and patching discipline as the container image, because it is part of what runs.
Weights as a separate artifact
There are three ways to get weights into a serving pod.
| Approach | Strength | Weakness |
|---|---|---|
| Bake into the image | One digest covers code and weights | Huge images, a rebuild for every patch, slow pulls on every node |
| Download at start from object storage | Small image, independent model releases | Needs its own verification; startup depends on storage |
| Mount an OCI artifact as an image volume | Registry tooling, digests and signatures for weights | Needs a recent Kubernetes and runtime |
Kubernetes image volumes let a pod mount the contents of an OCI artifact read-only, and the design document names model weights beside a model server as a target use. The feature arrived as alpha in 1.31 and became beta in 1.33; later releases are reported to make it generally available. Check your cluster version and container runtime before relying on it.
Whichever path you choose, the server should load only a verified copy. Choose the safetensors format, which stores raw tensors and metadata with no executable content, and refuse pickle-based files in production. Since PyTorch 2.6, torch.load defaults to weights_only=True, which restricts unpickling, but a format that cannot carry code is stronger than a loader setting someone can flip. Run the inference server without remote code; in vLLM that means never passing --trust-remote-code for models you have not reviewed and vendored.
Python dependencies: the widest door
The Python dependency tree is the widest door into a serving image. A typical vLLM or Transformers install resolves well over a hundred packages, several of them compiled extensions, and the machine learning ecosystem has seen typosquatted packages, hijacked maintainer accounts and malicious nightly builds. Three habits close most of it.
First, resolve once and lock with hashes. A lock file generated by pip-tools or uv records the exact version and hash of every wheel for your target platform; --require-hashes makes the install fail if any download differs. Regenerate the lock in a reviewed pull request, never inside the image build.
Second, control where packages come from. If you publish internal packages, serve them from a private index that also proxies the public one, and point --index-url at it alone. Adding the public index with --extra-index-url lets a public package with the same name and a higher version win, which is the classic dependency confusion attack.
Third, keep build and runtime separate. Compilers, headers and build tools belong in a builder stage; the runtime stage copies only installed packages. Fewer binaries in the final image means fewer findings to triage and fewer tools an attacker can use after a compromise. Accelerator wheels are large and often built against a specific CUDA version, so record the CUDA version in the image labels and the SBOM, and test the full matrix when either changes.
Building a signed image
The build turns pinned inputs into a signed image with evidence attached. Pin the base by digest, not tag, so a republished tag cannot change your image. Lock Python dependencies with hashes so a replaced wheel fails the install.
# syntax=docker/dockerfile:1
# pin bases by digest, not tag
FROM <cuda-devel-base>@sha256:<digest> AS build
COPY requirements.lock /tmp/
RUN python3 -m venv /opt/venv && \
/opt/venv/bin/pip install --no-cache-dir --require-hashes -r /tmp/requirements.lock
# runtime stage: no compilers in the final image
FROM <cuda-runtime-base>@sha256:<digest>
RUN useradd -u 10001 -m serve
COPY --from=build /opt/venv /opt/venv
COPY app/ /app/
ENV PATH=/opt/venv/bin:$PATH
USER 10001
# never fetch from a model hub at runtime
ENV HF_HUB_OFFLINE=1
ENTRYPOINT ["python", "-m", "app.serve"]In CI, after the build: generate an SBOM from the image, scan it, attach both as attestations, and sign the digest with a keyless identity tied to the workflow.
IMG=registry.example.com/llm/serve@${DIGEST}
syft "$IMG" -o spdx-json > sbom.spdx.json
grype "$IMG" --fail-on high
cosign attest --yes --predicate sbom.spdx.json --type spdxjson "$IMG"
cosign sign --yes "$IMG"
# anyone can verify later; identity and issuer must both match
cosign verify "$IMG" \
--certificate-identity "https://github.com/example/llm-serve/.github/workflows/build.yml@refs/heads/main" \
--certificate-oidc-issuer "https://token.actions.githubusercontent.com"Scanning a multi-gigabyte CUDA image produces long lists of findings in libraries your server never calls. Triage with VEX statements rather than ignoring the scanner wholesale, and rebuild on a schedule, weekly at least, so base patches land even when your code has not changed.
Signing and verifying weights
Weights need a signature of their own. The OpenSSF model signing project provides a library and CLI, installed with pip install model-signing, that hashes every file in a model directory and signs the resulting manifest as a Sigstore bundle. Verification fails if any file changes, is added or is removed.
# model release job (CI identity signs)
model_signing sign ./llama-8b-instruct --signature model.sig
# init container in the serving pod, before the server starts
model_signing verify ./llama-8b-instruct \
--signature model.sig \
--identity "$RELEASE_IDENTITY" \
--identity-provider "$OIDC_ISSUER"Run the verify step in an init container that writes to a volume the server mounts read-only. The server never sees unverified bytes, and a failure stops the pod before it accepts traffic. Record the weights digest in the deployment and in every response log, so an answer can be traced to the exact model files that produced it.
Admission and runtime controls
Admission control turns the signature into a rule. With Kyverno, an image verification policy can require that every pod in the serving namespace uses an image signed by your build workflow, and can rewrite tags to digests.
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: verify-llm-serving-images
spec:
validationFailureAction: Enforce # newer Kyverno versions set this per rule
rules:
- name: signed-by-build-workflow
match:
any:
- resources:
kinds: ["Pod"]
namespaces: ["llm-serving"]
verifyImages:
- imageReferences: ["registry.example.com/llm/*"]
mutateDigest: true
attestors:
- entries:
- keyless:
subject: "https://github.com/example/llm-serve/.github/workflows/build.yml@refs/heads/main"
issuer: "https://token.actions.githubusercontent.com"Pair this with runtime hardening from container security architecture: non-root user, read-only root filesystem, dropped capabilities and no privilege escalation. Inference servers rarely need outbound internet; deny egress by default so a compromised dependency cannot fetch a payload or send data out, and so a forgotten hub download fails loudly instead of silently pulling unverified weights.
Worked example: hardening a vLLM deployment
A platform team serves an 8B instruct model with vLLM on a GPU node pool. Before: a public image tag, weights pulled from a hub at startup, remote code enabled because a tutorial said so. Every pod start downloaded unverified files from the internet and ran whatever Python came with them.
After: the team builds its own serving image from a digest-pinned vendor base with hash-locked wheels, about 9 GB, signed by the CI workflow with an SBOM attestation. The weights are converted to safetensors once, reviewed, stored in an internal bucket and signed by a release job. Each pod runs an init container that downloads the weights, runs model_signing verify and writes to a shared volume; the server starts with HF_HUB_OFFLINE=1, no remote code and no egress. Kyverno rejects any image in the namespace not signed by the build workflow. The node pool's driver version is pinned in the node image and tracked with the same patch calendar. Cold start rises by the verification time, which is bounded by how fast the node can read and hash 16 GB from local disk; the team measures it once and accepts it. For serving details see the vLLM serving engine.
Failure modes
- Signed image, unsigned weights. Admission passes, and the server loads tampered weights from a bucket. Verify both artifacts.
- Tag drift. A base tag is republished with different contents. Pin digests and let the policy rewrite tags to digests.
- Pickle in production. A
.bincheckpoint slips through because conversion was skipped. Make the loader reject non-safetensors files. - Runtime hub download. A missing local file makes the server fetch from the internet. Set offline mode and deny egress.
- Scanner fatigue. Hundreds of findings in CUDA layers get blanket-ignored, including a real one. Use VEX and rebuild weekly.
- Driver blind spot. A vulnerable host driver is never patched because nobody inventories node images.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Build your own image | Known inputs, smaller attack surface | You own patching and CUDA compatibility |
| Vendor image by digest | Less work, vendor-tested | Opaque contents, their patch cadence |
| Weights separate from image | Independent releases, smaller images | A second artifact to sign and verify |
| Verify weights at every start | Tampering caught before serving | Longer cold starts |
What to do next
- Inventory every serving image, its base digest, its weights source and the node driver version.
- Pin bases by digest and lock Python dependencies with hashes.
- Generate an SBOM, scan and sign every image in CI with a keyless workflow identity.
- Convert weights to safetensors, sign them with model_signing, and verify in an init container.
- Remove remote code, set hub offline mode and deny egress in the serving namespace.
- Enforce signed images with an admission policy that rewrites tags to digests.
- Rebuild weekly, triage findings with VEX, and patch node images on the same calendar.
- Log the image digest and weights digest with every response.