A container that serves a large language model is a supply chain of its own. It stacks a Linux base, the CUDA runtime and math libraries, a deep tree of Python wheels, an inference server, your application code and, somewhere, tens of gigabytes of model weights. Each layer comes from a different publisher, changes on a different schedule and can carry its own attack. A compromised wheel runs with GPU access. A weights file in the wrong format can execute code when it loads. A server started with remote code enabled will run Python fetched from a model repository.

The general practices of software supply chain security still apply: pin inputs, build in CI, generate an SBOM, sign, verify at admission. They are covered in software supply chain security architecture and SBOMs and supply chain security. This article is about what is different for LLM serving: the size and shape of the images, weights as a separate artifact that needs its own signature, the formats that turn data into code, and the GPU driver boundary that no image scan can see. It ends with a worked rollout and a checklist.

Why LLM serving images are different

Four properties make LLM serving images unusual.

  • Size. The official vLLM OpenAI-compatible image is over 10 GB before any weights, mostly PyTorch and CUDA libraries built for several GPU architectures. Large images are slow to scan and tempting to rebuild rarely, so patches lag.
  • Native code everywhere. CUDA kernels, NCCL, cuDNN and compiled extensions in wheels are binaries that most scanners identify only by package metadata. If a wheel vendors a library without declaring it, the SBOM misses it.
  • Weights are data that can act like code. PyTorch's older checkpoint format is a Python pickle, and unpickling can run arbitrary code. Hugging Face's trust_remote_code option loads and runs Python modelling files shipped alongside the weights.
  • The host driver is inside the trust boundary. The NVIDIA Container Toolkit mounts the host's driver libraries into the container at start. They are not in your image, not in your SBOM and not covered by your signature, yet every kernel runs through them.

Anatomy of a serving image

Look at the image as layers with different owners and change rates. The table guides where to spend effort.

LayerTypical sourceSizeChangesMain risk
OS baseDistro or minimal baseTens of MBWeekly patchesKnown CVEs
CUDA runtime, cuDNN, NCCLVendor base imageSeveral GBPer CUDA releaseOpaque binaries, slow patching
PyTorch and wheelsPackage indexSeveral GBPer releaseTyposquats, compromised releases, undeclared vendored code
Inference servervLLM, TGI, TritonHundreds of MBEvery few weeksUnsafe flags, network listeners
Application codeYour repositorySmallDailyOrdinary bugs, secrets
Model weightsModel hub or internal1 to 400+ GBPer model releaseTampering, pickle execution, licence
Host GPU driverNode imageOutside the imagePer node updateInvisible to image scanning

Two conclusions follow. Weights should not live in the same artifact as code, because they differ in size, cadence and verification. And the node image, including the driver version, needs the same inventory and patching discipline as the container image, because it is part of what runs.

Weights as a separate artifact

Base + CUDApinned by digestLocked wheelshashes requiredServer + app codeCI buildSBOM, scan, provenancecosign signRegistryimage@sha256 + attestationsWeightssafetensors onlyModel release jobhash manifest, signmodel_signingModel storeweights + signatureimageClusteradmission verifies image; init verifies weightsTwo artifacts, two signatures,both checked before serving
Two signed artifacts: the serving image is verified at admission, and the weights are verified by an init step before the server loads them.

There are three ways to get weights into a serving pod.

ApproachStrengthWeakness
Bake into the imageOne digest covers code and weightsHuge images, a rebuild for every patch, slow pulls on every node
Download at start from object storageSmall image, independent model releasesNeeds its own verification; startup depends on storage
Mount an OCI artifact as an image volumeRegistry tooling, digests and signatures for weightsNeeds a recent Kubernetes and runtime

Kubernetes image volumes let a pod mount the contents of an OCI artifact read-only, and the design document names model weights beside a model server as a target use. The feature arrived as alpha in 1.31 and became beta in 1.33; later releases are reported to make it generally available. Check your cluster version and container runtime before relying on it.

Whichever path you choose, the server should load only a verified copy. Choose the safetensors format, which stores raw tensors and metadata with no executable content, and refuse pickle-based files in production. Since PyTorch 2.6, torch.load defaults to weights_only=True, which restricts unpickling, but a format that cannot carry code is stronger than a loader setting someone can flip. Run the inference server without remote code; in vLLM that means never passing --trust-remote-code for models you have not reviewed and vendored.

Python dependencies: the widest door

The Python dependency tree is the widest door into a serving image. A typical vLLM or Transformers install resolves well over a hundred packages, several of them compiled extensions, and the machine learning ecosystem has seen typosquatted packages, hijacked maintainer accounts and malicious nightly builds. Three habits close most of it.

First, resolve once and lock with hashes. A lock file generated by pip-tools or uv records the exact version and hash of every wheel for your target platform; --require-hashes makes the install fail if any download differs. Regenerate the lock in a reviewed pull request, never inside the image build.

Second, control where packages come from. If you publish internal packages, serve them from a private index that also proxies the public one, and point --index-url at it alone. Adding the public index with --extra-index-url lets a public package with the same name and a higher version win, which is the classic dependency confusion attack.

Third, keep build and runtime separate. Compilers, headers and build tools belong in a builder stage; the runtime stage copies only installed packages. Fewer binaries in the final image means fewer findings to triage and fewer tools an attacker can use after a compromise. Accelerator wheels are large and often built against a specific CUDA version, so record the CUDA version in the image labels and the SBOM, and test the full matrix when either changes.

Building a signed image

The build turns pinned inputs into a signed image with evidence attached. Pin the base by digest, not tag, so a republished tag cannot change your image. Lock Python dependencies with hashes so a replaced wheel fails the install.

# syntax=docker/dockerfile:1
# pin bases by digest, not tag
FROM <cuda-devel-base>@sha256:<digest> AS build
COPY requirements.lock /tmp/
RUN python3 -m venv /opt/venv && \
    /opt/venv/bin/pip install --no-cache-dir --require-hashes -r /tmp/requirements.lock

# runtime stage: no compilers in the final image
FROM <cuda-runtime-base>@sha256:<digest>
RUN useradd -u 10001 -m serve
COPY --from=build /opt/venv /opt/venv
COPY app/ /app/
ENV PATH=/opt/venv/bin:$PATH
USER 10001
# never fetch from a model hub at runtime
ENV HF_HUB_OFFLINE=1
ENTRYPOINT ["python", "-m", "app.serve"]

In CI, after the build: generate an SBOM from the image, scan it, attach both as attestations, and sign the digest with a keyless identity tied to the workflow.

IMG=registry.example.com/llm/serve@${DIGEST}
syft "$IMG" -o spdx-json > sbom.spdx.json
grype "$IMG" --fail-on high
cosign attest --yes --predicate sbom.spdx.json --type spdxjson "$IMG"
cosign sign --yes "$IMG"

# anyone can verify later; identity and issuer must both match
cosign verify "$IMG" \
  --certificate-identity "https://github.com/example/llm-serve/.github/workflows/build.yml@refs/heads/main" \
  --certificate-oidc-issuer "https://token.actions.githubusercontent.com"

Scanning a multi-gigabyte CUDA image produces long lists of findings in libraries your server never calls. Triage with VEX statements rather than ignoring the scanner wholesale, and rebuild on a schedule, weekly at least, so base patches land even when your code has not changed.

Signing and verifying weights

Weights need a signature of their own. The OpenSSF model signing project provides a library and CLI, installed with pip install model-signing, that hashes every file in a model directory and signs the resulting manifest as a Sigstore bundle. Verification fails if any file changes, is added or is removed.

# model release job (CI identity signs)
model_signing sign ./llama-8b-instruct --signature model.sig

# init container in the serving pod, before the server starts
model_signing verify ./llama-8b-instruct \
    --signature model.sig \
    --identity "$RELEASE_IDENTITY" \
    --identity-provider "$OIDC_ISSUER"

Run the verify step in an init container that writes to a volume the server mounts read-only. The server never sees unverified bytes, and a failure stops the pod before it accepts traffic. Record the weights digest in the deployment and in every response log, so an answer can be traced to the exact model files that produced it.

Admission and runtime controls

Admission control turns the signature into a rule. With Kyverno, an image verification policy can require that every pod in the serving namespace uses an image signed by your build workflow, and can rewrite tags to digests.

apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
  name: verify-llm-serving-images
spec:
  validationFailureAction: Enforce        # newer Kyverno versions set this per rule
  rules:
  - name: signed-by-build-workflow
    match:
      any:
      - resources:
          kinds: ["Pod"]
          namespaces: ["llm-serving"]
    verifyImages:
    - imageReferences: ["registry.example.com/llm/*"]
      mutateDigest: true
      attestors:
      - entries:
        - keyless:
            subject: "https://github.com/example/llm-serve/.github/workflows/build.yml@refs/heads/main"
            issuer: "https://token.actions.githubusercontent.com"

Pair this with runtime hardening from container security architecture: non-root user, read-only root filesystem, dropped capabilities and no privilege escalation. Inference servers rarely need outbound internet; deny egress by default so a compromised dependency cannot fetch a payload or send data out, and so a forgotten hub download fails loudly instead of silently pulling unverified weights.

Worked example: hardening a vLLM deployment

A platform team serves an 8B instruct model with vLLM on a GPU node pool. Before: a public image tag, weights pulled from a hub at startup, remote code enabled because a tutorial said so. Every pod start downloaded unverified files from the internet and ran whatever Python came with them.

After: the team builds its own serving image from a digest-pinned vendor base with hash-locked wheels, about 9 GB, signed by the CI workflow with an SBOM attestation. The weights are converted to safetensors once, reviewed, stored in an internal bucket and signed by a release job. Each pod runs an init container that downloads the weights, runs model_signing verify and writes to a shared volume; the server starts with HF_HUB_OFFLINE=1, no remote code and no egress. Kyverno rejects any image in the namespace not signed by the build workflow. The node pool's driver version is pinned in the node image and tracked with the same patch calendar. Cold start rises by the verification time, which is bounded by how fast the node can read and hash 16 GB from local disk; the team measures it once and accepts it. For serving details see the vLLM serving engine.

Failure modes

  • Signed image, unsigned weights. Admission passes, and the server loads tampered weights from a bucket. Verify both artifacts.
  • Tag drift. A base tag is republished with different contents. Pin digests and let the policy rewrite tags to digests.
  • Pickle in production. A .bin checkpoint slips through because conversion was skipped. Make the loader reject non-safetensors files.
  • Runtime hub download. A missing local file makes the server fetch from the internet. Set offline mode and deny egress.
  • Scanner fatigue. Hundreds of findings in CUDA layers get blanket-ignored, including a real one. Use VEX and rebuild weekly.
  • Driver blind spot. A vulnerable host driver is never patched because nobody inventories node images.

Trade-offs

ChoiceGainCost
Build your own imageKnown inputs, smaller attack surfaceYou own patching and CUDA compatibility
Vendor image by digestLess work, vendor-testedOpaque contents, their patch cadence
Weights separate from imageIndependent releases, smaller imagesA second artifact to sign and verify
Verify weights at every startTampering caught before servingLonger cold starts

What to do next

  1. Inventory every serving image, its base digest, its weights source and the node driver version.
  2. Pin bases by digest and lock Python dependencies with hashes.
  3. Generate an SBOM, scan and sign every image in CI with a keyless workflow identity.
  4. Convert weights to safetensors, sign them with model_signing, and verify in an init container.
  5. Remove remote code, set hub offline mode and deny egress in the serving namespace.
  6. Enforce signed images with an admission policy that rewrites tags to digests.
  7. Rebuild weekly, triage findings with VEX, and patch node images on the same calendar.
  8. Log the image digest and weights digest with every response.
Key takeaway: Treat an LLM serving deployment as two signed artifacts and a host: a digest-pinned, hash-locked, scanned and signed image verified at admission; safetensors weights signed at release and verified before load; and a node image whose GPU driver is inventoried and patched like everything else.