Kustomize builds Kubernetes manifests by layering plain YAML instead of templating it. You write one complete, valid base, and each environment or model variant is an overlay that says how it differs. It ships inside kubectl as kubectl kustomize and kubectl apply -k, and most GitOps controllers render it natively.

LLM serving stresses this model in specific ways. Every deployment is the same shape, an inference server with GPUs, a model volume and slow startup, but each one differs in GPU type and count, tensor-parallel degree, context length and model path, and a wrong value is an out-of-memory crash on a very expensive node. This article shows how to lay out a Kustomize tree for that, which patch types to use where, the traps that bite LLM deployments in particular, and how to check renders before they reach a cluster. For the templated alternative see Helm for LLM deployments; for delivery and rollback see GitOps for LLM deployments.

How Kustomize works

Kustomize's mental model is a pure function: a directory goes in, a stream of manifests comes out. A directory with a kustomization.yaml lists resources (files or other kustomization directories), then transformers that modify them: namespace, namePrefix, labels, images, patches, replacements, and generators that create ConfigMaps and Secrets. Nothing is templated, so every file in the base is valid Kubernetes YAML you can read and apply by itself.

Three kinds of directory matter. A base holds the complete resources. An overlay references a base and adds environment-specific changes. A component (kind: Component, API version kustomize.config.k8s.io/v1alpha1) is a reusable bundle of changes that several overlays can opt into. Components are what make Kustomize work for LLM fleets, because GPU class, shared memory and quantization are cross-cutting choices that do not fit a single inheritance chain.

One base, reusable components, thin overlays, one rendered manifest per targetbase/Deployment, Servicecomponents/gpu-h100nodeSelector, TP=2components/gpu-l41 GPU, short contextcomponents/shm/dev/shm volumeoverlays/dev-llama-8bbase + l4 + shmoverlays/prod-llama-70bbase + h100 + shmoverlays/prod-qwenbase + h100 + model cfgkustomize buildpure function: dir -> YAMLCI checksschema, policy, diffApplykubectl -k or GitOps syncOverlays never copy the base; a change to base/ shows up in every rendered diff
A Kustomize tree for an LLM fleet: the base is complete and valid, components carry cross-cutting GPU and runtime choices, overlays compose them per model and environment, and every render is checked before apply.

A base for an inference server

The base should be a deployable default with nothing environment-specific baked into the parts that overlays will change. Here is a trimmed vLLM base:

# base/deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: llm-server
spec:
  replicas: 1
  selector:
    matchLabels: {app: llm-server}
  template:
    metadata:
      labels: {app: llm-server}
    spec:
      containers:
        - name: server
          image: vllm/vllm-openai:pinned-by-overlay
          args: ["--model", "/models/current", "--port", "8000"]
          ports: [{containerPort: 8000, name: http}]
          resources:
            limits: {nvidia.com/gpu: 1}
          startupProbe:
            httpGet: {path: /health, port: http}
            periodSeconds: 10
            failureThreshold: 60      # 10 minutes to load weights
          readinessProbe:
            httpGet: {path: /health, port: http}
            periodSeconds: 5
          volumeMounts: [{name: models, mountPath: /models, readOnly: true}]
      volumes:
        - name: models
          persistentVolumeClaim: {claimName: model-store}
---
# base/kustomization.yaml
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources: [deployment.yaml, service.yaml]
labels:
  - pairs: {app.kubernetes.io/part-of: llm-serving}
    includeSelectors: false

Two choices in that file prevent the worst Kustomize incidents. The labels transformer is used with includeSelectors: false. The older commonLabels field, now deprecated, also rewrote spec.selector, and a Deployment's selector is immutable, so adding or changing a common label later makes every apply fail with a field-is-immutable error until you delete and recreate the Deployment, which for an LLM server means a full cold start. And the startup probe budget, periodSeconds x failureThreshold, is sized to weight loading time, which the probes article covers in detail.

Patching GPU settings without losing args

Kustomize offers two patch languages through the one patches field (the old patchesStrategicMerge and patchesJson6902 fields are deprecated, and kustomize edit fix migrates them). The choice matters most for container args, which is where LLM servers keep their configuration.

A strategic merge patch is a partial resource that is merged into the target. Lists of containers, ports, env vars and volumes are merged by key, typically name, so you can change one container or add one env var. But args is a plain list of strings with no merge key, so a strategic merge patch replaces the whole list. Patch the args to add a tensor-parallel flag and you silently drop --model and --port; the server then loads its default model or fails to start.

A JSON 6902 patch is a list of explicit operations on paths, and it can append to a list:

# components/gpu-h100/kustomization.yaml
apiVersion: kustomize.config.k8s.io/v1alpha1
kind: Component
patches:
  - target: {kind: Deployment, name: llm-server}
    patch: |-
      - op: add
        path: /spec/template/spec/containers/0/args/-
        value: "--tensor-parallel-size=2"
      - op: add
        path: /spec/template/spec/containers/0/args/-
        value: "--gpu-memory-utilization=0.90"
      - op: replace
        path: /spec/template/spec/containers/0/resources/limits/nvidia.com~1gpu
        value: 2
  - target: {kind: Deployment, name: llm-server}
    patch: |-
      apiVersion: apps/v1
      kind: Deployment
      metadata: {name: llm-server}
      spec:
        template:
          spec:
            nodeSelector:
              nvidia.com/gpu.product: NVIDIA-H100-80GB-HBM3   # use the value your nodes report

Note the ~1 in the path: JSON Pointer escapes / as ~1, so nvidia.com/gpu becomes nvidia.com~1gpu. Forgetting it gives a path-not-found error, which is at least loud. The quieter risk with JSON patches is the index containers/0: if someone adds a sidecar before the server container, the patch edits the wrong container. Keep the server first, and assert it in CI.

The tensor-parallel degree and the GPU limit must agree, so they belong in the same component. If vLLM is told to shard across two GPUs but the pod has one, it fails at start; if the pod has two GPUs and the server uses one, half an expensive node is idle. A separate components/shm component adds a memory-backed emptyDir at /dev/shm, which multi-GPU inference needs because the container default is far too small for NCCL's shared-memory transport.

Generators, pinned images and replacements

Generators produce ConfigMaps and Secrets from files or literals, and by default append a hash of the content to the name, then rewrite every reference to it. Change the content and the name changes, the pod template changes, and the Deployment rolls. That is exactly what you want for server configuration such as a chat template or a sampling-defaults file: configuration changes become rollouts, never silent drift in running pods.

# overlays/prod-llama-70b/kustomization.yaml
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
namespace: llm-prod
nameSuffix: -llama-70b
resources: [../../base]
components: [../../components/gpu-h100, ../../components/shm]
images:
  - name: vllm/vllm-openai
    digest: sha256:REPLACE_WITH_TESTED_DIGEST
configMapGenerator:
  - name: model-config
    literals:
      - MODEL_PATH=/models/llama-70b-instruct/rev-2026-09-30
replacements:
  - source: {kind: ConfigMap, name: model-config, fieldPath: data.MODEL_PATH}
    targets:
      - select: {kind: Deployment, name: llm-server}
        fieldPaths: [spec.template.spec.containers.0.args.1]

The images transformer pins the server by digest, so the rendered manifest names an exact image; tags can be repointed, digests cannot. The replacements block, which supersedes the deprecated vars, copies the model path from the generated ConfigMap into position 1 of the args list, the value after --model. That keeps the model path in one place. It has the same positional fragility as JSON patches, so test it.

When the generator is wrong for you, for example an external system reads the ConfigMap by a fixed name, set generatorOptions: {disableNameSuffixHash: true} for that generator only, and accept that changes will no longer roll the pods on their own.

Worked example: adding a 70B model

Walk through adding a new model, a 70B instruct model on H100s, to a fleet that already serves an 8B model on L4s. You create overlays/prod-llama-70b with the file above, which references the shared base, the H100 and shm components, a pinned image digest, and a ConfigMap with the model path.

Running kustomize build overlays/prod-llama-70b renders a Deployment named llm-server-llama-70b in llm-prod with args --model /models/llama-70b-instruct/rev-2026-09-30 --port 8000 --tensor-parallel-size=2 --gpu-memory-utilization=0.90, a GPU limit of 2, the H100 node selector, a memory-backed /dev/shm volume, and a ConfigMap named model-config-llama-70b- followed by a content hash. CI renders every overlay, not just the changed one, because a change to the base or a component affects all of them, and posts the rendered diff on the pull request.

Reviewers then see the real change. If someone had used a strategic merge patch for the tensor-parallel flag, the diff would show --model disappearing from the args, which is easy to spot in rendered YAML and nearly invisible in the patch file itself. That is the core operational habit: review renders, not patches.

Checking renders in CI

Because the build is a pure function, the checks are cheap and run before anything touches a cluster:

#!/usr/bin/env bash
set -euo pipefail
for o in overlays/*/; do
  out="rendered/$(basename "$o").yaml"
  kustomize build "$o" > "$out"
  kubeconform -strict -summary "$out"            # schema validation
  python3 checks/llm_invariants.py "$out"        # project rules, below
done
git diff --no-index --stat rendered.main/ rendered/ || true   # diff against main's renders
# checks/llm_invariants.py
import sys, yaml

def flag(args, name):
    for i, a in enumerate(args):
        if a.startswith(name + "="):
            return a.split("=", 1)[1]
        if a == name and i + 1 < len(args):
            return args[i + 1]
    return None

for doc in yaml.safe_load_all(open(sys.argv[1])):
    if not doc or doc.get("kind") != "Deployment":
        continue
    c = doc["spec"]["template"]["spec"]["containers"][0]
    assert c["name"] == "server", "server container must be first"
    args = c.get("args", [])
    assert flag(args, "--model"), "args lost --model (strategic merge on args?)"
    gpus = int(c["resources"]["limits"].get("nvidia.com/gpu", 0))
    tp = int(flag(args, "--tensor-parallel-size") or 1)
    assert gpus == tp, f"GPU limit {gpus} != tensor parallel {tp}"
    assert "@sha256:" in c["image"], "image must be pinned by digest"
    if tp > 1:
        mounts = {m["mountPath"] for m in c.get("volumeMounts", [])}
        assert "/dev/shm" in mounts, "multi-GPU pod without /dev/shm volume"
print("ok", sys.argv[1])

Those few assertions encode the failures that actually happen with LLM manifests. Add a policy engine on top if you have one, but keep the project-specific invariants in code that reviewers can read.

Promotion then becomes a one-line change. CI that has tested an image runs kustomize edit set image vllm/vllm-openai@sha256:... in the next environment's overlay and opens a pull request, so the digest moves from dev to staging to production as reviewed commits rather than as manual edits. Model weights move the same way, through the model path in the generated ConfigMap. Where these manifests sit in the wider serving stack, alongside gateways, autoscaling and observability, is laid out in the LLM deployment pattern overview.

Failure modes

  • Args list replaced by a strategic merge patch. The server starts with defaults or crashes. Use JSON 6902 appends and assert required flags in CI.
  • Selector changed by a label transformer. Apply fails as immutable, or worse, a delete-and-recreate takes all replicas down at once. Keep includeSelectors: false for labels added after first deploy.
  • Positional paths hit the wrong container. A new sidecar shifts index 0. Keep the server first and assert its name.
  • Version skew. The Kustomize embedded in kubectl lags the standalone release, so a field that renders locally may fail in a pipeline. Pin one version and use it everywhere.
  • Hash-suffix surprises. An external reference to a generated ConfigMap by its unsuffixed name breaks; and every config edit restarts GPU pods, which may mean minutes of reloading. Plan surge capacity for it.
  • Overlay sprawl. Copying the base into overlays instead of patching it brings back the drift Kustomize exists to prevent. Overlays should be short.

Trade-offs

Kustomize wins when deployments are mostly the same shape and you want every rendered manifest to be readable, diffable YAML with no template logic. Helm wins when you need real parameterization, loops over model lists, values schemas, packaged releases for other teams, or lifecycle hooks. Many teams combine them: a vendor chart rendered once, then Kustomize overlays for local policy, either by rendering the chart to files or with the helmCharts field, which needs the --enable-helm flag. The cost of Kustomize is positional fragility in JSON patches and replacements and a lack of validation of inputs, which the CI checks above make up for.

What to do next

  1. Write one complete vLLM base that applies on its own, with a startup probe sized to load time.
  2. Move GPU class, tensor parallel degree and /dev/shm into components; keep GPU limit and tensor parallel together.
  3. Replace strategic merge patches on args with JSON 6902 appends, and migrate deprecated fields with kustomize edit fix.
  4. Switch commonLabels to labels with includeSelectors false for any label added after first deploy.
  5. Pin the server image by digest through the images transformer.
  6. Render every overlay in CI, validate schemas, run the invariant checks and post rendered diffs on pull requests.
  7. Pin one Kustomize version across laptops, CI and the GitOps controller.
Key takeaway: Kustomize suits LLM fleets because every server is the same shape with a few risky differences. Keep one valid base, put GPU and runtime choices in components, append to args with JSON patches instead of replacing them, keep selectors out of label transformers, pin images by digest, let config generators roll pods, and check every rendered overlay in CI before it reaches a GPU node.