Helm is the package manager most teams use to put applications on Kubernetes. A chart is a directory of templated manifests plus default values; helm upgrade --install renders the templates with your values, applies them to the cluster, and records the result as a numbered revision of a release. That record is what makes helm rollback and helm history work.

LLM serving stresses every default Helm and Kubernetes assume. A pod may take ten minutes to become ready because it is loading tens of gigabytes of weights. A rolling update may be impossible because there is no spare GPU node to surge onto. Engine flags such as tensor parallel size are tied to the GPU type the pod lands on. And the most important artefact, the model, usually lives outside the chart entirely.

This page shows how to design a chart for an LLM server so that those facts are handled explicitly: a values layout built around model profiles, a schema that rejects impossible combinations, templates with correct GPU resources, shared memory and probes, upgrade settings that match real load times, and a smoke test. The examples use vLLM, whose engine behaviour is covered in vLLM on GPU, in Depth, but the chart patterns apply to any serving engine.

What Helm does and does not manage

Three Helm facts drive most of the design.

  • Releases are stored in the cluster. By default each revision is a Secret of type helm.sh/release.v1 in the release namespace, containing the rendered manifest and values. Rollback re-applies an earlier stored manifest. Nothing else is restored.
  • Helm applies, then optionally waits. Without --wait, Helm reports success as soon as the API server accepts the objects, while pods may still be crash-looping. With --wait, it waits for resources to become ready, up to --timeout, which defaults to five minutes. Five minutes is shorter than many model loads.
  • Helm 4 changed some defaults and flags. Helm 4 uses server-side apply for new releases, while releases created under Helm 3 keep their original apply method unless you opt in. The --atomic flag, which rolls back a failed upgrade automatically, is named --rollback-on-failure in Helm 4, with the old name deprecated. Check helm version in CI and use the flag your version expects.
From values to a serving release: what Helm renders and what it waits forvalues.yaml+ model profile filevalues.schema.jsonreject bad input earlyhelm upgrade --installrender templatesDeployment / StatefulSetGPU requests, probes, shmService + PodDisruptionBudgetSecret ref (hub token)PVC / weight cacheRelease Secretrevision N, manifestrecord--wait --timeoutmust exceed model load timehelm testone real completionHelm rolls manifests back, not data: weights on a PVC and the GPU queue are outside its control.
Figure: values and a schema feed rendering; the release records revision N; --wait holds the command until pods are ready; a test pod sends one real request.

Values built around model profiles

The common mistake is a flat values file with dozens of independent knobs: model name, GPU count, tensor parallel size, max context length, memory utilisation, node selector. Those are not independent. A 70B-parameter model in 16-bit weights needs about 140 GB just for weights, so it cannot run with tensor parallel size 1 on an 80 GB GPU, and its GPU count must equal tensor parallel size times pipeline parallel size. Expose a small number of profiles that are known to work, and keep the free-form knobs for overrides.

# values.yaml
image:
  repository: vllm/vllm-openai
  tag: ""                 # required; pinned per release, never "latest"

profile: llama70b-h100x4  # selects an entry below
profiles:
  llama70b-h100x4:
    model: meta-llama/Llama-3.1-70B-Instruct
    gpus: 4
    tensorParallel: 4
    maxModelLen: 32768
    gpuMemoryUtilization: 0.90
    nodeSelector: { nvidia.com/gpu.product: NVIDIA-H100-80GB-HBM3 }
    loadSeconds: 600      # measured cold start, used for probes and --timeout
  qwen7b-l4x1:
    model: Qwen/Qwen2.5-7B-Instruct
    gpus: 1
    tensorParallel: 1
    maxModelLen: 16384
    gpuMemoryUtilization: 0.90
    nodeSelector: { nvidia.com/gpu.product: NVIDIA-L4 }
    loadSeconds: 180

replicas: 2
weights:
  pvc: model-cache        # pre-populated; see "Getting weights onto the node"
extraArgs: []

Node labels such as nvidia.com/gpu.product come from NVIDIA's GPU feature discovery, normally installed by the GPU Operator; check the exact strings on your nodes with kubectl get nodes --show-labels. Then add a values.schema.json so helm lint, helm template and helm upgrade reject bad input before anything reaches the cluster:

{
  "$schema": "https://json-schema.org/draft-07/schema#",
  "type": "object",
  "required": ["image", "profile", "replicas"],
  "properties": {
    "image": {"type": "object", "required": ["tag"],
              "properties": {"tag": {"type": "string", "minLength": 1, "not": {"const": "latest"}}}},
    "replicas": {"type": "integer", "minimum": 1},
    "profile": {"type": "string", "enum": ["llama70b-h100x4", "qwen7b-l4x1"]}
  }
}

The empty default tag is deliberate: every install must supply a pinned tag, so CI runs lint and template with --set image.tag=....

JSON Schema cannot easily check that GPU count equals tensor parallel size, so add a fail guard in a template helper for cross-field rules:

{{- define "llm.profile" -}}
{{- $p := index .Values.profiles .Values.profile -}}
{{- if not $p }}{{ fail (printf "unknown profile %s" .Values.profile) }}{{ end -}}
{{- if ne (int $p.gpus) (int $p.tensorParallel) -}}
{{- fail (printf "profile %s: gpus (%v) must equal tensorParallel (%v)" .Values.profile $p.gpus $p.tensorParallel) -}}
{{- end -}}
{{- toYaml $p -}}
{{- end -}}

A GPU-aware Deployment template

The Deployment template is where GPU-specific details live. The excerpt below shows the parts that matter; labels and selectors are omitted.

{{- $p := include "llm.profile" . | fromYaml }}
apiVersion: apps/v1
kind: Deployment
spec:
  replicas: {{ .Values.replicas }}
  strategy:
    type: RollingUpdate
    rollingUpdate: { maxSurge: 0, maxUnavailable: 1 }   # no spare GPUs to surge onto
  template:
    spec:
      nodeSelector: {{ toYaml $p.nodeSelector | nindent 8 }}
      terminationGracePeriodSeconds: 120                 # let in-flight streams finish
      containers:
        - name: engine
          image: "{{ .Values.image.repository }}:{{ .Values.image.tag }}"
          args: ["--model", "{{ $p.model }}",
                 "--tensor-parallel-size", "{{ $p.tensorParallel }}",
                 "--max-model-len", "{{ $p.maxModelLen }}",
                 "--gpu-memory-utilization", "{{ $p.gpuMemoryUtilization }}"
                 {{- range .Values.extraArgs }}, {{ . | quote }}{{ end }}]
          resources:
            limits: { nvidia.com/gpu: {{ $p.gpus }} }
          ports: [{ containerPort: 8000, name: http }]
          startupProbe:
            httpGet: { path: /health, port: http }
            periodSeconds: 10
            failureThreshold: {{ add (div $p.loadSeconds 10) 30 }}   # load time + 5 min margin
          readinessProbe:
            httpGet: { path: /health, port: http }
            periodSeconds: 5
          livenessProbe:
            httpGet: { path: /health, port: http }
            periodSeconds: 15
            failureThreshold: 4
          volumeMounts:
            - { name: shm, mountPath: /dev/shm }
            - { name: weights, mountPath: /root/.cache/huggingface }
      volumes:
        - name: shm
          emptyDir: { medium: Memory, sizeLimit: 16Gi }
        - name: weights
          persistentVolumeClaim: { claimName: {{ .Values.weights.pvc }} }

Each choice has a reason:

  • GPUs as limits. Extended resources like nvidia.com/gpu are requested through limits and are not overcommitted; requests default to the limit. For fractional sharing see NVIDIA MIG architecture.
  • maxSurge 0. The default rolling update creates a new pod before removing an old one. On a cluster with no free GPU node, that new pod stays Pending and the upgrade hangs until --timeout. With maxSurge 0 the rollout removes one replica, then replaces it, at the cost of temporarily reduced capacity.
  • Startup probe sized from measured load time. Without it the liveness probe kills the pod mid-load and it restarts forever. The startup probe disables the others until it passes.
  • Shared memory. Multi-GPU engines use /dev/shm for inter-process communication, and the container default is small; a memory-backed emptyDir fixes it.
  • Checksum annotation if you add a ConfigMap. A change to a ConfigMap the pod mounts, such as a custom chat template, does not restart pods on its own. Add a pod annotation like checksum/config: {{ include (print $.Template.BasePath "/configmap.yaml") . | sha256sum }} so that it does.
  • extraArgs appended last. Overrides go after the profile's flags, so a one-off flag can be tried without editing a tested profile.
  • Secrets by reference. A model-hub token for gated models belongs in an existing Kubernetes Secret named in values and mounted as an environment variable, never in the values file itself, because values are stored in every release revision.
  • A PodDisruptionBudget with maxUnavailable: 1 stops node drains from taking every replica at once.

Getting weights onto the node

There are three common ways to get weights onto a node, and Helm treats them differently.

ApproachHowHelm caveat
Download at startEngine pulls from a model hub or bucket on startupLoad time includes download; every restart re-pays it unless cached
Shared PVCA ReadOnlyMany volume populated once by a separate jobRollback does not restore old weights if the volume was overwritten; use a versioned path per model
Init containerCopies from object storage to a local disk before the engine startsInit time counts against --timeout but not the startup probe

Helm hooks such as pre-upgrade can run a download Job, but hook resources are not part of the release, are not rolled back, and block the upgrade while they run. Prefer a separate pipeline step that populates a versioned path, and make the chart reference that path. Pinning the model revision alongside the image tag is the core of GitOps for LLM Deployments, in depth.

Upgrading, testing and rolling back

A safe upgrade for an LLM release has four steps: see the change, apply with a realistic timeout, verify with a real request, and know how to undo.

# 1. see exactly what will change (helm-diff plugin)
helm diff upgrade chat ./llm-chart -f values-prod.yaml --set image.tag=v0.11.0

# 2. apply; timeout covers (replicas x per-pod load time) because maxSurge is 0
helm upgrade --install chat ./llm-chart -f values-prod.yaml \
  --set image.tag=v0.11.0 \
  --wait --timeout 25m \
  --rollback-on-failure          # Helm 3: --atomic

# 3. smoke test
helm test chat --logs

# 4. if needed
helm history chat
helm rollback chat 41 --wait --timeout 25m

Worked timing example. Two replicas of the 70B profile, each measured at 600 s to load. With maxSurge 0 and maxUnavailable 1, replicas are replaced one at a time: roughly 2 x 600 s plus scheduling and image pull, around 22 minutes. A 25-minute timeout is reasonable; the default five minutes would mark a healthy upgrade as failed and, with automatic rollback, start a second ten-minute-per-pod rollout in the other direction. Note too that a rollback is itself a full rollout, with the same load times.

The test is an ordinary pod with the helm.sh/hook: test annotation. Make it send one real completion, not just a health check, because a server can be healthy while serving the wrong model:

{{- $p := include "llm.profile" . | fromYaml }}
apiVersion: v1
kind: Pod
metadata:
  name: "{{ .Release.Name }}-smoke"
  annotations: { "helm.sh/hook": test, "helm.sh/hook-delete-policy": before-hook-creation }
spec:
  restartPolicy: Never
  containers:
    - name: smoke
      image: curlimages/curl:8.10.1
      command: ["sh", "-c"]
      args:
        - >
          curl -sf http://{{ .Release.Name }}:8000/v1/models | grep -q '{{ $p.model }}' &&
          curl -sf http://{{ .Release.Name }}:8000/v1/completions
          -H 'Content-Type: application/json'
          -d '{"model":"{{ $p.model }}","prompt":"2+2=","max_tokens":4}'

For traffic-weighted rollouts, where a new model takes a small share of requests before the rest, Helm alone is not enough; see LLM Canary Deployment.

Failure modes

FailureSymptomFix
Timeout shorter than loadUpgrade reported failed while pods were still loadingDerive --timeout from measured load time and replica count
Surge onto full clusterNew pod Pending, upgrade hangsmaxSurge 0, or reserve one spare node
No startup probeCrashLoopBackOff during load; liveness kills the podStartup probe with failureThreshold from loadSeconds
Small /dev/shmMulti-GPU engine errors or hangs at startMemory-backed emptyDir
Config change without restartNew chat template in ConfigMap, old behaviour in podsChecksum annotation
Weights overwritten in placeRollback restores old image with new weightsVersioned weight paths, referenced from values
Floating image tagTwo replicas run different engine versionsSchema forbids empty or latest tags

What to do next

  1. Measure cold-start time for each model and GPU combination you serve and record it in the profile.
  2. Replace free-form engine knobs with a few tested profiles and add a values schema plus a fail guard.
  3. Set maxSurge 0 or reserve spare GPU capacity, and add a PodDisruptionBudget.
  4. Add startup, readiness and liveness probes, shared memory and a checksum annotation to the template.
  5. Move weight population into a separate, versioned pipeline step.
  6. In CI, run helm lint and helm template for every profile with a pinned image.tag set, and helm diff on every change.
  7. Use --wait with a computed timeout, the rollback-on-failure flag your Helm version expects, and a helm test that sends a real completion.
  8. Practise a rollback in staging and time it.
Key takeaway: A good LLM chart encodes the facts of GPU serving: tested model profiles instead of free knobs, a schema that rejects impossible combinations, GPU limits, startup probes and shared memory, rollouts that do not need spare GPUs, and upgrade timeouts computed from measured load time. Keep weights versioned outside the chart, verify every release with a real completion, and remember that Helm rolls back manifests, not data.