Helm is the package manager most teams use to put applications on Kubernetes. A chart is a directory of templated manifests plus default values; helm upgrade --install renders the templates with your values, applies them to the cluster, and records the result as a numbered revision of a release. That record is what makes helm rollback and helm history work.
LLM serving stresses every default Helm and Kubernetes assume. A pod may take ten minutes to become ready because it is loading tens of gigabytes of weights. A rolling update may be impossible because there is no spare GPU node to surge onto. Engine flags such as tensor parallel size are tied to the GPU type the pod lands on. And the most important artefact, the model, usually lives outside the chart entirely.
This page shows how to design a chart for an LLM server so that those facts are handled explicitly: a values layout built around model profiles, a schema that rejects impossible combinations, templates with correct GPU resources, shared memory and probes, upgrade settings that match real load times, and a smoke test. The examples use vLLM, whose engine behaviour is covered in vLLM on GPU, in Depth, but the chart patterns apply to any serving engine.
What Helm does and does not manage
Three Helm facts drive most of the design.
- Releases are stored in the cluster. By default each revision is a Secret of type
helm.sh/release.v1in the release namespace, containing the rendered manifest and values. Rollback re-applies an earlier stored manifest. Nothing else is restored. - Helm applies, then optionally waits. Without
--wait, Helm reports success as soon as the API server accepts the objects, while pods may still be crash-looping. With--wait, it waits for resources to become ready, up to--timeout, which defaults to five minutes. Five minutes is shorter than many model loads. - Helm 4 changed some defaults and flags. Helm 4 uses server-side apply for new releases, while releases created under Helm 3 keep their original apply method unless you opt in. The
--atomicflag, which rolls back a failed upgrade automatically, is named--rollback-on-failurein Helm 4, with the old name deprecated. Checkhelm versionin CI and use the flag your version expects.
Values built around model profiles
The common mistake is a flat values file with dozens of independent knobs: model name, GPU count, tensor parallel size, max context length, memory utilisation, node selector. Those are not independent. A 70B-parameter model in 16-bit weights needs about 140 GB just for weights, so it cannot run with tensor parallel size 1 on an 80 GB GPU, and its GPU count must equal tensor parallel size times pipeline parallel size. Expose a small number of profiles that are known to work, and keep the free-form knobs for overrides.
# values.yaml
image:
repository: vllm/vllm-openai
tag: "" # required; pinned per release, never "latest"
profile: llama70b-h100x4 # selects an entry below
profiles:
llama70b-h100x4:
model: meta-llama/Llama-3.1-70B-Instruct
gpus: 4
tensorParallel: 4
maxModelLen: 32768
gpuMemoryUtilization: 0.90
nodeSelector: { nvidia.com/gpu.product: NVIDIA-H100-80GB-HBM3 }
loadSeconds: 600 # measured cold start, used for probes and --timeout
qwen7b-l4x1:
model: Qwen/Qwen2.5-7B-Instruct
gpus: 1
tensorParallel: 1
maxModelLen: 16384
gpuMemoryUtilization: 0.90
nodeSelector: { nvidia.com/gpu.product: NVIDIA-L4 }
loadSeconds: 180
replicas: 2
weights:
pvc: model-cache # pre-populated; see "Getting weights onto the node"
extraArgs: []Node labels such as nvidia.com/gpu.product come from NVIDIA's GPU feature discovery, normally installed by the GPU Operator; check the exact strings on your nodes with kubectl get nodes --show-labels. Then add a values.schema.json so helm lint, helm template and helm upgrade reject bad input before anything reaches the cluster:
{
"$schema": "https://json-schema.org/draft-07/schema#",
"type": "object",
"required": ["image", "profile", "replicas"],
"properties": {
"image": {"type": "object", "required": ["tag"],
"properties": {"tag": {"type": "string", "minLength": 1, "not": {"const": "latest"}}}},
"replicas": {"type": "integer", "minimum": 1},
"profile": {"type": "string", "enum": ["llama70b-h100x4", "qwen7b-l4x1"]}
}
}The empty default tag is deliberate: every install must supply a pinned tag, so CI runs lint and template with --set image.tag=....
JSON Schema cannot easily check that GPU count equals tensor parallel size, so add a fail guard in a template helper for cross-field rules:
{{- define "llm.profile" -}}
{{- $p := index .Values.profiles .Values.profile -}}
{{- if not $p }}{{ fail (printf "unknown profile %s" .Values.profile) }}{{ end -}}
{{- if ne (int $p.gpus) (int $p.tensorParallel) -}}
{{- fail (printf "profile %s: gpus (%v) must equal tensorParallel (%v)" .Values.profile $p.gpus $p.tensorParallel) -}}
{{- end -}}
{{- toYaml $p -}}
{{- end -}}
A GPU-aware Deployment template
The Deployment template is where GPU-specific details live. The excerpt below shows the parts that matter; labels and selectors are omitted.
{{- $p := include "llm.profile" . | fromYaml }}
apiVersion: apps/v1
kind: Deployment
spec:
replicas: {{ .Values.replicas }}
strategy:
type: RollingUpdate
rollingUpdate: { maxSurge: 0, maxUnavailable: 1 } # no spare GPUs to surge onto
template:
spec:
nodeSelector: {{ toYaml $p.nodeSelector | nindent 8 }}
terminationGracePeriodSeconds: 120 # let in-flight streams finish
containers:
- name: engine
image: "{{ .Values.image.repository }}:{{ .Values.image.tag }}"
args: ["--model", "{{ $p.model }}",
"--tensor-parallel-size", "{{ $p.tensorParallel }}",
"--max-model-len", "{{ $p.maxModelLen }}",
"--gpu-memory-utilization", "{{ $p.gpuMemoryUtilization }}"
{{- range .Values.extraArgs }}, {{ . | quote }}{{ end }}]
resources:
limits: { nvidia.com/gpu: {{ $p.gpus }} }
ports: [{ containerPort: 8000, name: http }]
startupProbe:
httpGet: { path: /health, port: http }
periodSeconds: 10
failureThreshold: {{ add (div $p.loadSeconds 10) 30 }} # load time + 5 min margin
readinessProbe:
httpGet: { path: /health, port: http }
periodSeconds: 5
livenessProbe:
httpGet: { path: /health, port: http }
periodSeconds: 15
failureThreshold: 4
volumeMounts:
- { name: shm, mountPath: /dev/shm }
- { name: weights, mountPath: /root/.cache/huggingface }
volumes:
- name: shm
emptyDir: { medium: Memory, sizeLimit: 16Gi }
- name: weights
persistentVolumeClaim: { claimName: {{ .Values.weights.pvc }} }Each choice has a reason:
- GPUs as limits. Extended resources like
nvidia.com/gpuare requested through limits and are not overcommitted; requests default to the limit. For fractional sharing see NVIDIA MIG architecture. - maxSurge 0. The default rolling update creates a new pod before removing an old one. On a cluster with no free GPU node, that new pod stays Pending and the upgrade hangs until
--timeout. With maxSurge 0 the rollout removes one replica, then replaces it, at the cost of temporarily reduced capacity. - Startup probe sized from measured load time. Without it the liveness probe kills the pod mid-load and it restarts forever. The startup probe disables the others until it passes.
- Shared memory. Multi-GPU engines use
/dev/shmfor inter-process communication, and the container default is small; a memory-backed emptyDir fixes it. - Checksum annotation if you add a ConfigMap. A change to a ConfigMap the pod mounts, such as a custom chat template, does not restart pods on its own. Add a pod annotation like
checksum/config: {{ include (print $.Template.BasePath "/configmap.yaml") . | sha256sum }}so that it does. - extraArgs appended last. Overrides go after the profile's flags, so a one-off flag can be tried without editing a tested profile.
- Secrets by reference. A model-hub token for gated models belongs in an existing Kubernetes Secret named in values and mounted as an environment variable, never in the values file itself, because values are stored in every release revision.
- A PodDisruptionBudget with
maxUnavailable: 1stops node drains from taking every replica at once.
Getting weights onto the node
There are three common ways to get weights onto a node, and Helm treats them differently.
| Approach | How | Helm caveat |
|---|---|---|
| Download at start | Engine pulls from a model hub or bucket on startup | Load time includes download; every restart re-pays it unless cached |
| Shared PVC | A ReadOnlyMany volume populated once by a separate job | Rollback does not restore old weights if the volume was overwritten; use a versioned path per model |
| Init container | Copies from object storage to a local disk before the engine starts | Init time counts against --timeout but not the startup probe |
Helm hooks such as pre-upgrade can run a download Job, but hook resources are not part of the release, are not rolled back, and block the upgrade while they run. Prefer a separate pipeline step that populates a versioned path, and make the chart reference that path. Pinning the model revision alongside the image tag is the core of GitOps for LLM Deployments, in depth.
Upgrading, testing and rolling back
A safe upgrade for an LLM release has four steps: see the change, apply with a realistic timeout, verify with a real request, and know how to undo.
# 1. see exactly what will change (helm-diff plugin)
helm diff upgrade chat ./llm-chart -f values-prod.yaml --set image.tag=v0.11.0
# 2. apply; timeout covers (replicas x per-pod load time) because maxSurge is 0
helm upgrade --install chat ./llm-chart -f values-prod.yaml \
--set image.tag=v0.11.0 \
--wait --timeout 25m \
--rollback-on-failure # Helm 3: --atomic
# 3. smoke test
helm test chat --logs
# 4. if needed
helm history chat
helm rollback chat 41 --wait --timeout 25mWorked timing example. Two replicas of the 70B profile, each measured at 600 s to load. With maxSurge 0 and maxUnavailable 1, replicas are replaced one at a time: roughly 2 x 600 s plus scheduling and image pull, around 22 minutes. A 25-minute timeout is reasonable; the default five minutes would mark a healthy upgrade as failed and, with automatic rollback, start a second ten-minute-per-pod rollout in the other direction. Note too that a rollback is itself a full rollout, with the same load times.
The test is an ordinary pod with the helm.sh/hook: test annotation. Make it send one real completion, not just a health check, because a server can be healthy while serving the wrong model:
{{- $p := include "llm.profile" . | fromYaml }}
apiVersion: v1
kind: Pod
metadata:
name: "{{ .Release.Name }}-smoke"
annotations: { "helm.sh/hook": test, "helm.sh/hook-delete-policy": before-hook-creation }
spec:
restartPolicy: Never
containers:
- name: smoke
image: curlimages/curl:8.10.1
command: ["sh", "-c"]
args:
- >
curl -sf http://{{ .Release.Name }}:8000/v1/models | grep -q '{{ $p.model }}' &&
curl -sf http://{{ .Release.Name }}:8000/v1/completions
-H 'Content-Type: application/json'
-d '{"model":"{{ $p.model }}","prompt":"2+2=","max_tokens":4}'For traffic-weighted rollouts, where a new model takes a small share of requests before the rest, Helm alone is not enough; see LLM Canary Deployment.
Failure modes
| Failure | Symptom | Fix |
|---|---|---|
| Timeout shorter than load | Upgrade reported failed while pods were still loading | Derive --timeout from measured load time and replica count |
| Surge onto full cluster | New pod Pending, upgrade hangs | maxSurge 0, or reserve one spare node |
| No startup probe | CrashLoopBackOff during load; liveness kills the pod | Startup probe with failureThreshold from loadSeconds |
| Small /dev/shm | Multi-GPU engine errors or hangs at start | Memory-backed emptyDir |
| Config change without restart | New chat template in ConfigMap, old behaviour in pods | Checksum annotation |
| Weights overwritten in place | Rollback restores old image with new weights | Versioned weight paths, referenced from values |
| Floating image tag | Two replicas run different engine versions | Schema forbids empty or latest tags |
What to do next
- Measure cold-start time for each model and GPU combination you serve and record it in the profile.
- Replace free-form engine knobs with a few tested profiles and add a values schema plus a fail guard.
- Set maxSurge 0 or reserve spare GPU capacity, and add a PodDisruptionBudget.
- Add startup, readiness and liveness probes, shared memory and a checksum annotation to the template.
- Move weight population into a separate, versioned pipeline step.
- In CI, run helm lint and helm template for every profile with a pinned image.tag set, and helm diff on every change.
- Use --wait with a computed timeout, the rollback-on-failure flag your Helm version expects, and a helm test that sends a real completion.
- Practise a rollback in staging and time it.