GitOps means the desired state of a system lives in a git repository, and an agent in the cluster keeps pulling that state and making the cluster match it. Argo CD and Flux are the two common agents. For a stateless web service this is routine. For an LLM server it gets harder: the artifact is tens or hundreds of gigabytes of weights that cannot go in git, a pod can take minutes to become ready, every replica needs a GPU that may not be free, and a release is not one version but three. Those are the engine image, the model revision and the serving flags.
This article shows how to lay out a deploy repo for model servers, pin a release so it can be reproduced, get weights to pods without breaking the git model, set health checks that survive slow starts, and what rollback by git revert does and does not undo. Traffic splitting and abort thresholds are covered in LLM canary deployment. This article covers the delivery mechanics underneath them.
Why model servers strain GitOps
Three properties of LLM serving break naive GitOps:
- The artifact is not in the manifest. A Deployment for a web service names an image, and the image is the whole release. A Deployment for vLLM names an engine image and a model id. If the model id is a moving name such as a repository's main branch, the same manifest can serve different weights next week. Git then no longer describes what is running, which defeats the point of GitOps.
- Readiness is slow and variable. A pod has to pull a multi-gigabyte image, load weights into GPU memory, and let the engine warm up (vLLM, for example, profiles memory and may capture CUDA graphs; vLLM on GPU explains where that memory goes) before it can serve. That is minutes, not seconds, and it varies with cache state and storage bandwidth.
- Capacity is scarce. A rolling update with the default surge needs spare GPUs for the new pod while the old one still runs. On a full node pool the new pod stays Pending, and the sync never finishes.
The fix is not a different tool. It is to be precise about what git records, and to tune health and rollout settings for slow, expensive pods. Where GitOps sits among the other serving choices is covered in LLM deployment patterns, and multi-node expert-parallel launches, where one release spans many pods, in MoE expert parallelism deployment.
A deploy repo for model servers
Keep application code and deployment state in separate repositories. CI builds the engine image and runs evaluations in the app repo. A promotion is a pull request to the deploy repo that changes an image digest, a model revision or a flag. Reviewers can then see exactly what changed in a release, and the GitOps agent only watches the deploy repo.
deploy-repo/
base/llm-server/
deployment.yaml # vLLM container, probes, GPU requests
service.yaml
weights-pvc.yaml
prefetch-job.yaml # fills the PVC, runs before the Deployment
kustomization.yaml
models/
support-chat.yaml # model repo id + immutable revision + flags
envs/
staging/kustomization.yaml # image digest, replicas, overrides
prod-eu/kustomization.yaml
prod-us/kustomization.yamlEach environment overlay pins the image by digest. A tag such as :latest or even :v0.x can be re-pushed. A digest cannot, so the overlay alone describes the engine binary:
# envs/prod-eu/kustomization.yaml
resources:
- ../../base/llm-server
images:
- name: vllm/vllm-openai
digest: sha256:<digest of the engine image CI tested>
patches:
- path: model-support-chat.yaml
Pinning the release tuple
The model side needs the same rule. Hugging Face style repositories are git repositories, and a commit hash is immutable where a branch name is not. vLLM accepts --revision for exactly this. Put the model id, the revision and every flag that changes output or memory in one patch file, so one PR diff shows the whole release:
# model-support-chat.yaml (strategic merge patch)
apiVersion: apps/v1
kind: Deployment
metadata:
name: llm-server
spec:
template:
metadata:
annotations:
release.example.com/model-revision: "<40-char commit hash>"
spec:
containers:
- name: vllm
args:
- "--model=/models/support-chat/<rev>" # per-revision dir on the PVC
- "--served-model-name=support-chat"
- "--tensor-parallel-size=2"
- "--max-model-len=16384"
- "--gpu-memory-utilization=0.90"
resources:
limits:
nvidia.com/gpu: 2The tuple is engine digest, model revision, flags. Treat it as one versioned unit. A tokenizer fix in a new model revision, a changed default in a new engine release and a different --max-model-len can each change outputs. When a quality regression appears, a single PR that changed only one element of the tuple is what lets you bisect. Use --served-model-name to keep the client-facing name stable while the revision underneath changes, and put the revision in an annotation and in your request logs so every response can be traced to its weights.
Getting weights to the pod
Weights must reach the node without being stored in git and without every pod downloading them from the internet on start. The options trade start time against operational complexity:
| Approach | Start time | Reproducible | Main risk |
|---|---|---|---|
| Download in the serving container at start | Slowest, repeated per pod | Only with a pinned revision | Rate limits and outages of the model host block scale-up |
| Pre-sync job fills a shared PVC | Fast after the first fill | Yes, if the job pins the revision | ReadWriteMany storage bandwidth sets the load time |
| Node-local cache filled by a DaemonSet | Fastest on warm nodes | Yes, keyed by revision | Disk pressure and stale copies on nodes |
| Weights baked into an image or OCI artifact | Image pull time | Yes, by digest | Very large images and registry limits |
A good default is a pre-sync job. In Argo CD, a resource with a lower argocd.argoproj.io/sync-wave annotation is applied and must become healthy before resources in later waves, so the job that downloads the pinned revision runs before the Deployment changes:
apiVersion: batch/v1
kind: Job
metadata:
name: prefetch-support-chat-<short-rev> # Jobs are immutable: new name per release
annotations:
argocd.argoproj.io/sync-wave: "-1"
spec:
backoffLimit: 2
template:
spec:
restartPolicy: Never
containers:
- name: fetch
image: python:3.12-slim
command: ["sh", "-c"]
args:
- pip install -q huggingface_hub &&
python -c "from huggingface_hub import snapshot_download;
snapshot_download('org/support-chat', revision='$(REV)',
local_dir='/models/support-chat/$(REV)')"
env:
- name: REV
value: "<40-char commit hash>"
volumeMounts:
- { name: weights, mountPath: /models }
volumes:
- name: weights
persistentVolumeClaim: { claimName: llm-weights }Give the PVC wave -2 so it exists before the job on a fresh install. Name the job per revision, because a Job's pod template cannot be changed after creation. Write each revision to its own directory if you need two versions at once during a canary. Then old pods keep reading the old weights while new pods load the new ones. Overwriting one directory in place means a restarted old pod silently loads new weights. Flux has no sync waves, but it gets the same ordering with a separate Kustomization for the prefetch step and a dependsOn reference from the serving Kustomization.
Health checks that survive slow starts
Slow starts cause more failed GitOps syncs for LLM servers than anything else. Kubernetes and the GitOps agent each have a clock, and both must be longer than the slowest realistic start.
Estimate that start first. Say a model is 140 GB of bf16 weights. Read from a node-local NVMe disk at about 2 GB/s, that is roughly 70 seconds. From shared network storage at 500 MB/s, it is nearly five minutes. Add image pull and engine warm-up on top. These are example figures; measure your own on a cold node, because the cold case is the one that happens during an incident.
startupProbe:
httpGet: { path: /health, port: 8000 }
periodSeconds: 10
failureThreshold: 90 # up to 15 minutes to start
readinessProbe:
httpGet: { path: /health, port: 8000 }
periodSeconds: 5
failureThreshold: 3
livenessProbe:
httpGet: { path: /health, port: 8000 }
periodSeconds: 15
failureThreshold: 4The startup probe holds off liveness checks until the server first answers, so a slow load is not killed and restarted in a loop. Then raise the Deployment's progressDeadlineSeconds above the same worst case. Its default is 600 seconds. When it expires, the Deployment reports that its progress deadline was exceeded, and Argo CD shows the application as Degraded, even if the pods would have become ready a minute later.
For capacity, choose the rollout shape from the GPUs you actually have free. With no spare GPUs, maxSurge: 0 and maxUnavailable: 1 replace pods one at a time and accept reduced capacity during the rollout. With spare GPUs, maxSurge: 1 keeps full capacity. What fails is the default surge on a full node pool. The new pod waits in Pending, the sync shows as progressing forever, and someone eventually syncs again by hand.
Metric-gated promotion
A plain rolling update promotes a new model as soon as its pods pass a health check, and a health check says nothing about answer quality or latency. Argo Rollouts and Flagger add promotion steps that query metrics. vLLM exposes Prometheus metrics such as vllm:time_to_first_token_seconds and vllm:num_requests_waiting, which are good inputs for an automated gate:
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
name: llm-ttft
spec:
args:
- name: pod-template-hash
metrics:
- name: ttft-p95
interval: 2m
count: 5
failureLimit: 1
successCondition: result[0] <= 1.5
provider:
prometheus:
address: http://prometheus.monitoring:9090
query: |
histogram_quantile(0.95, sum by (le) (rate(
vllm:time_to_first_token_seconds_bucket{
pod_template_hash="{{args.pod-template-hash}}"}[2m])))The label you filter on depends on how your Prometheus scrape config relabels pods, so check that the query returns data for a canary before trusting it. A query that returns nothing does not prove the canary is healthy. Latency gates catch engine and configuration regressions. They do not catch a model that answers quickly and wrongly. Run the offline evaluation in CI before the PR is opened, and use the canary for what only live traffic shows. With Argo Rollouts, the Deployment becomes a Rollout (or a Rollout that references it through workloadRef); the args and probes stay the same. The canary article covers thresholds, sticky routing and bake time.
Drift, self-heal and pruning
With syncPolicy.automated and selfHeal: true, Argo CD reverts manual changes to the cluster. That is the point of GitOps, but it surprises people during an incident. An on-call engineer who runs kubectl edit to lower --gpu-memory-utilization at 3 a.m. sees the change undone within minutes. Decide the policy in advance: either hotfixes go through a fast-tracked PR, or the runbook says to pause auto-sync on the application first and record that it was paused.
Some fields should be owned by a controller rather than git. If a HorizontalPodAutoscaler or KEDA scales replicas, leave replicas out of the manifest, or tell Argo CD to ignore that field with ignoreDifferences. Otherwise each sync resets the replica count.
prune: true deletes resources that were removed from git. Protect the weights PVC with the argocd.argoproj.io/sync-options: Prune=false annotation.
What rollback by revert does not undo
Reverting the promotion commit restores the previous tuple in git, and the agent rolls the Deployment back. That is a real advantage: the rollback is reviewed, logged and uses the same path as a release. But some things are not restored:
- Weights that were deleted. If the old revision's directory was cleaned up, the rollback has to download it again before pods can start. Keep at least the previous revision cached.
- Capacity. Rolling back is another rollout with the same slow starts and GPU needs. Plan the time it takes, and do not assume it is instant.
- Caches and state outside the pod. Prefix caches start cold, and conversations that began on the new model continue on the old one.
- Data written by the bad version. Summaries, labels or embeddings the bad model produced and stored are still there. Tag stored outputs with the model revision so you can find and reprocess them.
Worked example: a revision that queues
A team serves a support chat model on two GPUs per replica with four replicas. They want to ship a new fine-tuned revision. CI runs the offline evaluation set against the new revision with the same engine digest and flags, and passes. The bot opens a PR that changes only the revision hash, in three places: the annotation, the prefetch job and the model path.
On merge, Argo CD runs wave -1. The prefetch job writes the revision into a new directory on the PVC. That takes six minutes. Then wave 0 updates the Rollout. The first canary pod starts, and its startup probe passes after about four minutes. Analysis runs for ten minutes at 10 percent of traffic. The p95 time to first token stays under the threshold, but the queue metric climbs. The new revision's chat template produces longer prompts, so the KV cache fills sooner and requests wait. The failure limit trips and the Rollout aborts. Traffic returns to the stable pods.
Nobody ran kubectl. The revert PR records why the release was rejected. The next attempt changes one more element of the tuple, a lower --max-num-seqs plus one extra replica. The reviewer can see that the release now costs more GPUs, which is a decision that should happen in review, not at 3 a.m.
What to do next
- Write down the release tuple for each model you serve: engine image digest, model revision hash, flags. Anything still referenced by a tag or branch is a bug.
- Move deployment state into a deploy repo watched by Argo CD or Flux, with environments as overlays.
- Add a prefetch step ordered before the Deployment (a sync wave or
dependsOn), writing each revision to its own directory. - Measure a cold start on a fresh node, then set the startup probe and
progressDeadlineSecondsabove it. - Choose
maxSurgeandmaxUnavailablefrom the GPUs actually free in the pool. - Protect the weights PVC from pruning, and remove HPA-owned fields from git.
- Add a metrics-based analysis step, and check that its query returns data for a canary pod.
- Rehearse a rollback by revert in staging and time it from merge to full traffic.