An LLM inference server is a strange thing to health-check. It takes minutes to start, it holds tens of gigabytes of weights in GPU memory, it is deliberately run near saturation, and the most common way it fails is not by exiting but by becoming slow, stuck or silently wrong. The Kubernetes probe defaults were designed for small web processes that start in seconds and answer a health URL in milliseconds. Applied unchanged to a GPU server they kill pods during model load, remove every replica from the load balancer during a traffic spike, and restart healthy processes in a loop while the actual broken GPU stays in service.

This article builds probe design from first principles: what each probe type decides, why model servers break the usual assumptions, how to derive the numbers from measurements, a small health shim that separates cheap checks from expensive ones, and where node-level GPU checks have to take over. It assumes you already run a server such as vLLM or Triton; vLLM on GPU explains what that engine does during startup, and the serving architecture overview places the server in the full request path.

Three probes, three questions

Kubernetes offers three probes, and each one answers a different question with a different consequence. Mixing up the questions is the root of almost every probe incident.

ProbeQuestionOn failureMust be
startupProbeHas the process finished initialising?Container restarted once the budget is spentGenerous: covers the slowest legitimate start
readinessProbeShould new requests be sent here right now?Pod removed from Service endpoints; not restartedHonest about capacity, quick to recover
livenessProbeIs the process wedged beyond self-recovery?Container killed and restartedCheap, conservative, independent of load

Two mechanics matter. First, while a startup probe is defined and has not yet succeeded, the kubelet does not run the liveness or readiness probes at all. The startup budget is therefore failureThreshold x periodSeconds, and that product, not initialDelaySeconds, is the right way to give a slow server time to load. Second, the defaults are tight: periodSeconds is 10, timeoutSeconds is 1 and failureThreshold is 3. A one-second timeout is fine for a process that answers from memory and dangerous for anything that has to wait behind a busy event loop.

The asymmetry of consequences is the design principle. A wrong readiness verdict costs seconds of capacity and corrects itself. A wrong liveness verdict throws away minutes of loading and shifts load onto the surviving replicas, making their checks more likely to fail too.

Why model servers break the defaults

Four properties of model servers break web-service habits.

  • Startup is long and variable. Pulling a multi-gigabyte image, fetching weights from object storage or a node cache, copying them into GPU memory, profiling memory for the KV cache, capturing CUDA graphs and possibly compiling kernels can take anywhere from under a minute to well over ten, and the same deployment varies with cache warmth and storage throughput.
  • The server is designed to run full. Continuous batching keeps the GPU saturated; a long queue is normal operation. A check that equates queueing with sickness fires exactly when the fleet is most needed.
  • Real work is expensive. The only end-to-end proof that a model works is a generation, which competes with user traffic for the same GPU. Putting that on a probe path ties the probe result to load.
  • Hardware faults hide behind a live process. A GPU that has thrown an uncorrectable memory error or fallen off the bus can leave the HTTP server answering. vLLM's /health returns 503 when the engine has died with EngineDeadError and 200 otherwise; users have reported cases where a GPU fault leaves the process up and the endpoint still green while inference fails. Treat any single endpoint as evidence about the process, not proof about the GPU.

Architecture: probes, shim and node agent

Who asks which question: probes, the health shim and node-level GPU checkskubeletstartup, readiness, livenessService / gatewayroutes to ready pods onlyInference podHealth shim (sidecar)/startup /ready /liveInference server/health, /v1/modelsHTTP proberequestsCanary loop1-token generation, asynccached resultGPU node agentDCGM watches, Xid, ECCNode controllercordon, taint, drainunhealthy GPUevict podProbes judge the process. Node agents judge the hardware. Only the second can move work off a bad GPU.
The kubelet probes a small shim, which combines the server's own endpoints with a cached canary result. A separate node agent watches the GPUs and is the only component that can cordon a bad machine.

The architecture splits the problem into three layers. The inference server exposes its native endpoints: for vLLM, /health and the OpenAI-compatible /v1/models; for Triton, /v2/health/live and /v2/health/ready. A health shim, either a sidecar or a thread in a wrapper process, turns those into three probe endpoints with deliberately different logic, and runs a background canary that sends a tiny generation every so often and caches the result with a timestamp. The probe handlers only read cached state, so they answer in microseconds regardless of load.

Below the pod, a GPU node agent (DCGM health watches, Xid log scraping, or Kubernetes Node Problem Detector with a custom GPU check) judges the hardware and acts on the node: taint it, cordon it, drain it, and hand it to repair. This layer exists because restarting a container is useless when the fault is in the device; the new container lands on the same GPU and fails the same way. The DCGM article covers the health watches and diagnostics that feed this layer.

Designing each probe

Startup. Point it at something that is true only once the model is servable. /v1/models is a reasonable choice for vLLM because it succeeds only once the model is loaded and servable. Derive the budget from measurement: record startup time over a few dozen cold and warm starts, take the worst cold start, and add 50 percent. Keep periodSeconds modest so a fast start is detected quickly; the budget comes from the threshold.

Readiness. Ready means the engine is alive, the last canary succeeded recently, and the pod is not draining. Be careful with load signals. Marking a pod unready because its queue is long removes it from rotation, which pushes its share onto the others, which lengthens their queues. Load shedding belongs in admission control (a 429 with a retry hint), not in readiness. The legitimate load-related use of readiness is a hard ceiling that means the pod cannot accept even one more request without failing, such as KV cache exhaustion with preemption thrashing, and even that should be guarded by a fleet-wide floor.

Liveness. Ask only whether the process can still make progress: the event loop answers, the engine has not died, and the canary has not been failing for a long time. Give it a generous timeout and a high failure threshold. A liveness probe that would trigger within thirty seconds of trouble will trigger on garbage-collection pauses, slow first requests after a deploy, and CPU contention from the tokenizer.

containers:
- name: server
  image: vllm/vllm-openai:pinned-tag
  args: ["--model", "/models/llama-70b", "--served-model-name", "llama-70b",
         "--tensor-parallel-size", "4"]
  ports: [{containerPort: 8000}]
  # Probes live on the server container but target the shim's port (shared pod network),
  # so a liveness failure restarts the inference process, not the shim.
  startupProbe:
    httpGet: {path: /startup, port: 8081}
    periodSeconds: 10
    failureThreshold: 90        # 15 minutes: worst measured cold start 9.5 min + 50%
  readinessProbe:
    httpGet: {path: /ready, port: 8081}
    periodSeconds: 5
    timeoutSeconds: 2
    failureThreshold: 2         # out of rotation within ~10 s
    successThreshold: 1
  livenessProbe:
    httpGet: {path: /live, port: 8081}
    periodSeconds: 20
    timeoutSeconds: 5
    failureThreshold: 6         # two minutes of sustained failure before a restart
- name: health-shim
  image: registry.example.com/llm-health-shim:1.4
  env: [{name: UPSTREAM, value: "http://127.0.0.1:8000"}]
  ports: [{containerPort: 8081}]

Note where the probes sit. A liveness probe restarts the container it is defined on, so probes placed on the shim container would restart the shim and leave a wedged server running. The example defines them on the server container and points them at the shim's port, which works because containers in a pod share a network namespace. The served model name is set explicitly so the shim's canary can address it.

A health shim in code

The shim is small enough to read in one screen. The important property is that handlers never block on the GPU; only the background loop does.

import asyncio, time, httpx
from fastapi import FastAPI, Response

UP = "http://127.0.0.1:8000"
app = FastAPI()
state = {"started": False, "engine_ok": False, "canary_ok_at": 0.0,
         "canary_fail_since": None, "draining": False}

async def canary_loop():
    async with httpx.AsyncClient(timeout=30) as http:
        while True:
            try:
                r = await http.get(f"{UP}/health", timeout=3)
                state["engine_ok"] = r.status_code == 200
                if not state["started"]:
                    state["started"] = (await http.get(f"{UP}/v1/models", timeout=3)).status_code == 200
                if state["started"] and state["engine_ok"]:
                    g = await http.post(f"{UP}/v1/completions", json={
                        "model": "llama-70b", "prompt": "2+2=", "max_tokens": 1, "temperature": 0})
                    if g.status_code == 200 and g.json().get("choices"):
                        state["canary_ok_at"] = time.time()
                        state["canary_fail_since"] = None
                    else:
                        raise RuntimeError(f"canary status {g.status_code}")
            except Exception:
                state["canary_fail_since"] = state["canary_fail_since"] or time.time()
            await asyncio.sleep(15)

@app.on_event("startup")
async def _start():
    asyncio.create_task(canary_loop())

@app.get("/startup")
def startup():
    return Response(status_code=200 if state["started"] else 503)

@app.get("/ready")
def ready():
    fresh = time.time() - state["canary_ok_at"] < 60
    ok = state["started"] and state["engine_ok"] and fresh and not state["draining"]
    return Response(status_code=200 if ok else 503)

@app.get("/live")
def live():
    since = state["canary_fail_since"]
    wedged = since is not None and time.time() - since > 300
    return Response(status_code=503 if (state["started"] and wedged) else 200)

@app.post("/drain")          # called from preStop
def drain():
    state["draining"] = True
    return {"draining": True}

Three choices are deliberate. The canary uses a one-token completion at temperature zero, so it costs one prefill of a few tokens and proves the full path from tokenizer to sampler. Liveness tolerates five minutes of canary failure, because a canary timing out under heavy load is not a reason to restart. And /drain lets a preStop hook take the pod out of rotation before termination, so in-flight streams can finish within terminationGracePeriodSeconds.

Worked example: three incidents on a 70B deployment

Take a 70B model served with tensor parallelism across four GPUs. Twenty measured starts give warm starts (weights in the node's local cache) of 3 to 4 minutes and cold starts (weights from object storage) of up to 9.5 minutes. The original configuration had only a liveness probe on /health with initialDelaySeconds: 300. Every cold start was killed at about the 5.5-minute mark, restarted, found the weights still uncached because the download had been interrupted, and was killed again: a crash loop that looked like a broken image.

With the startup budget of 15 minutes from the YAML above, cold starts complete. The next incident came from readiness. A team had added a rule marking a pod unready when more than 64 requests were waiting. During a launch, all eight replicas crossed 64 within a minute; each flapped out of the endpoints list, the remaining ready pods absorbed the traffic and crossed the threshold themselves, and for about forty seconds the Service had no ready endpoints at all, so every request failed instead of queueing. Plain Kubernetes Services have no panic mode that routes to unready pods when all are unready; some Envoy-based gateways do, via a panic threshold. Removing the queue rule and adding a 429 at the gateway when the estimated wait exceeded the latency objective turned a total outage into a few percent of shed requests.

The third incident is the one probes cannot fix. One node logged an Xid error indicating a GPU memory fault. The canary began failing, readiness removed the pod, and five minutes later liveness restarted it, onto the same GPUs, where it failed again. The fix was a node-level rule: on specific Xid codes or DCGM health failures, taint the node, let the pod reschedule elsewhere, and send the node to diagnostics.

Failure modes

SymptomCauseFix
CrashLoopBackOff only on cold nodesLiveness or startup budget shorter than cold loadStartup probe sized from measured worst case
All replicas unready during spikesLoad signal in readinessAdmission control and 429; fleet floor on readiness
Restarts with no errors in server logs1 s timeout behind a busy event loopProbe a shim that reads cached state; raise timeout
Pod green, users get errorsHealth URL checks the process, not inferenceCached canary generation feeds readiness
Same pod fails after every restartFaulty GPU or NVLinkNode agent taints and drains; restart is not repair
Dropped streams on every deployNo drain before SIGTERMpreStop calls /drain; grace period covers longest stream
Canary hides partial failureOne tiny prompt never exercises long contextsOccasional larger canary off the probe path, alert only

Trade-offs

Every probe is a trade between detection speed and false positives, and the costs are not symmetric. Fast readiness is cheap to get wrong; fast liveness is expensive. A deeper check (a real generation) catches more faults but costs GPU time and ties the verdict to load, which is why it belongs in a background loop whose result is read, not executed, by the probe. A sidecar adds a container; an in-process shim avoids that but shares the event loop it is judging.

Some teams drop liveness entirely and rely on the server exiting on fatal errors plus node-level checks. Probes are also not a substitute for fleet-level signals: a pod that is merely slower than its peers is a scoring and routing problem, covered in LLM system health scoring, and finding the load at which the server degrades is a job for load testing, whose results should set your admission thresholds.

What to do next

  1. Measure twenty cold and warm starts per model and set the startup budget to the worst cold start plus 50 percent.
  2. Remove every load signal from readiness; move shedding to the gateway as a 429 with a retry hint.
  3. Add a health shim whose probe handlers read cached state and never call the GPU directly.
  4. Run a one-token canary every 15 seconds and feed its freshness into readiness.
  5. Raise liveness timeouts and thresholds so only minutes of sustained failure cause a restart.
  6. Decide which container the probes live on, so a liveness failure restarts the process you meant.
  7. Add a preStop drain and a grace period longer than your longest expected stream.
  8. Wire DCGM or Xid-based node checks to taint and drain bad GPUs, and test it by injecting a fault in staging.
Key takeaway: Give each probe one question. Size the startup budget from measured cold starts, keep load out of readiness and shed at the gateway instead, make liveness cheap and slow to anger, and let probe handlers read a cached canary result rather than touching the GPU. Restarting a container never repairs a GPU, so pair probes with node-level checks that cordon and drain bad hardware.