vLLM, SGLang and TensorRT-LLM solve the GPU half of LLM serving: batching, paged KV cache, kernels. They do not decide how many replicas run, where the weights come from, which node gets a GPU, how traffic reaches a pod, or what happens during a rollout. On Kubernetes, KServe is one of the common answers to those questions. It is a CNCF project that turns a single custom resource into pods, services, routes and autoscalers, and for LLMs it ships a Hugging Face serving runtime that runs vLLM underneath and speaks OpenAI-style HTTP.

This article treats KServe as what it is, a control plane, and explains the two resource types you can deploy LLMs with, how a replica gets its weights, why default autoscaling signals fail for LLMs and what to scale on instead, how multi-GPU and multi-node serving are expressed, and the failure modes that show up in production. KServe moves quickly, so every field name below comes from the current KServe documentation; check kubectl explain against your installed version before copying YAML.

Control plane and data plane

Think of the system in two planes. The data plane is the inference container: a vLLM engine, wrapped by KServe's Hugging Face runtime, serving requests on port 8080. Each request lands on one pod, and inside that pod the engine's scheduler decides which sequences share a forward pass. Everything about tokens per second, KV cache pressure and time to first token lives here, and the vLLM engine deep dive covers it.

The control plane is the KServe controller plus Kubernetes. It watches your custom resource and reconciles it into lower-level objects: a Deployment (or a LeaderWorkerSet for multi-node), a Service, an HTTP route through an ingress or Gateway, and an autoscaler. It never sees a token. Most production pain with KServe is a control-plane mismatch with data-plane reality: scaling on the wrong signal, a probe that passes before the model loads, or a rollout that drains a pod with thirty streams still open.

InferenceService has two deployment modes. Standard mode creates plain Deployments and Services; the Knative mode (the docs also call it Serverless) uses Knative Serving revisions and its request-based autoscaler. For GPU-backed LLMs, Standard mode with an explicit autoscaler is the usual choice, because Knative's concurrency-based scaling and scale-to-zero assume replicas that start in seconds.

Two resource types for LLMs

KServe for LLMs: a control plane that turns one custom resource into a serving stackInferenceServicev1beta1, huggingface runtimeLLMInferenceServicev1alpha1, GenAI-first CRDKServe controllerreconciles the specDeployment or LWSpods with GPU requestsService, route, gatewayOpenAI-style endpointsAutoscalerKEDA ScaledObject or HPAPodstorage initializer + engineModel storagehf://, s3://, pvc://, cacheweightsPrometheusvllm:num_requests_runningmetricThe engine (vLLM) does the GPU work; KServe owns everything around it: pods, storage, routing, scaling, rollout.
Both resource types feed the same controller, which creates the workload, the route and the autoscaler. The engine inside the pod is vLLM.

There are two resource types for LLMs, and choosing between them is the first design decision.

InferenceService (serving.kserve.io/v1beta1) is KServe's long-standing API. For LLMs you set modelFormat.name: huggingface and a storageUri. The runtime picks a backend automatically: vLLM for supported generative models, the plain Hugging Face backend for predictive tasks such as classification, with a fallback if vLLM does not support the architecture. It exposes /openai/v1/completions, /openai/v1/chat/completions, /openai/v1/embeddings and /openai/v1/rerank; the openai prefix is configurable with the KSERVE_OPENAI_ROUTE_PREFIX environment variable. Engine flags such as --tensor_parallel_size, --max_model_len, --quantization and --gpu_memory_utilization pass through as container args.

LLMInferenceService (serving.kserve.io/v1alpha1 in the current docs; some newer material shows a later alpha version, so check your CRD) is the GenAI-first resource. Its spec has model.uri, replicas, a pod template for single-node or decode workloads, worker for multi-node workers (which produces a LeaderWorkerSet), prefill for a separate prefill pool in disaggregated serving, parallelism for tensor, data and expert parallel layouts, and router for gateway, route and scheduler configuration. The router integrates llm-d's endpoint picker, which can route on engine state such as queue depth and prefix-cache locality instead of round robin.

Rule of thumb: a single-node model behind ordinary load balancing fits InferenceService well and uses a stable API. Reach for LLMInferenceService when you need multi-node workers, prefill/decode disaggregation or cache-aware routing, and accept an alpha API that may change between releases.

Worked example: one model on one GPU

A worked deployment: an 8B instruct model on one GPU in Standard mode. Size the resources from the model, not from a quick-start. Eight billion parameters in bf16 is about 16 GB of weights in GPU memory before any KV cache, so the GPU needs at least 24 GB to leave room for useful batch sizes, and the container needs enough host memory to stage the weights while loading; the 2 GiB memory requests found in some minimal examples are illustrative, not production sizing.

apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
  name: llama3-8b
  annotations:
    serving.kserve.io/deploymentMode: "Standard"
spec:
  predictor:
    minReplicas: 1
    maxReplicas: 4
    model:
      modelFormat:
        name: huggingface
      args:
        - --model_name=llama3
        - --max_model_len=8192
        - --gpu_memory_utilization=0.90
      storageUri: "hf://meta-llama/meta-llama-3-8b-instruct"
      resources:
        requests:
          cpu: "4"
          memory: 32Gi
          nvidia.com/gpu: "1"
        limits:
          cpu: "8"
          memory: 32Gi
          nvidia.com/gpu: "1"

Gated Hugging Face models need a token: the KServe docs wire an HF_TOKEN secret through a ClusterStorageContainer so the storage initializer can authenticate. The model field in each request must match --model_name:

curl http://${INGRESS_HOST}:${INGRESS_PORT}/openai/v1/chat/completions \
  -H "Host: ${SERVICE_HOSTNAME}" \
  -H "Content-Type: application/json" \
  -d '{"model": "llama3",
       "messages": [{"role": "user", "content": "Name three uses of a KV cache."}],
       "max_tokens": 100, "stream": false}'

Because the endpoint is OpenAI-shaped, existing clients work by changing the base URL to http://<host>/openai/v1. Setting --max_model_len explicitly matters: left unset, the engine uses the model's full context window, and a long default context reserves KV space per sequence that a 24 GB card cannot afford. The capacity estimation guide shows the arithmetic for choosing it.

Model loading and cold start

Where a new LLM replica spends its cold start, in orderNodeprovision GPU nodeImagepull engine imageWeightsdownload or cache hitLoadweights into HBMWarm-upgraphs, KV poolPre-provisioned nodes remove stage 1, image pre-pulls remove stage 2, a local model cache removes most of stage 3.Stages 4 and 5 remain, so a replica is never instant: autoscaling must start before the queue is full.
Cold start is five serial stages. Caching attacks the middle three; the last two are physics.

When a pod starts, a storage initializer resolves storageUri (hf://, s3://, gs://, pvc:// and others) and downloads the weights into a volume the engine reads. For a 16 GB checkpoint at an effective 200 MB/s from object storage, that stage alone is about 80 seconds; add a multi-gigabyte engine image pull on a fresh node, the weight load into GPU memory and the engine's warm-up, and a new replica on a new node routinely takes minutes. If the cluster autoscaler must add the GPU node first, add that provisioning time too.

KServe offers several ways to cut this. A LocalModelCache pre-stages model files on node-local storage for a group of nodes, so new pods on those nodes skip the download. A PVC holding the weights (pvc://) avoids repeated object-store reads. Packaging weights as an OCI image (the modelcar approach, oci:// URIs where enabled) lets the node's image cache and registry mirrors do the work. Pre-pulling the engine image with a DaemonSet removes the image stage. None of these remove the load and warm-up, so the operational rule stands: keep minReplicas at the floor your latency SLO needs and scale ahead of demand, not in response to a full queue.

Readiness must reflect the model, not the process. A pod whose HTTP server is up but whose weights are still loading must not receive traffic. Confirm your runtime's readiness probe succeeds only once the engine is serving, and give the startup probe enough failures to cover the slowest cold start you measured.

Autoscaling on engine signals

The default Kubernetes signals are wrong for LLMs. GPU utilisation sits near 100% as soon as there is any batch, so it cannot distinguish "busy and fine" from "overloaded". CPU is irrelevant. Requests per second ignores that one request may generate ten tokens and the next four thousand. The useful signals come from the engine: how many sequences are running, how many are waiting, and how full the KV cache is.

KServe's documented pattern uses KEDA with a Prometheus query against vLLM's own metrics:

metadata:
  annotations:
    serving.kserve.io/deploymentMode: "Standard"
    serving.kserve.io/autoscalerClass: "keda"
    serving.kserve.io/enable-prometheus-scraping: "true"
spec:
  predictor:
    minReplicas: 1
    maxReplicas: 5
    autoScaling:
      metrics:
        - type: External
          external:
            metric:
              backend: "prometheus"
              serverAddress: "http://prometheus.istio-system.svc.cluster.local:9090"
              query: vllm:num_requests_running
            target:
              type: Value
              value: "2"

The documentation's target of 2 running requests is a demo value for a tiny model. Derive yours: load-test one replica, find the concurrency at which p95 time to first token or inter-token latency crosses your SLO, and set the target to roughly 70% of that so scaling starts with headroom for the minutes a new replica needs. Scope the query to this service's pods with label filters, or it scales on other teams' traffic.

Waiting requests (vllm:num_requests_waiting) is often the better trigger because it is zero in steady state and rises only when the engine cannot admit more work. Pair a scale-up trigger on it with a slow scale-down, since scaling down a replica evicts its prefix cache and every stream it was serving must drain. Scale to zero is possible but rarely worth it for interactive LLMs: the first request after idle pays the full cold start.

Multi-GPU, multi-node and routing

A model that does not fit on one GPU is split with tensor parallelism inside a node: request nvidia.com/gpu: "4" and pass --tensor_parallel_size=4. Keep tensor parallelism inside one node, because every layer does an all-reduce that NVLink handles well and an Ethernet hop does not.

When the model needs more than one node, the pods must start, fail and restart as a unit, and plain Deployments cannot express that. LLMInferenceService's worker section produces a LeaderWorkerSet, which groups a leader and its workers into one replica with a shared lifecycle; the parallelism section declares the tensor, data or expert layout. The prefill section adds a separate prefill pool so long prompts do not stall decode on the same GPUs; the disaggregated serving article explains when that split pays for its KV transfer cost.

Routing is the other half. Round robin across replicas ignores which replica already has a request's prefix in its paged KV cache, so a chat whose next turn lands on a different pod recomputes the whole history. The LLMInferenceService router can use llm-d's endpoint picker to score replicas on load and cache locality. InferenceService gets ordinary Service load balancing, fine for short, stateless prompts.

Rollouts and canaries

Rollouts are where LLM serving differs most from web services. A stream can stay open for a minute, so a pod being replaced must stop receiving new requests, then finish the ones it has. Set terminationGracePeriodSeconds above your longest expected generation and make sure the engine stops admitting work on SIGTERM rather than exiting at once. Roll one replica at a time with maxUnavailable of zero when capacity is tight, because a new replica takes minutes to become ready and surge capacity needs spare GPUs you may not have.

For canaries, Knative mode supports canaryTrafficPercent on the predictor. In Standard mode, deploy the candidate as a second resource and split traffic at the gateway. Compare the canary on time to first token, inter-token latency, error rates and a fixed evaluation prompt set: a new model can be fast and wrong.

Failure modes

  • OOM at startup. --max_model_len or --gpu_memory_utilization too high for the card; the pod restarts forever. Read the engine log for the KV block count it tried to allocate.
  • Ready but not serving. A probe on a process-level endpoint routes traffic to a pod still loading weights; clients see connection resets or 503s after every scale-up.
  • Startup probe kills slow loads. A short failure threshold restarts pods mid-download, which looks like a crash loop with no error in the model code.
  • Scaling on GPU utilisation. It reads near 100% at any load, so the autoscaler either never scales or scales to the maximum.
  • Pending pods. maxReplicas exceeds the GPUs the cluster can actually provide, so new replicas sit in Pending while the dashboard claims capacity is being added.
  • Rollout drops streams. A grace period shorter than generation time cuts responses mid-answer; users see truncated text, not an error code.
  • Model name mismatch. Clients send a model that does not match --model_name and receive a 404 that looks like a routing problem.

Trade-offs

ChoiceGainCostPick it when
InferenceService + huggingface runtimeStable v1beta1 API, OpenAI routes, simpleSingle-node replicas, plain load balancingModel fits one node; prompts mostly short
LLMInferenceServiceMulti-node workers, prefill/decode pools, cache-aware routingAlpha API, more moving partsLarge models, long shared prefixes, disaggregation
Standard mode + KEDAEngine-level scaling signals, no Knative dependencyYou run Prometheus and KEDAMost GPU LLM deployments
Knative modeRevisions, built-in canary, scale to zeroConcurrency scaling assumes fast startsSmall models, bursty internal tools
Local model cache or PVCMinutes off cold startNode storage to manage and invalidateAny autoscaled LLM
Plain vLLM Deployment, no KServeFewest layersYou build routing, scaling and rollout yourselfOne model, one team, fixed capacity

What to do next

  1. Run kubectl get crd | grep serving.kserve.io and kubectl explain on both resource types to learn which API versions and fields your cluster really has.
  2. Deploy the single-GPU InferenceService above with an explicit --max_model_len, call it with curl, and time each cold-start stage from the pod events and engine log.
  3. Load-test one replica, find the concurrency where p95 time to first token breaks your SLO, and set the KEDA target to about 70% of it.
  4. Prove readiness is model-aware: scale from one to two replicas under load and confirm no request reaches the new pod before the weights load.
  5. Set the termination grace period above your longest generation and verify a rolling update finishes every in-flight stream.
  6. Add a local model cache or PVC, re-measure cold start, and lower minReplicas only if the new number fits inside your scale-up headroom.
Key takeaway: KServe is the control plane around an LLM engine: it turns a custom resource into pods, storage, routes and autoscalers while vLLM does the GPU work. Use InferenceService with the huggingface runtime for single-node models and LLMInferenceService when you need multi-node workers, prefill pools or cache-aware routing. Size memory and context from the model, attack cold start with caching, scale on engine queue metrics with headroom, make readiness model-aware, and drain streams on every rollout.