A Kubernetes node with a GPU in it is not yet a GPU node. Before a pod can request nvidia.com/gpu: 1 and get a working CUDA device, the host needs a kernel driver that matches its kernel, a container runtime that knows how to inject device files and driver libraries into containers, a device plugin that tells the kubelet how many GPUs exist, and ideally labels describing the GPU model and a metrics exporter. Doing that by hand means baking it all into a machine image and rebuilding the image whenever the kernel, driver or CUDA version moves. The NVIDIA GPU Operator does it from inside the cluster instead: each layer runs as a DaemonSet, and one custom resource describes the whole stack.
This article explains what the operator installs and in which order, how a node goes from bare to schedulable, how to control it per node, how it hooks into GPU sharing, how it upgrades drivers without killing running jobs, and what breaks. Version-specific details here were checked against NVIDIA's GPU Operator documentation in October 2026, when the current chart was v26.7.1; pin your own version and read its release notes. The general shape of a reconciling operator is covered in the Kubernetes operator pattern article.
The software stack a GPU node needs
Think of a GPU node as five layers, each depending on the one below. The operator calls each managed component an operand.
| Operand | What it does | Depends on |
|---|---|---|
| Node Feature Discovery (NFD) | labels nodes by hardware, including the NVIDIA PCI vendor ID, so the other DaemonSets know where to run | nothing |
| Driver container | builds or loads the NVIDIA kernel modules on the host and exposes the driver libraries | a supported kernel and OS |
| Container toolkit | configures containerd or CRI-O so containers receive GPU devices; with CDI, writes device specs instead of swapping the default runtime | driver |
| Operator validator | checks driver and toolkit, and runs a short CUDA workload (the nvidia-cuda-validator pod) | driver, toolkit |
| Device plugin | registers nvidia.com/gpu with the kubelet and hands out devices to pods | validation passed |
| GPU Feature Discovery | adds GPU labels such as nvidia.com/gpu.product | validation passed |
| DCGM exporter | publishes GPU metrics for Prometheus | validation passed |
| MIG manager | applies MIG layouts requested by a node label, on MIG-capable GPUs | driver |
Ordering is the part people underestimate. If the device plugin started before the driver was loaded, it would advertise zero GPUs, or worse, pods would schedule and then fail inside CUDA initialisation. The operator therefore holds the upper layers back until the validator reports the lower layers ready, and a node only starts advertising nvidia.com/gpu capacity after that. When something is stuck, the validator is the first pod to look at.
Installing it
Installation is a Helm chart into its own namespace. The operator pods need privileged access to the host, so if Pod Security Admission is enforced, label the namespace first.
kubectl create namespace gpu-operator
kubectl label --overwrite ns gpu-operator pod-security.kubernetes.io/enforce=privileged
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia && helm repo update
helm install --wait --generate-name -n gpu-operator \
nvidia/gpu-operator --version=v26.7.1
# nodes whose image already ships a driver and toolkit:
# --set driver.enabled=false --set toolkit.enabled=falseAfter a few minutes, kubectl get pods -n gpu-operator should show the operator, the three NFD pods, and per GPU node a driver, toolkit, device plugin, feature discovery, DCGM exporter and validator pod, all Running, plus a CUDA validator pod in Completed. Then check that the node advertises capacity:
kubectl get nodes -o custom-columns=NAME:.metadata.name,GPUS:.status.allocatable.nvidia\.com/gpu
kubectl get clusterpolicy # should report the policy as readyTwo chart defaults are worth knowing. cdi.enabled is true by default from release 25.10.0: the operator no longer makes the nvidia runtime the default runtime handler and relies on the container runtime's native Container Device Interface support instead. And dcgmExporter.enabled is on, so metrics appear without extra work; what to do with them is covered in the DCGM article.
ClusterPolicy as the source of truth
Helm values become a single cluster-scoped custom resource, a ClusterPolicy named cluster-policy. The operator reconciles it continuously: it compares the DaemonSets that should exist for the spec with those that do, and fixes the difference. This has two consequences that surprise people. First, editing an operand DaemonSet by hand does not stick; the next reconcile reverts it. Change the ClusterPolicy (or the Helm values, then upgrade the release) instead. Second, the ClusterPolicy's status is the single readiness signal for the whole stack, so alert on it rather than on individual pods.
For clusters with different driver needs per node pool, the operator also offers an NVIDIADriver custom resource that manages the driver for a subset of nodes, so one pool can run a newer branch while another stays put. Treat it as an advanced option and keep one driver version per pool, which keeps the debugging matrix small.
Worked example: a node joins the cluster
Follow one node from cluster join to its first training pod. The autoscaler adds a VM with a single GPU-capable machine type. NFD's worker labels it with PCI information within seconds, which makes it match the operand DaemonSets' node selectors. The driver pod starts first and loads the kernel modules; this is the slow step, often a few minutes when modules must be compiled for the running kernel. The toolkit pod then configures the runtime. The validator confirms the driver responds, confirms the toolkit is in place and runs a tiny CUDA program to completion. Only then do the device plugin, feature discovery and exporter start, and the device plugin registers the GPUs with the kubelet. The scheduler now sees allocatable nvidia.com/gpu and the pending training pod lands.
The practical lesson is the time budget. A node that boots in one minute may take several more before it can host GPU work, and a cluster autoscaler or job queue that assumes otherwise will mark scale-ups as failures. Pre-built driver images or nodes with a pre-installed driver shorten this, at the cost of the flexibility discussed below. If you gate admission with Kueue, account for this delay in its provisioning settings.
Per-node control
Not every node should get every operand. The operator honours per-node labels:
# keep all GPU Operator operands off this node (e.g. a GPU node managed another way)
kubectl label nodes gpu-node-17 nvidia.com/gpu.deploy.operands=false
# run everything except the driver container (driver baked into this node's image)
kubectl label nodes gpu-node-18 nvidia.com/gpu.deploy.driver=falseThis is how mixed fleets stay sane: a pool built from a vendor image with a driver already installed gets the driver label, and the operator still provides the toolkit, plugin and monitoring. Apply these labels through the node pool's configuration rather than by hand, or a replaced node comes back without them and the operator tries to install a second driver on top of the first.
Time-slicing and MIG through the operator
GPU sharing is configured through the operator rather than separately. Time-slicing is a device plugin setting: put the configuration in a ConfigMap in the operator namespace and point the ClusterPolicy at it.
apiVersion: v1
kind: ConfigMap
metadata:
name: time-slicing-config
namespace: gpu-operator
data:
any: |-
version: v1
flags:
migStrategy: none
sharing:
timeSlicing:
renameByDefault: false
failRequestsGreaterThanOne: false
resources:
- name: nvidia.com/gpu
replicas: 4At install time pass --set devicePlugin.config.name=time-slicing-config; on a running cluster patch the same field in the ClusterPolicy, and set a default key or the configuration will not apply to nodes automatically. With four replicas, each physical GPU is advertised as four nvidia.com/gpu units, and by default the product label gains a -SHARED suffix so workloads can select or avoid shared nodes. There is no memory isolation between those four pods; see the time-slicing article for when that is acceptable.
MIG is driven by a node label that the MIG manager watches:
kubectl label nodes gpu-node-21 nvidia.com/mig.config=all-1g.10gb --overwrite
kubectl get node gpu-node-21 -o jsonpath='{.metadata.labels.nvidia\.com/mig\.config\.state}'
# pending while it works, success when doneChanging a MIG layout terminates GPU pods on that node, so drain work first or let your scheduler move it. Layout planning is covered in the multi-instance GPU deployment article.
Driver upgrades without lost jobs
Upgrading a GPU driver means unloading kernel modules, which cannot happen while any process holds the GPU. The operator's driver upgrade controller turns that into a rolling procedure. You change one field and it walks each node through a state machine recorded in the nvidia.com/gpu-driver-upgrade-state node label.
# ClusterPolicy (or Helm values under driver.upgradePolicy)
driver:
upgradePolicy:
autoUpgrade: true
maxParallelUpgrades: 1 # nodes upgraded at once
maxUnavailable: 25% # cap on nodes out of service
waitForCompletion:
timeoutSeconds: 0 # 0 waits forever for matching pods
podSelector: "app=trainer" # pods allowed to finish first
gpuPodDeletion:
force: false
timeoutSeconds: 300
deleteEmptyDir: false
drain:
enable: false # full drain only if pod deletion is not enoughkubectl patch clusterpolicies.nvidia.com/cluster-policy --type=json \
-p='[{"op":"replace","path":"/spec/driver/version","value":"<new-version>"}]'
kubectl get nodes -L nvidia.com/gpu-driver-upgrade-state -wEach node moves through upgrade-required, cordon-required, wait-for-jobs-required, pod-deletion-required, pod-restart-required, validation-required, uncordon-required and finally upgrade-done, with drain-required as an optional step and upgrade-failed when something goes wrong. The settings that matter are the waiting ones. With waitForCompletion pointed at your training pods and a timeout of zero, a node running a three-day job simply waits; the upgrade takes longer but loses no work. With a short timeout and force, the upgrade is fast and jobs without checkpoints restart from scratch. Pick deliberately, and upgrade a canary pool first.
Failure modes
| Symptom | Likely cause | First check |
|---|---|---|
| driver pod in CrashLoopBackOff | kernel headers unavailable for the node kernel, unsupported OS, or nouveau loaded | driver pod logs; blacklist nouveau in the node image |
| modules refuse to load | Secure Boot rejects unsigned modules | use signed or pre-installed drivers on Secure Boot nodes |
| validator stuck in Init | driver or toolkit not ready | logs of the init containers in the validator pod |
| node shows 0 allocatable GPUs | device plugin not started or not registered | device plugin logs; ClusterPolicy status |
| two drivers fighting | pre-installed driver plus driver container | set nvidia.com/gpu.deploy.driver=false on that pool |
| pods start but CUDA fails | runtime not injecting devices (CDI or runtime class mismatch) | toolkit pod logs; container runtime config |
| upgrade never finishes | waitForCompletion waiting on a long job | node upgrade-state label; the selector's matching pods |
| manual DaemonSet edit vanished | reconcile reverted it | make the change in ClusterPolicy instead |
Most incidents come down to two things: the host kernel or image changed under the driver container, or two systems both think they own the driver. Track kernel versions alongside driver versions and keep ownership explicit per pool.
Trade-offs
The big decision is who owns the driver. Operator-managed drivers make the node image generic and turn a driver upgrade into a rolling, observable cluster operation, but they add minutes to node start-up and fail when the kernel moves ahead of what the driver container supports. Pre-installed drivers in the machine image boot faster and are fixed at build time, but each driver change means a new image and a node pool rotation. Managed Kubernetes services often ship their own GPU drivers and device plugin; running the full operator there can duplicate components, so check the provider's guidance and usually disable the driver and toolkit operands.
The operator is not a scheduler. It makes GPUs appear as resources and keeps the software stack healthy; quotas, gang scheduling and fair sharing belong to tools such as Kueue or a dedicated GPU scheduler on top.
What to do next
- Decide driver ownership per node pool and record it as node pool labels.
- Install the chart with a pinned version in a staging cluster and time a node from join to first allocatable GPU.
- Alert on ClusterPolicy status and on nodes whose allocatable
nvidia.com/gpudrops below their hardware count. - Configure
driver.upgradePolicywithwaitForCompletionmatching your long-running jobs, then rehearse a driver upgrade on one canary node. - If you share GPUs, put time-slicing or MIG configuration under version control as ConfigMaps and labels, not ad hoc edits.
- Wire the DCGM exporter into dashboards before the first production job, not after the first incident.