A Kubernetes node with a GPU in it is not yet a GPU node. Before a pod can request nvidia.com/gpu: 1 and get a working CUDA device, the host needs a kernel driver that matches its kernel, a container runtime that knows how to inject device files and driver libraries into containers, a device plugin that tells the kubelet how many GPUs exist, and ideally labels describing the GPU model and a metrics exporter. Doing that by hand means baking it all into a machine image and rebuilding the image whenever the kernel, driver or CUDA version moves. The NVIDIA GPU Operator does it from inside the cluster instead: each layer runs as a DaemonSet, and one custom resource describes the whole stack.

This article explains what the operator installs and in which order, how a node goes from bare to schedulable, how to control it per node, how it hooks into GPU sharing, how it upgrades drivers without killing running jobs, and what breaks. Version-specific details here were checked against NVIDIA's GPU Operator documentation in October 2026, when the current chart was v26.7.1; pin your own version and read its release notes. The general shape of a reconciling operator is covered in the Kubernetes operator pattern article.

The software stack a GPU node needs

Think of a GPU node as five layers, each depending on the one below. The operator calls each managed component an operand.

OperandWhat it doesDepends on
Node Feature Discovery (NFD)labels nodes by hardware, including the NVIDIA PCI vendor ID, so the other DaemonSets know where to runnothing
Driver containerbuilds or loads the NVIDIA kernel modules on the host and exposes the driver librariesa supported kernel and OS
Container toolkitconfigures containerd or CRI-O so containers receive GPU devices; with CDI, writes device specs instead of swapping the default runtimedriver
Operator validatorchecks driver and toolkit, and runs a short CUDA workload (the nvidia-cuda-validator pod)driver, toolkit
Device pluginregisters nvidia.com/gpu with the kubelet and hands out devices to podsvalidation passed
GPU Feature Discoveryadds GPU labels such as nvidia.com/gpu.productvalidation passed
DCGM exporterpublishes GPU metrics for Prometheusvalidation passed
MIG managerapplies MIG layouts requested by a node label, on MIG-capable GPUsdriver

Ordering is the part people underestimate. If the device plugin started before the driver was loaded, it would advertise zero GPUs, or worse, pods would schedule and then fail inside CUDA initialisation. The operator therefore holds the upper layers back until the validator reports the lower layers ready, and a node only starts advertising nvidia.com/gpu capacity after that. When something is stuck, the validator is the first pod to look at.

What the GPU Operator puts on one GPU node, in dependency ordergpu-operatorreconciles ClusterPolicyNode Feature Discoverylabels: PCI vendor 10deGPU node foundDaemonSets scheduledonly on GPU-labelled nodes1. driverkernel modules2. container toolkitruntime + CDI specs3. validatordriver, toolkit, CUDA test4a. device pluginadvertises nvidia.com/gpu4b. GPU feature discoveryproduct, count labels4c. DCGM exportermetrics for Prometheus4d. MIG managerapplies mig.configreadykubeletschedules pods requesting nvidia.com/gpu
Operand order on one node. Layers 4a to 4c wait for the validator, the MIG manager needs only the driver; the kubelet only sees nvidia.com/gpu capacity once the device plugin registers.

Installing it

Installation is a Helm chart into its own namespace. The operator pods need privileged access to the host, so if Pod Security Admission is enforced, label the namespace first.

kubectl create namespace gpu-operator
kubectl label --overwrite ns gpu-operator pod-security.kubernetes.io/enforce=privileged

helm repo add nvidia https://helm.ngc.nvidia.com/nvidia && helm repo update
helm install --wait --generate-name -n gpu-operator \
  nvidia/gpu-operator --version=v26.7.1

# nodes whose image already ships a driver and toolkit:
#   --set driver.enabled=false --set toolkit.enabled=false

After a few minutes, kubectl get pods -n gpu-operator should show the operator, the three NFD pods, and per GPU node a driver, toolkit, device plugin, feature discovery, DCGM exporter and validator pod, all Running, plus a CUDA validator pod in Completed. Then check that the node advertises capacity:

kubectl get nodes -o custom-columns=NAME:.metadata.name,GPUS:.status.allocatable.nvidia\.com/gpu
kubectl get clusterpolicy        # should report the policy as ready

Two chart defaults are worth knowing. cdi.enabled is true by default from release 25.10.0: the operator no longer makes the nvidia runtime the default runtime handler and relies on the container runtime's native Container Device Interface support instead. And dcgmExporter.enabled is on, so metrics appear without extra work; what to do with them is covered in the DCGM article.

ClusterPolicy as the source of truth

Helm values become a single cluster-scoped custom resource, a ClusterPolicy named cluster-policy. The operator reconciles it continuously: it compares the DaemonSets that should exist for the spec with those that do, and fixes the difference. This has two consequences that surprise people. First, editing an operand DaemonSet by hand does not stick; the next reconcile reverts it. Change the ClusterPolicy (or the Helm values, then upgrade the release) instead. Second, the ClusterPolicy's status is the single readiness signal for the whole stack, so alert on it rather than on individual pods.

For clusters with different driver needs per node pool, the operator also offers an NVIDIADriver custom resource that manages the driver for a subset of nodes, so one pool can run a newer branch while another stays put. Treat it as an advanced option and keep one driver version per pool, which keeps the debugging matrix small.

Worked example: a node joins the cluster

Follow one node from cluster join to its first training pod. The autoscaler adds a VM with a single GPU-capable machine type. NFD's worker labels it with PCI information within seconds, which makes it match the operand DaemonSets' node selectors. The driver pod starts first and loads the kernel modules; this is the slow step, often a few minutes when modules must be compiled for the running kernel. The toolkit pod then configures the runtime. The validator confirms the driver responds, confirms the toolkit is in place and runs a tiny CUDA program to completion. Only then do the device plugin, feature discovery and exporter start, and the device plugin registers the GPUs with the kubelet. The scheduler now sees allocatable nvidia.com/gpu and the pending training pod lands.

The practical lesson is the time budget. A node that boots in one minute may take several more before it can host GPU work, and a cluster autoscaler or job queue that assumes otherwise will mark scale-ups as failures. Pre-built driver images or nodes with a pre-installed driver shorten this, at the cost of the flexibility discussed below. If you gate admission with Kueue, account for this delay in its provisioning settings.

Per-node control

Not every node should get every operand. The operator honours per-node labels:

# keep all GPU Operator operands off this node (e.g. a GPU node managed another way)
kubectl label nodes gpu-node-17 nvidia.com/gpu.deploy.operands=false

# run everything except the driver container (driver baked into this node's image)
kubectl label nodes gpu-node-18 nvidia.com/gpu.deploy.driver=false

This is how mixed fleets stay sane: a pool built from a vendor image with a driver already installed gets the driver label, and the operator still provides the toolkit, plugin and monitoring. Apply these labels through the node pool's configuration rather than by hand, or a replaced node comes back without them and the operator tries to install a second driver on top of the first.

Time-slicing and MIG through the operator

GPU sharing is configured through the operator rather than separately. Time-slicing is a device plugin setting: put the configuration in a ConfigMap in the operator namespace and point the ClusterPolicy at it.

apiVersion: v1
kind: ConfigMap
metadata:
  name: time-slicing-config
  namespace: gpu-operator
data:
  any: |-
    version: v1
    flags:
      migStrategy: none
    sharing:
      timeSlicing:
        renameByDefault: false
        failRequestsGreaterThanOne: false
        resources:
          - name: nvidia.com/gpu
            replicas: 4

At install time pass --set devicePlugin.config.name=time-slicing-config; on a running cluster patch the same field in the ClusterPolicy, and set a default key or the configuration will not apply to nodes automatically. With four replicas, each physical GPU is advertised as four nvidia.com/gpu units, and by default the product label gains a -SHARED suffix so workloads can select or avoid shared nodes. There is no memory isolation between those four pods; see the time-slicing article for when that is acceptable.

MIG is driven by a node label that the MIG manager watches:

kubectl label nodes gpu-node-21 nvidia.com/mig.config=all-1g.10gb --overwrite
kubectl get node gpu-node-21 -o jsonpath='{.metadata.labels.nvidia\.com/mig\.config\.state}'
# pending while it works, success when done

Changing a MIG layout terminates GPU pods on that node, so drain work first or let your scheduler move it. Layout planning is covered in the multi-instance GPU deployment article.

Driver upgrades without lost jobs

Upgrading a GPU driver means unloading kernel modules, which cannot happen while any process holds the GPU. The operator's driver upgrade controller turns that into a rolling procedure. You change one field and it walks each node through a state machine recorded in the nvidia.com/gpu-driver-upgrade-state node label.

# ClusterPolicy (or Helm values under driver.upgradePolicy)
driver:
  upgradePolicy:
    autoUpgrade: true
    maxParallelUpgrades: 1        # nodes upgraded at once
    maxUnavailable: 25%           # cap on nodes out of service
    waitForCompletion:
      timeoutSeconds: 0           # 0 waits forever for matching pods
      podSelector: "app=trainer"  # pods allowed to finish first
    gpuPodDeletion:
      force: false
      timeoutSeconds: 300
      deleteEmptyDir: false
    drain:
      enable: false               # full drain only if pod deletion is not enough
kubectl patch clusterpolicies.nvidia.com/cluster-policy --type=json \
  -p='[{"op":"replace","path":"/spec/driver/version","value":"<new-version>"}]'
kubectl get nodes -L nvidia.com/gpu-driver-upgrade-state -w

Each node moves through upgrade-required, cordon-required, wait-for-jobs-required, pod-deletion-required, pod-restart-required, validation-required, uncordon-required and finally upgrade-done, with drain-required as an optional step and upgrade-failed when something goes wrong. The settings that matter are the waiting ones. With waitForCompletion pointed at your training pods and a timeout of zero, a node running a three-day job simply waits; the upgrade takes longer but loses no work. With a short timeout and force, the upgrade is fast and jobs without checkpoints restart from scratch. Pick deliberately, and upgrade a canary pool first.

Failure modes

SymptomLikely causeFirst check
driver pod in CrashLoopBackOffkernel headers unavailable for the node kernel, unsupported OS, or nouveau loadeddriver pod logs; blacklist nouveau in the node image
modules refuse to loadSecure Boot rejects unsigned modulesuse signed or pre-installed drivers on Secure Boot nodes
validator stuck in Initdriver or toolkit not readylogs of the init containers in the validator pod
node shows 0 allocatable GPUsdevice plugin not started or not registereddevice plugin logs; ClusterPolicy status
two drivers fightingpre-installed driver plus driver containerset nvidia.com/gpu.deploy.driver=false on that pool
pods start but CUDA failsruntime not injecting devices (CDI or runtime class mismatch)toolkit pod logs; container runtime config
upgrade never finisheswaitForCompletion waiting on a long jobnode upgrade-state label; the selector's matching pods
manual DaemonSet edit vanishedreconcile reverted itmake the change in ClusterPolicy instead

Most incidents come down to two things: the host kernel or image changed under the driver container, or two systems both think they own the driver. Track kernel versions alongside driver versions and keep ownership explicit per pool.

Trade-offs

The big decision is who owns the driver. Operator-managed drivers make the node image generic and turn a driver upgrade into a rolling, observable cluster operation, but they add minutes to node start-up and fail when the kernel moves ahead of what the driver container supports. Pre-installed drivers in the machine image boot faster and are fixed at build time, but each driver change means a new image and a node pool rotation. Managed Kubernetes services often ship their own GPU drivers and device plugin; running the full operator there can duplicate components, so check the provider's guidance and usually disable the driver and toolkit operands.

The operator is not a scheduler. It makes GPUs appear as resources and keeps the software stack healthy; quotas, gang scheduling and fair sharing belong to tools such as Kueue or a dedicated GPU scheduler on top.

What to do next

  1. Decide driver ownership per node pool and record it as node pool labels.
  2. Install the chart with a pinned version in a staging cluster and time a node from join to first allocatable GPU.
  3. Alert on ClusterPolicy status and on nodes whose allocatable nvidia.com/gpu drops below their hardware count.
  4. Configure driver.upgradePolicy with waitForCompletion matching your long-running jobs, then rehearse a driver upgrade on one canary node.
  5. If you share GPUs, put time-slicing or MIG configuration under version control as ConfigMaps and labels, not ad hoc edits.
  6. Wire the DCGM exporter into dashboards before the first production job, not after the first incident.
Key takeaway: The GPU Operator runs the GPU software stack as ordered DaemonSets driven by one ClusterPolicy: driver, toolkit, validation, then device plugin and monitoring. Decide who owns the driver per pool, change the policy rather than the DaemonSets, and configure the upgrade controller to wait for jobs.