Google Kubernetes Engine is where many teams on Google Cloud run their own training and inference, for one reason: accelerators are scarce and expensive, and Kubernetes gives one place to share them across many jobs and teams. But a GKE cluster that runs web services well will run ML badly by default. Jobs start half their workers and wait forever, data loading starves GPUs, and serving pods autoscale on CPU metrics that mean nothing for a model server.
This article covers the ML-specific layer on top of GKE: accelerator node pools, getting capacity with Dynamic Workload Scheduler and Kueue, running multi-node jobs, feeding data from Cloud Storage, TPU slices, and serving. Flag and label names were checked against the GKE documentation; general cluster mechanics are in GKE architecture in depth.
The ML platform on GKE
Accelerator node pools
Accelerators live in dedicated node pools. You choose the GPU type and count, and let GKE install the NVIDIA driver by setting gpu-driver-version to default or latest (or disabled if you manage the driver yourself, for example with the GPU Operator):
gcloud container node-pools create h100-train \
--cluster ml --location us-central1 --node-locations us-central1-a \
--machine-type a3-highgpu-8g \
--accelerator type=nvidia-h100-80gb,count=8,gpu-driver-version=latest \
--enable-autoscaling --num-nodes 0 --min-nodes 0 --max-nodes 4Pods ask for nvidia.com/gpu in their resource limits and select hardware with the cloud.google.com/gke-accelerator node label, whose values are strings like nvidia-l4, nvidia-h100-80gb or nvidia-b200. When you add a GPU pool to a cluster that already has CPU pools, GKE taints the GPU nodes with nvidia.com/gpu=present:NoSchedule so ordinary pods stay off them. Keep CPU-only system workloads in a separate small pool; a stray logging agent that requests 4 CPUs can make an 8-GPU node unschedulable for a job that needs the whole machine.
Getting capacity: flex-start and Kueue
There are four ways to get GPU capacity, and they trade price against certainty. On-demand is simplest but often unavailable for large GPUs. Reservations guarantee capacity you pay for whether used or not. Spot VMs are much cheaper but can be reclaimed at short notice, so they need frequent checkpoints. Flex-start, part of Dynamic Workload Scheduler, queues your request and provisions all nodes together when capacity frees up, for runs of up to seven days.
| Option | Price | Certainty | Best for |
|---|---|---|---|
| On-demand | List price | Low for large GPUs | Small or short experiments |
| Reservation | Paid whether used or not | High | Teams that keep accelerators busy all day |
| Spot | Lowest | Can be reclaimed any time | Fault-tolerant jobs with frequent checkpoints |
| Flex-start | Discounted | Start time uncertain; runs up to 7 days once started | Scheduled fine-tunes and batch jobs |
Flex-start with queued provisioning uses a node pool created with --enable-queued-provisioning, --flex-start, --enable-autoscaling, --num-nodes=0, --location-policy=ANY and --no-enable-autorepair. Kueue then drives it. You define an AdmissionCheck whose controller is kueue.x-k8s.io/provisioning-request, a ProvisioningRequestConfig with provisioning class queued-provisioning.gke.io managing nvidia.com/gpu, and a ClusterQueue and LocalQueue that use them. Copy those manifests from the GKE page for your Kueue version, because Kueue's API version has been moving. The job itself is plain Kubernetes:
apiVersion: batch/v1
kind: Job
metadata:
name: finetune-8b
labels:
kueue.x-k8s.io/queue-name: dws-local-queue
annotations:
provreq.kueue.x-k8s.io/maxRunDurationSeconds: "86400"
spec:
suspend: true # Kueue unsuspends it once every node is provisioned
parallelism: 2
completions: 2
completionMode: Indexed
template:
spec:
nodeSelector:
cloud.google.com/gke-accelerator: nvidia-h100-80gb
restartPolicy: Never
containers:
- name: trainer
image: us-docker.pkg.dev/my-proj/ml/trainer:2026-10-01
resources:
limits:
nvidia.com/gpu: 8The important property is all-or-nothing: Kueue keeps the job suspended until the ProvisioningRequest succeeds for both nodes, so you never pay for one idle H100 node waiting for its partner. The maximum run duration cannot be changed after the request is made, so set it from measured epoch times plus margin.
Multi-node training jobs
Distributed training needs every worker to find the others and to fail together. An Indexed Job gives stable ranks; JobSet adds a headless service for discovery and treats a group of Jobs as one unit that restarts together:
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
name: fsdp-run
labels:
kueue.x-k8s.io/queue-name: dws-local-queue
spec:
failurePolicy:
maxRestarts: 3 # restart the whole set from the last checkpoint
replicatedJobs:
- name: workers
replicas: 1
template:
spec:
parallelism: 2
completions: 2
completionMode: Indexed
template:
spec: { ... same pod spec as above, running torchrun with rank from the index ... }On A3-class machines, high-bandwidth GPU-to-GPU networking across nodes needs the machine-specific networking setup that GKE documents for each family; without it collective operations fall back to the ordinary VPC path and scale poorly. Measure NCCL all-reduce bandwidth on two nodes before launching a long run. For sharding strategy inside the job see FSDP for large-model training.
Feeding data from Cloud Storage
Training data and checkpoints usually live in Cloud Storage. The Cloud Storage FUSE CSI driver, enabled with the GcsFuseCsiDriver add-on, mounts a bucket as a file system through a sidecar that GKE injects when the pod has the annotation gke-gcsfuse/volumes: "true". Access goes through Workload Identity Federation, so the pod's Kubernetes service account needs read access to the bucket.
metadata:
annotations:
gke-gcsfuse/volumes: "true"
gke-gcsfuse/memory-limit: "4Gi"
spec:
serviceAccountName: trainer-ksa
volumes:
- name: data
csi:
driver: gcsfuse.csi.storage.gke.io
readOnly: true
volumeAttributes:
bucketName: my-training-data
mountOptions: "implicit-dirs,file-cache:max-size-mb:-1"The file cache is what makes multi-epoch training fast: the first epoch reads from the bucket, later epochs from local disk, so back the cache with local SSD on GPU nodes. Many small files are the classic slowdown, so shard datasets into large files such as WebDataset tar shards or Parquet. Size the sidecar with the gke-gcsfuse/cpu-limit and gke-gcsfuse/memory-limit annotations; an under-sized sidecar throttles reads. An unlimited file cache also needs room: size the sidecar's ephemeral storage with gke-gcsfuse/ephemeral-storage-limit or give the cache its own volume. Details are in Cloud Storage FUSE in depth.
TPU slices on GKE
GKE also schedules Cloud TPUs. Pods request google.com/tpu chips and select a slice with the cloud.google.com/gke-tpu-accelerator and cloud.google.com/gke-tpu-topology labels (a topology looks like 2x4; read the accelerator value for your generation from the docs or kubectl get nodes -L cloud.google.com/gke-tpu-accelerator). Multi-host slice node pools are atomic: if GKE cannot create one node none are created, and repairing one node recreates the whole slice, evicting every pod on it. Checkpoint as if each repair were a preemption. Chip generations are compared in TPU v5 and v6.
Serving models
Serving is a different shape: long-running, latency-sensitive, often on smaller GPUs such as L4. Three habits matter. First, autoscale on model-server signals such as queue depth or KV-cache utilisation exported to Cloud Monitoring, not on CPU, which stays low while the GPU saturates. Second, share GPUs for small models: time-sharing is enabled per pool with gpu-sharing-strategy=time-sharing,max-shared-clients-per-gpu=N in the accelerator flag (up to 48), and MIG partitions with gpu-partition-size such as 1g.5gb. Time-sharing gives no memory isolation, so cap each server's GPU memory itself. Third, cut cold starts: model weights of tens of gigabytes dominate startup, so cache them close to the node and keep a minimum replica count during business hours.
apiVersion: apps/v1
kind: Deployment
metadata:
name: assistant-vllm
spec:
replicas: 2
selector: { matchLabels: { app: assistant-vllm } }
template:
metadata: { labels: { app: assistant-vllm } }
spec:
nodeSelector:
cloud.google.com/gke-accelerator: nvidia-l4
containers:
- name: vllm
image: vllm/vllm-openai:v0.x.y # pin an exact tag you have tested
args: ["--model", "/models/assistant-2026-10-01", "--max-model-len", "8192"]
resources:
limits: { nvidia.com/gpu: 1 }
readinessProbe:
httpGet: { path: /health, port: 8000 }
periodSeconds: 10The readiness probe matters more than it looks: a vLLM pod takes minutes to load weights, and without the probe the Service sends traffic to a server that cannot answer yet. Point the model path at a cached copy rather than pulling from the Hub at start.
For LLMs, GKE Inference Gateway adds model-aware load balancing on top of Gateway API, routing requests by model server load and LoRA adapter. Engine configuration is covered in vLLM on GPUs.
Observability and cost control
Accelerators are the most expensive thing in the cluster, so measure them directly. Node CPU and memory graphs say nothing about whether a GPU is busy. Collect NVIDIA DCGM metrics (GPU utilisation, memory used, SM activity, tensor-core activity, XID errors), either through GKE's managed collection where your version supports it or by running the DCGM exporter yourself, and put them next to training throughput in the same dashboard.
Three alerts pay for themselves. An idle-GPU alert fires when a node with allocated GPUs shows near-zero utilisation for 30 minutes, which catches hung jobs and forgotten notebooks. A stalled-progress alert fires when a training job's step counter stops advancing, which catches NCCL hangs that leave pods Running. A capacity-wait alert fires when a Kueue workload has been pending longer than its team expects, so someone can switch it to another capacity option.
For cost, label every workload with team and project, enable GKE cost allocation, and give each team a Kueue ClusterQueue with an explicit GPU quota, so contention is decided by policy rather than by whoever submits first. Borrowing between queues lets idle quota be used without giving up the guarantee.
Worked example: weekly fine-tune and serve
A team fine-tunes an 8B model weekly and serves it to an internal assistant. The training pool is two a3-highgpu-8g nodes under flex-start; the weekly Job above waits in Kueue, typically between minutes and a few hours, then runs FSDP across 16 H100s for about six hours, reading 300 GB of sharded data through GCS FUSE with a local-SSD cache and checkpointing to the bucket every 30 minutes. The run duration is capped at 24 hours. Nodes scale back to zero when the job finishes, so the pool costs nothing between runs.
Serving runs on an L4 pool with autoscaling on vLLM's waiting-request metric, minimum two replicas. An evaluation Job gates each new checkpoint; only a passing one is rolled out by updating the Deployment's model path. When they added a second, smaller classifier model, they put it on a time-shared L4 with explicit memory limits instead of buying another node.
Failure modes
- Partial scheduling. Half the workers start and block in NCCL init while the rest wait for nodes. Use Kueue with all-or-nothing admission rather than bare Jobs.
- Starved GPUs. Utilisation hovers at 30 percent because data loading cannot keep up. Profile reads, enable the FUSE file cache and use large shards.
- Run duration too short. Flex-start nodes are reclaimed at the max run duration. Checkpoint well inside it and resume.
- Slice-wide eviction. One TPU host repair restarts the slice. Make restart from checkpoint routine.
- Driver mismatch. A container built for a newer CUDA fails on the node's driver. Pin
gpu-driver-versiondeliberately and test images on the pool.
Trade-offs
GKE gives control and sharing at the cost of operating a cluster; Vertex AI training and prediction remove that work but constrain scheduling, images and cost tricks. Flex-start lowers the price of large GPUs and avoids holding idle nodes, but start time is uncertain, which suits weekly fine-tunes rather than urgent ones. Reservations buy certainty for teams that keep accelerators busy. GPU sharing raises utilisation for small models but trades away isolation, so keep latency-critical servers on whole GPUs.
What to do next
- Create separate node pools for CPU system work, GPU training, and serving, with the GPU taint left in place.
- Install Kueue, define queues per team, and route every training Job through a LocalQueue.
- Try one flex-start queued-provisioning pool and set max run duration from measured epoch times.
- Move datasets into large shards in Cloud Storage and mount them with GCS FUSE and a local-SSD file cache.
- Measure two-node NCCL bandwidth before the first long multi-node run.
- Checkpoint to the bucket on a schedule and rehearse a restart from the last checkpoint.
- Autoscale serving on model-server metrics and decide which small models can share a GPU.