Kueue is the Kubernetes SIG Scheduling project for job queueing. It answers a question the default scheduler never asks: should this job start at all right now, given the quota each team is owed? The kube-scheduler places pods one at a time as soon as they exist. For GPU training that is the wrong behaviour, because a 32-GPU job with 20 pods placed and 12 pending holds 20 expensive GPUs doing nothing, and two such jobs can deadlock each other. Kueue keeps whole jobs suspended until their full request fits in quota, then releases them together.
This article explains the objects and the admission loop, then works through quota, borrowing, preemption, flavors, readiness timeouts and topology, with YAML for the current kueue.x-k8s.io/v1beta2 API. Kueue moves quickly: release v0.20 stopped serving v1beta1, so older blog posts and manifests need migrating. Field names here were checked against the upstream API reference; check the reference for your installed version before copying anything.
The object model and admission loop
Five objects carry the model. A ResourceFlavor names a kind of capacity, usually a GPU model or a pricing class, and maps it to nodes through nodeLabels, nodeTaints and tolerations. A ClusterQueue is cluster-scoped and holds quota: for each flavor and resource, a nominalQuota. A LocalQueue lives in a namespace and points at one ClusterQueue, so teams submit to their own namespace without seeing cluster-level policy. A cohort groups ClusterQueues that may borrow each other's unused quota. A Workload is Kueue's internal record for one job: its pod sets, their counts and resource requests, its priority and its admission status.
The integration point is the suspend flag that batch Jobs, JobSet, Kubeflow training jobs, Ray jobs and others share. A webhook suspends any job carrying the kueue.x-k8s.io/queue-name label, Kueue creates the matching Workload, and when the Workload is admitted Kueue injects the flavor's node selectors into the pod template and unsuspends the job. Pods only then exist and the normal scheduler binds them.
Keep one consequence in mind throughout: admission is quota accounting over resource requests, not a proof that nodes have room. A job can be admitted and still have pods pending because free GPUs are scattered across nodes, or a node is cordoned. The readiness timeout and topology features below exist to close that gap.
A GPU setup in YAML
A minimal GPU setup has one flavor per GPU type, a ClusterQueue per team in a shared cohort and a LocalQueue per team namespace.
apiVersion: kueue.x-k8s.io/v1beta2
kind: ResourceFlavor
metadata:
name: h100
spec:
nodeLabels:
nvidia.com/gpu.product: NVIDIA-H100-80GB-HBM3
---
apiVersion: kueue.x-k8s.io/v1beta2
kind: ClusterQueue
metadata:
name: team-a
spec:
namespaceSelector: {}
cohortName: research
queueingStrategy: BestEffortFIFO
preemption:
reclaimWithinCohort: Any
withinClusterQueue: LowerPriority
resourceGroups:
- coveredResources: ["cpu", "memory", "nvidia.com/gpu"]
flavors:
- name: h100
resources:
- name: "cpu"
nominalQuota: 384
- name: "memory"
nominalQuota: 3Ti
- name: "nvidia.com/gpu"
nominalQuota: 16
borrowingLimit: 16
lendingLimit: 12
---
apiVersion: kueue.x-k8s.io/v1beta2
kind: LocalQueue
metadata:
namespace: team-a
name: training
spec:
clusterQueue: team-aThe node label value depends on how your GPU operator or feature discovery labels nodes; read it from kubectl get nodes --show-labels rather than copying this one. Every resource a job requests must be covered by some resource group, otherwise the job is never admitted, so include CPU and memory alongside GPUs. A training job then opts in with a single label:
metadata:
labels:
kueue.x-k8s.io/queue-name: training
Queueing order and gang admission
Inside a ClusterQueue, pending Workloads are ordered by priority and then by creation time. StrictFIFO always tries the head first and blocks everything behind it if the head does not fit. BestEffortFIFO, the default, lets a smaller job behind a blocked big one be admitted if it fits. BestEffortFIFO gives better utilisation but can starve a large job indefinitely while a stream of small ones keeps slipping past it. StrictFIFO is fair to the big job but leaves GPUs idle while it waits. A common compromise is BestEffortFIFO plus a higher priority class for jobs that have waited too long, applied by your own controller or by hand.
Each scheduling cycle Kueue takes the heads of the ClusterQueues, tries to fit each one into its own quota and then into borrowable cohort capacity, and decides whether preemption could make room. Admission is all or nothing per Workload: either every pod set gets a flavor assignment for its full request, or nothing is admitted. That is the gang-admission property that protects distributed training.
Cohorts, borrowing and preemption
Borrowing is what makes quota efficient. When ClusterQueues share a cohort, unused nominal quota in one can be used by another. borrowingLimit caps how much a queue may take beyond its nominal quota, and lendingLimit caps how much of its own quota it will lend, which protects a floor that is always available immediately. Borrowed capacity is not owned: when the lender needs it back, reclaimWithinCohort decides whether the lender may preempt borrowers. Values are Never, LowerPriority and Any. withinClusterQueue controls preemption among a queue's own workloads by priority, and borrowWithinCohort controls whether a queue that is itself borrowing may preempt others.
Preemption evicts whole Workloads. The evicted job is suspended, its pods are deleted and it goes back into the queue. For training this means lost work since the last checkpoint, so preemption policy and checkpoint frequency must be designed together. A job that checkpoints every two hours should not run on borrowed capacity that can be reclaimed at any minute.
Put a number on it. If reclaim can arrive at any time, the expected work lost per eviction is half the checkpoint interval times the GPUs held. A 24-GPU job checkpointing every two hours loses on average one hour of 24 GPUs, so 24 GPU-hours, plus restart time for image pull, data loading and collective setup. Count evictions per week from Kueue's metrics and multiply: if that figure is a noticeable share of the borrowed GPU-hours, borrowing is costing more than it earns for that team, and the fix is shorter checkpoint intervals, a lower borrowingLimit or a higher lendingLimit floor on the lender so fewer reclaims are needed.
Worked example: a borrower gets evicted
Two teams share a cohort. Each ClusterQueue has 16 H100s nominal, lendingLimit 12 and borrowingLimit 16, with reclaimWithinCohort: Any. The table follows the GPU counts.
| Event | Team A using | Team B using | Outcome |
|---|---|---|---|
| B runs a 4-GPU job | 0 | 4 | admitted on B quota |
| A submits a 24-GPU job | 24 | 4 | 16 own plus 8 borrowed from B |
| B submits a 12-GPU job | 24 | 4 | B needs its quota back |
| Kueue reclaims for B | 0 | 16 | A job evicted whole, requeued |
| A job waits | 0 | 16 | A has 16, job needs 24 |
Team B's new job needs 12, its own quota has 16 minus 4 in use, so it is entitled to 12, but 8 of those are lent to A. Because A borrowed, A's 24-GPU job is the preemption candidate and the whole job is evicted, freeing 24 GPUs to reclaim 8. A's job now waits until B's usage drops enough to lend 8 again. Two remedies: reduce A's borrowingLimit so it only admits jobs that mostly fit its own quota, or split the work into jobs no larger than nominal quota. This interaction, large borrowers evicted wholesale by small reclaims, is the single most common surprise in a new Kueue deployment.
Flavors and fungibility
A ClusterQueue can list several flavors in one resource group, for example reserved H100s first and spot H100s second. Kueue tries flavors in order. flavorFungibility decides when to stop searching: whenCanBorrow and whenCanPreempt take MayStopSearch or TryNextFlavor. With TryNextFlavor on borrowing, a job prefers an unborrowed later flavor over borrowing on the first one, which keeps cohort lending free for teams that have no other flavor. Flavors also carry taints and tolerations, so spot nodes can repel pods that were not admitted onto the spot flavor. For partitioned GPUs, a MIG profile is just another extended resource name you can cover in a resource group.
Closing the gap between admission and placement
Because admission does not check node fit, configure the readiness timeout in the Kueue controller Configuration. The fragment below goes in the controller manager configuration, not in a queue.
waitForPodsReady:
timeout: 10m
recoveryTimeout: 5m
blockAdmission: false
requeuingStrategy:
timestamp: Eviction
backoffLimitCount: 5
backoffBaseSeconds: 60
backoffMaxSeconds: 3600If an admitted job's pods are not all ready within timeout, Kueue cancels admission, suspends the job and requeues it with exponential backoff, so a job that cannot be placed does not hold quota forever. recoveryTimeout applies the same rule to a running job that loses a pod. Pick the timeout from your slowest legitimate start: large image pulls and model downloads can take many minutes, and a timeout shorter than that causes eviction loops.
Topology-aware scheduling, beta and on by default since v0.14, attacks the same gap. A Topology object lists node label levels such as block and rack, a ResourceFlavor references it with topologyName, and a pod template annotation such as kueue.x-k8s.io/podset-required-topology asks that all pods land within one domain at that level. Kueue then admits only if it has found concrete room inside one domain, which matters for collective-heavy training where crossing racks costs bandwidth.
Operating Kueue
Operating Kueue is mostly watching queues. Useful habits:
kubectl get workloads -Ashows admission status and the reason a job is pending; check it before debugging pods, because a suspended job has no pods to inspect.- Export Kueue's Prometheus metrics for pending workloads, admission wait time and quota usage per ClusterQueue, and alert on wait time per queue, not just cluster utilisation.
- Version quota objects in Git and review them like code; a misplaced borrowingLimit changes who gets evicted.
- Set
stopPolicyon a ClusterQueue or LocalQueue to drain it for maintenance instead of deleting it. - Write down for each team: nominal quota, lending floor, preemption exposure and the checkpoint interval their jobs need.
Kueue sits beside other tools rather than replacing them. Kubeflow training jobs integrate with it directly. Compared with Slurm, Kueue provides the queue and quota layer that Slurm has natively, while leaving placement to Kubernetes. The broader picture of inventory, gang scheduling and failure handling is in GPU orchestration.
Failure modes
- Uncovered resource. The job requests ephemeral storage or a second extended resource not in any resource group, and stays pending forever with a quota message.
- Missing label. A job without the queue-name label bypasses Kueue entirely unless you configure Kueue to manage unlabelled jobs, which quietly consumes GPUs outside quota.
- Admitted but unschedulable. Quota fits but nodes are fragmented or cordoned; without waitForPodsReady the job holds quota indefinitely.
- Eviction storms. Large borrowers repeatedly evicted by small reclaims lose hours of work. Cap borrowing or checkpoint more often.
- Head-of-line starvation. BestEffortFIFO lets small jobs bypass a big one forever.
- Flavor and node mismatch. The flavor's node labels match no nodes, so every admitted job is unschedulable.
- Stale API version. Manifests still using v1beta1 fail on v0.20 or later; run the upstream migration first.
Trade-offs
Kueue's quota model gives teams guarantees and lets idle capacity flow, at the cost of preemption risk for anyone who borrows. Gang admission avoids partial allocations but can leave GPUs idle while a large job waits for a full fit. Keeping placement in the kube-scheduler preserves the Kubernetes ecosystem but means admission and placement can disagree, which you pay for with timeouts or topology constraints. Compared with an all-in-one commercial scheduler, Kueue is open and composable but needs more assembly: priority policy, aging, dashboards and quota governance are yours to build.
What to do next
- Inventory GPU node labels and create one ResourceFlavor per GPU type or pricing class.
- Create one ClusterQueue per team in a cohort, covering GPU, CPU and memory, with a lendingLimit that protects each team's floor.
- Create LocalQueues and require the queue-name label on training jobs.
- Turn on waitForPodsReady with a timeout above your slowest real start.
- Decide reclaim and priority policy together with each team's checkpoint interval.
- Rehearse the worked example in a test cluster and confirm who gets evicted.
- Dashboard pending time per queue and review quotas monthly.