Run:ai is a GPU orchestration layer for Kubernetes. It adds the pieces the stock Kubernetes scheduler lacks for AI work: queues with guaranteed quotas per team, lending of idle GPUs to teams over their quota, preemption to take them back, gang scheduling for distributed jobs, and fractional GPUs for notebooks and small inference services. NVIDIA acquired the company, the product is now sold as NVIDIA Run:ai, and in 2025 NVIDIA released its scheduler as the open-source KAI Scheduler under the Apache 2.0 licence.

This page explains the model underneath: how quota, over-quota and fair share decide who runs, what preemptibility means for your training code, how fractions split a GPU and what they do not split, and how to size a cluster for three teams in a worked example. It sticks to behaviour documented by NVIDIA and to the open-source scheduler's resource model. Product details change between releases, so check the documentation for your version before relying on any default. For the hardware-level sharing options this builds on, see Multi-Instance GPU and GPU time-slicing.

Why Kubernetes needs help with GPUs

In plain Kubernetes, a GPU is an integer extended resource such as nvidia.com/gpu: 1, advertised by the device plugin. The default scheduler places one pod at a time, first come, first served. That causes three problems for AI clusters.

  • No sharing policy between teams. ResourceQuota is a hard cap per namespace. Set it high and one team can take the whole cluster; set it low and GPUs sit idle while another team waits.
  • No gangs. A 16-GPU data-parallel job is 16 pods that are useless unless all start together. Scheduled one by one, two such jobs can each grab half the GPUs and deadlock, both waiting forever.
  • Whole GPUs only. A notebook that needs 6 GB of memory and runs a kernel a minute occupies an 80 GB accelerator all day.

Run:ai addresses all three with a scheduler that sees the whole queue of pending work and the whole cluster at once, plus an organisational model that turns budget decisions into scheduler inputs. Compare it with the batch-HPC answer to the same problem in Slurm GPU scheduling.

Departments, projects, queues and node pools

The organisational model has two levels. A department groups projects, typically one per business unit. A project is the unit that submits workloads, usually one per team or initiative, mapped to a Kubernetes namespace. Each project gets, per node pool, a deserved quota of GPUs it is guaranteed, an over-quota weight or priority that decides its share of idle capacity, and optionally a hard limit.

In the open-source KAI Scheduler the same ideas appear as hierarchical Queue objects. Each queue has a quota, a limit, an over-quota weight and a priority, and child queues point at a parent:

apiVersion: scheduling.run.ai/v2
kind: Queue
metadata:
  name: research
spec:
  parentQueue: ml-department
  resources:
    gpu:
      quota: 16          # guaranteed GPUs
      limit: 32          # hard ceiling, including borrowed GPUs
      overQuotaWeight: 2 # share of idle GPUs relative to sibling queues
    # cpu: and memory: blocks take the same keys; set them too, or CPU and
    # memory requests from this queue may not be admitted.

Node pools partition the cluster by hardware, for example H100 training nodes, L40S inference nodes and a pool of older cards for notebooks. Quotas are set per pool, so a team can hold 16 H100s and 4 L40S separately, and a workload can list the pools it accepts in preference order.

Run:ai on Kubernetes: a control plane for policy, a cluster-side scheduler for placementUsers and CIUI, CLI v2, APIControl planedepartments, projects, quotas, policysubmitCluster: schedulerqueues, fair share, gangs, preemptionqueue configKubernetes APIpods, podgroupsbind podsNode pool A: 8x GPU nodeswhole GPUs, bin-packedNode pool B: shared nodesfractions, time-slicingNode pool C: MIG nodesno Run:ai fractions hereNVIDIA device stack: driver, container toolkit, device plugin or GPU Operator, DCGM metrics
Policy lives in the control plane; placement happens in the cluster scheduler. Node pools separate hardware types, and Run:ai fractions cannot share a node with MIG.

How the scheduler decides

On each cycle the scheduler works out every queue's fair share and places pending work against it. The documented order of reasoning is:

  1. Satisfy deserved quotas first. A project asking for no more than its quota is entitled to those GPUs, even if another project has borrowed them.
  2. Divide the remaining idle GPUs among projects that want more, in proportion to over-quota weight, respecting each project's limit. Fair share is recalculated continuously as demand changes.
  3. When a project below its quota has pending work and no free GPUs exist, reclaim: preempt workloads of projects running above their fair share until the deficit is covered.
  4. Within a project, a higher-priority workload can preempt a lower-priority preemptible one.

Priority and preemptibility are attached to each workload through a priority class. NVIDIA's documentation lists a dictionary from very-low (25) through low (40), medium-low (65), medium (80) and medium-high (90), all preemptible, up to high (125) and very-high (150), which are non-preemptible. Non-preemptible workloads must fit inside the project's deserved quota, cannot use over-quota GPUs and are not interrupted once running. Preemptible workloads can borrow idle GPUs beyond quota and can be stopped at any time. By default, training workloads are low priority and preemptible, while interactive build workloads are high priority and non-preemptible. Defaults have moved between releases, so read them for your version.

Distributed jobs are scheduled as gangs. Pods belonging to one job are grouped into a podgroup, and the scheduler binds either the minimum number the job needs or none of them, which removes the half-placed deadlock. KAI's podgrouper builds these groups automatically for common frameworks such as the Kubeflow Training Operator, Ray and Argo. Placement then follows the node pool's strategy: bin-packing fills nodes to keep whole 8-GPU nodes free for large jobs, while spreading reduces contention for smaller ones.

Fractional GPUs

Run:ai fractions let several pods share one physical GPU. A pod requests either a portion of memory or an amount in MiB through annotations, and the platform enforces that limit:

apiVersion: v1
kind: Pod
metadata:
  name: notebook-small
  labels:
    kai.scheduler/queue: research          # open-source KAI queue label
  annotations:
    gpu-fraction: "0.25"                   # 25% of one GPU's memory
    # gpu-memory: "8192"                   # alternative: MiB, not a fraction
    # gpu-fraction-num-devices: "2"        # Run:ai platform docs: slice on 2 GPUs
spec:
  schedulerName: kai-scheduler
  containers:
    - name: jupyter
      image: nvcr.io/nvidia/pytorch:24.08-py3

The documentation is explicit about what a fraction is. gpu-fraction: "0.5" means half of the GPU's memory, not half of its compute. Each pod gets its own address space and cannot use more memory than it asked for or touch another pod's memory. Compute is shared by time-slicing, either NVIDIA's default round-robin or Run:ai's own time-slicing modes, which can weight compute between pods. Run:ai fractions and MIG cannot be used on the same node, and splitting memory into fractions can fragment it. Dynamic fractions add a request and a higher limit, so a pod can grow into unused memory while only its request is guaranteed.

From the CLI v2 the same requests look like this:

runai project set research
# 8 whole GPUs for a training job (preemptible by default)
runai training submit llama-ft -i registry.local/llama-ft:1.4 -g 8
# half a GPU's memory for an interactive workspace
runai workspace submit eda-notebook -i jupyter/base-notebook --gpu-portion-request 0.5

Worked example: three teams, 32 GPUs

Take a cluster of 4 nodes with 8 H100 GPUs each, 32 GPUs in one pool, shared by three projects. Research gets a deserved quota of 16, product fine-tuning 8 and an applied team 8, all with equal over-quota weight. Quotas sum to the cluster size, which is the safe default; over-committing quotas means guarantees that cannot all be honoured at once.

TimeDemand (R / P / A)Allocation (R / P / A)What happened
09:0024 / 0 / 424 / 0 / 4Research borrows 8 idle GPUs for one preemptible 8-GPU job
11:0024 / 8 / 416 / 8 / 4Product claims its quota; the 8-GPU gang is reclaimed whole, leaving 4 GPUs idle until a smaller job fits
14:0024 / 8 / 1216 / 8 / 8Every project is at or above quota; quotas use all 32, so nobody borrows
16:008 / 8 / 128 / 8 / 12Research shrinks; Applied borrows 4 over quota

The arithmetic is simple because quotas sum to capacity: once every project asks for at least its quota, the whole cluster is spoken for and Research's extra 8 GPUs of demand waits. At 16:00 Research drops to 8, leaving 8 idle; only Applied wants more, so it borrows 4. Real allocations are lumpier than this table, because an 8-GPU gang cannot use 4 free GPUs scattered over two nodes. That is why placement strategy matters, and why you should read the scheduler's own reports rather than predict them.

The lesson for training code is in the 11:00 row. Product needed only 4 GPUs beyond the idle ones, but the borrowed work was a single 8-GPU gang, and a gang cannot shrink, so all 8 GPUs of Research's job were stopped. If that job checkpointed every two hours, up to two hours of 8-GPU time was lost. Over-quota capacity is only worth having if jobs survive interruption.

Writing training code that survives preemption

Make training code preemption-tolerant before you rely on borrowed GPUs:

import os, signal, torch

stop = False
def on_sigterm(signum, frame):          # Kubernetes sends SIGTERM before SIGKILL
    global stop
    stop = True
signal.signal(signal.SIGTERM, on_sigterm)

ckpt = os.environ.get("CKPT", "/checkpoints/latest.pt")
step = 0
if os.path.exists(ckpt):
    state = torch.load(ckpt, map_location="cpu")
    model.load_state_dict(state["model"]); opt.load_state_dict(state["opt"])
    step = state["step"]

while step < total_steps:
    train_step(model, opt, next(batches))
    step += 1
    if stop or step % ckpt_every == 0:
        torch.save({"model": model.state_dict(), "opt": opt.state_dict(),
                    "step": step}, ckpt + ".tmp")
        os.replace(ckpt + ".tmp", ckpt)  # atomic: no half-written checkpoint
        if stop:
            break

Size ckpt_every so a checkpoint finishes well inside the pod's termination grace period, write to shared storage rather than the node's disk, and make data loaders resumable from the step counter. For multi-node jobs, only rank 0 should write the checkpoint, with every rank honouring the stop flag at the same step.

Operating it: the numbers to watch

Run the cluster from four numbers per node pool and per project. Allocation is the share of GPUs bound to pods; utilisation is how busy those GPUs actually are, from DCGM metrics. A gap between them means reserved but idle accelerators, usually notebooks or data-starved jobs. Pending time, split by job size, shows whether large gangs are starving behind fragmentation. Preemption count and the work lost per preemption show whether borrowing is paying off or just burning restarts.

Review those numbers with team leads monthly and change quotas from evidence rather than requests. The general scheduling principles behind these metrics are covered in GPU scheduling.

Failure modes

  • Over-committed quotas. If deserved quotas add up to more than the pool, two teams can both be under quota with no GPUs to reclaim. Keep the sum at or below capacity per node pool.
  • Idle non-preemptible workspaces. Notebooks left open hold quota all weekend. Set idle timeouts and move interactive users to fractions.
  • Fractions treated as compute isolation. A pod with 0.25 of the memory can still saturate the SMs when its neighbour's latency-sensitive service needs them. Keep latency-critical inference on whole GPUs or MIG slices.
  • Fragmentation. Many 1-GPU jobs spread across nodes leave no node with 8 free GPUs, so large gangs wait while utilisation looks healthy. Use bin-packing for training pools and watch pending gang time.
  • Preemption without checkpoints. The scheduler is working as designed; the job loses hours anyway.
  • Two schedulers, one node. Pods placed by the default scheduler on Run:ai-managed GPUs use GPUs without being charged to any project's quota. Route all GPU pods through one scheduler.

Trade-offs

OptionTeam quotas and borrowingGangsGPU sharingCost of adoption
Default scheduler + device pluginHard caps onlyNoWhole GPUsNone
Kueue or VolcanoYesYesVia MIG or time-slicingOperate it yourself
KAI Scheduler (open source)Hierarchical queuesYesFractionsOperate it yourself
NVIDIA Run:ai platformDepartments, projects, poolsYesFractions, dynamic, time-slicingLicence; adds UI, policy, reporting
SlurmFair-share accountsNativeMIG, MPSSeparate from Kubernetes

Choose the full platform when you need the UI, policies, chargeback reporting and support; start with KAI or Kueue when you mostly need fair queues and gangs and can run the scheduler yourself. Whatever you choose, the economics come from utilisation; GPU cost optimisation covers measuring it.

What to do next

  • Export a month of per-team GPU hours and peak demand per node pool before setting any quota.
  • Set deserved quotas that sum to no more than each pool's capacity, and weights that reflect priorities.
  • Make every training job checkpoint atomically on SIGTERM before allowing it to run over quota.
  • Move notebooks to fractional GPUs with idle timeouts; keep latency-critical inference on whole GPUs or MIG.
  • Choose bin-packing for pools that host multi-node training, then track pending time for large gangs.
  • Pilot the open-source KAI Scheduler on a test pool if you have not yet decided whether you need the full platform.
  • Read your version's documentation for priority and preemptibility defaults, and record them in your runbook.
Key takeaway: Run:ai turns team budgets into scheduler inputs: deserved quotas are guaranteed, idle GPUs are lent by weight, and borrowed GPUs are reclaimed by preemption. Gangs keep distributed jobs from deadlocking, and fractions split GPU memory but not compute. Keep quotas within capacity, make training checkpoint on SIGTERM before it borrows, keep latency-critical serving off shared GPUs, and verify defaults against your version's documentation.