For most of its life YARN scheduled two things: memory and virtual cores. That was enough when every container ran a JVM doing map, reduce or Spark work. It stops being enough the moment a cluster gains a rack of GPU servers. A training job that needs two GPUs per container has no way to say so in a two-dimensional model, and a scheduler that cannot see GPUs will happily put four such containers on a node that has two, or put a CPU-only job on the GPU node and leave its accelerators idle.
The YARN resource model fixes that by making the resource vector extensible. You declare extra resource types, NodeManagers advertise how much of each they have, applications ask for them alongside memory and vcores, and the scheduler packs containers against every dimension at once. GPUs and FPGAs are built-in types with NodeManager plugins that discover the devices and isolate them per container. This article explains the model from first principles, shows the configuration that makes GPUs and FPGAs schedulable and isolated, works through a mixed CPU and GPU cluster, and lists the failure modes operators meet. Property names come from the Apache Hadoop documentation pages ResourceModel, UsingGpus and UsingFPGA as read on 2026-10-03 (labelled 3.3.5); check them against the release you run.
From request to isolated device
The resource model: types, units and limits
A YARN resource is a named, countable quantity. Memory (memory-mb) and vcores are always present. Everything else is declared in yarn.resource-types, a comma-separated list set in resource-types.xml or yarn-site.xml. Names must start with a letter and contain only letters, digits, dot, underscore or hyphen, may carry a namespace prefix such as com.acme/licence, and may not reuse memory, memory-mb or vcores. The two built-in accelerator types are yarn.io/gpu and yarn.io/fpga.
Each type can carry three more properties. yarn.resource-types.<name>.units sets the default unit, from a fixed list that runs from p and n through k, M, G, T, P and the binary Ki, Mi, Gi, Ti, Pi. yarn.resource-types.<name>.minimum-allocation and .maximum-allocation bound what one container may ask for. Requests are normalised: the documentation warns that a request may be adjusted to fit the configured minimum and maximum or rounded to allocation increments, so the container you get is not always the container you asked for. For a device count such as GPUs, set the maximum to the number of devices in your largest node; anything above it can never be placed.
Each NodeManager advertises its capacity per type. For arbitrary types that is yarn.nodemanager.resource-type.<name> in node-resources.xml or yarn-site.xml, with a value that may include a unit, for example 5G; YARN converts units when the node and the ResourceManager disagree. For GPUs and FPGAs you normally do not hard-code a number, because the device plugin discovers it. That difference matters: a hand-declared custom resource is pure bookkeeping, a counter the scheduler respects and nothing else enforces, while the GPU and FPGA plugins also assign specific devices to containers and fence the rest off.
Why the DominantResourceCalculator is required
Declaring a type is half the job. The Capacity Scheduler compares containers and queue usage through a resource calculator, and its default calculator only looks at memory. With it in place a GPU request is not treated as a scheduling constraint. Both the GPU and FPGA pages require yarn.scheduler.capacity.resource-calculator to be set to org.apache.hadoop.yarn.util.resource.DominantResourceCalculator. Forgetting it is a common reason a freshly configured GPU cluster schedules containers that then fail to find their devices.
The DominantResourceCalculator implements dominant resource fairness. For every container or queue it computes the share of the cluster used in each dimension and takes the largest as the dominant share. Fairness and capacity checks then compare dominant shares. Take a cluster of 20 nodes, each with 256 GiB and 64 vcores, eight of which also have four GPUs: 5,120 GiB, 1,280 vcores and 32 GPUs in total. A training container asking for 120 GiB, 16 vcores and two GPUs uses 2.3 percent of memory, 1.25 percent of vcores and 6.25 percent of GPUs, so its dominant share is 6.25 percent and GPU is its dominant resource. An ETL container with 16 GiB and 4 vcores has a dominant share of about 0.31 percent. A queue capped at 25 percent can therefore hold four such training containers, eight GPUs, however little memory they use. Under a memory-only calculator that same queue would look 9 percent full and keep accepting GPU containers until the hardware ran out.
GPUs on the NodeManager: discovery and isolation
GPU support is a NodeManager resource plugin. Enable it with yarn.nodemanager.resource-plugins set to yarn.io/gpu. On start-up the plugin runs nvidia-smi to find the devices; yarn.nodemanager.resource-plugins.gpu.path-to-discovery-executables points at the binary when it is not on the default path, for example /usr/local/bin/nvidia-smi. yarn.nodemanager.resource-plugins.gpu.allowed-gpu-devices defaults to auto, meaning every discovered GPU; to reserve devices for something else, list the ones YARN may use as index:minor_number pairs. The documentation states that only NVIDIA GPUs are supported and that NodeManagers must have the NVIDIA driver pre-installed.
Discovery tells the ResourceManager how many GPUs a node has. Isolation makes sure a container uses only the ones it was given. Without Docker, the plugin uses the cgroups devices controller to deny each container access to the device files of GPUs it was not assigned. That needs the LinuxContainerExecutor with cgroups configured, yarn.nodemanager.linux-container-executor.cgroups.mount (default true), and a [gpu] section in container-executor.cfg with module.enabled=true. The documented examples use the cgroup v1 layout; if your hosts run a unified cgroup v2 hierarchy, verify device isolation on your release rather than assuming it. The cgroup mechanics themselves are covered in YARN and Linux cgroups.
With the Docker runtime, yarn.nodemanager.resource-plugins.gpu.docker-plugin chooses how GPUs are passed in: nvidia-docker-v1 (the default, which calls a local endpoint at http://localhost:3476/v1.0/docker/cli) or nvidia-docker-v2. The [docker] section of container-executor.cfg must allow the device files and runtime, for example docker.allowed.devices=/dev/nvidiactl,/dev/nvidia-uvm,/dev/nvidia-uvm-tools,/dev/nvidia0,/dev/nvidia1 and docker.allowed.runtimes=nvidia. Image and user-identity concerns are the same as for any Docker container on YARN; see Docker containers on YARN.
Two consequences shape how you use it. Allocation is per device: yarn.io/gpu is an integer count of whole GPUs, so there is no built-in way to give two containers half a GPU each. And YARN isolates devices, not GPU memory or compute: a container that owns GPU 0 owns all of it.
FPGAs: discovery and programming the board
FPGA support has the same shape and narrower scope. Declare yarn.io/fpga, set yarn.nodemanager.resource-plugins to include it, and leave yarn.nodemanager.resource-plugins.fpga.allowed-fpga-devices at auto or list devices. The vendor plugin class defaults to IntelFpgaOpenclPlugin, which discovers boards with Intel's aocl tool, located through yarn.nodemanager.resource-plugins.fpga.path-to-discovery-executables. The [fpga] section of container-executor.cfg enables the module and names the device major number (documented default 246) and optionally the allowed minor numbers.
The interesting difference is programming. A GPU runs whatever kernel the process launches; an FPGA must be loaded with a bitstream, called an IP, before it is useful. An application names the IP it needs in the REQUESTED_FPGA_IP_ID environment variable and ships the matching .aocx file as a localised resource; the plugin finds it in the container's local directory and programs the assigned board before the process starts. The documented limits are real constraints: only the Intel OpenCL SDK for FPGA is supported, Docker is not, only one major device number may be configured, and one IP is flashed across all of a container's devices. Reprogramming takes time, so mixing jobs that need different IPs on the same boards costs start-up latency every time the bitstream changes.
Configuring the cluster and requesting accelerators
The configuration below is the minimum for a GPU cluster. The resource type and calculator go to the ResourceManager and every NodeManager; the plugin and the container-executor section go to GPU nodes.
<!-- resource-types.xml (ResourceManager and NodeManagers) -->
<property>
<name>yarn.resource-types</name>
<value>yarn.io/gpu</value>
</property>
<property>
<name>yarn.resource-types.yarn.io/gpu.maximum-allocation</name>
<value>4</value>
</property>
<!-- yarn-site.xml -->
<property>
<name>yarn.scheduler.capacity.resource-calculator</name>
<value>org.apache.hadoop.yarn.util.resource.DominantResourceCalculator</value>
</property>
<property>
<name>yarn.nodemanager.resource-plugins</name>
<value>yarn.io/gpu</value>
</property>
<property>
<name>yarn.nodemanager.resource-plugins.gpu.allowed-gpu-devices</name>
<value>auto</value>
</property>
# container-executor.cfg
[gpu]
module.enabled=trueApplications request accelerators the same way they request memory. Distributed shell takes a -container_resources list. MapReduce reads mapreduce.map.resource.<name>, mapreduce.reduce.resource.<name> and yarn.app.mapreduce.am.resource.<name>. Spark on YARN translates spark.executor.resource.gpu.amount into a yarn.io/gpu request by itself; the mapping can be changed with spark.yarn.resourceGpuDeviceName if you declared GPUs under a custom name. Spark's documentation adds a requirement that surprises people: YARN does not tell Spark which device addresses a container received, so each executor runs a discovery script at start-up that prints the GPUs it can see as JSON. With isolation configured it should report only the assigned devices; check that from inside a container. Without isolation, Spark makes you responsible for a script that keeps executors apart, or two executors on one node can collide on the same GPU.
# Distributed shell: two containers, two GPUs each
yarn jar hadoop-yarn-applications-distributedshell.jar \
-jar hadoop-yarn-applications-distributedshell.jar \
-shell_command nvidia-smi \
-container_resources memory-mb=3072,vcores=1,yarn.io/gpu=2 \
-num_containers 2
# MapReduce: one GPU per map task
-Dmapreduce.map.resource.yarn.io/gpu=1
# Spark on YARN: gpu maps to yarn.io/gpu automatically
--conf spark.executor.resource.gpu.amount=2 \
--conf spark.executor.resource.gpu.discoveryScript=/opt/spark/getGpusResources.sh \
--conf spark.task.resource.gpu.amount=1
Worked example: stranded GPUs on a shared cluster
Use the cluster from the calculator section: twelve CPU nodes and eight GPU nodes, each with 256 GiB, 64 vcores and, on the GPU nodes, four GPUs. Two teams share it. Analytics runs Spark ETL in 16 GiB, 4-vcore executors. Research trains models in containers of 120 GiB, 16 vcores and two GPUs, so a GPU node fits exactly two training containers with 16 GiB and 32 vcores spare.
On day one the operators enable the GPU plugin and the DominantResourceCalculator and do nothing else. The scheduler is correct and the result is still bad. GPU nodes are ordinary nodes to the ETL queue, so ETL executors land on them whenever they have free memory. After a busy morning a typical GPU node carries ten ETL executors, 160 GiB, leaving 96 GiB free. A training container needs 120 GiB, so neither of the node's training slots can be placed, and four GPUs sit allocatable but unusable. This is the stranded-accelerator problem: the scarce resource is blocked by a plentiful one.
The fix is to keep CPU work off GPU nodes. Put the eight GPU nodes in an exclusive node partition, give only the research queue access to it, and leave analytics on the default partition. Now the GPU partition holds 32 GPUs, 2,048 GiB and 512 vcores, and its dominant resource for every training container is GPU. Sixteen training containers fill it exactly. Research can also borrow CPU capacity in the default partition for preprocessing. Partitions, exclusivity and queue access are covered in YARN node labels, and the queue arithmetic in the Capacity Scheduler.
Finally, yarn.scheduler.maximum-allocation-mb must be at least 122,880, or the 120 GiB request is rejected before GPUs are even considered.
Failure modes
| Symptom | Likely cause | What to check |
|---|---|---|
| Containers start but see no GPU, or see all of them | Isolation not active: no [gpu] module, cgroups not mounted, or Docker devices not allowed | container-executor.cfg sections and NodeManager log at container launch |
| Node reports zero GPUs | nvidia-smi not found or driver not loaded | path-to-discovery-executables, run nvidia-smi as the NodeManager user |
| GPU requests ignored, nodes oversubscribed | Memory-only resource calculator | yarn.scheduler.capacity.resource-calculator is the DominantResourceCalculator |
| Request rejected as invalid | Above a maximum-allocation for memory, vcores or the custom type | Scheduler maximums versus the request shape |
| Training containers pending with GPUs free | Memory or vcores on GPU nodes taken by CPU work | Per-node usage; move GPU nodes to an exclusive partition |
| Two Spark executors share one GPU | No isolation, discovery script lists all devices | Enable isolation; check the discovery output in executor logs |
| FPGA job slow to start | Board reprogrammed for a different IP | REQUESTED_FPGA_IP_ID across jobs sharing boards |
| Unknown resource type at start-up | resource-types.xml differs between RM and NodeManagers | Distribute one file to every node |
Operating accelerator nodes
Treat resource-types.xml as cluster-wide configuration and roll it out to the ResourceManager and every NodeManager together; a node that advertises a type the ResourceManager does not know, or the reverse, is a start-up failure. Upgrading the NVIDIA driver is a node operation: drain the node with graceful decommission so running training is not killed, upgrade, run nvidia-smi as the NodeManager user, then recommission. The NodeManager's role in advertising capacity and draining is covered in the NodeManager article.
Per partition, watch GPUs allocated against GPUs still placeable given free memory and vcores on the same node; the gap is stranding. YARN does not measure GPU utilisation, so pair its metrics with nvidia-smi or DCGM exporters and chase containers that hold GPUs they do not use.
Trade-offs
Whole-device allocation is simple and safe but wastes capacity on small inference or notebook workloads that need a fraction of a GPU. Exclusive partitions remove stranding but make GPU-node CPU capacity unavailable to everyone else; non-exclusive partitions share idle capacity and risk stranding when demand returns. Docker images allow a CUDA version per job at the cost of image management. The FPGA path is narrow enough that you should prototype one job end to end before planning capacity around it.
What to do next
- Write down the resource shapes your jobs need, including memory and vcores per GPU, and check they pack into your nodes.
- Set the DominantResourceCalculator, declare
yarn.io/gpuwith a maximum allocation, and roll resource-types.xml to every node. - Enable the GPU plugin and the [gpu] container-executor module on GPU nodes; confirm nvidia-smi runs as the NodeManager user.
- Prove isolation: run distributed shell with two containers of two GPUs each and check each sees exactly two devices.
- Move GPU nodes into an exclusive partition and grant it to the queues that need accelerators.
- Raise scheduler memory maximums to fit your largest training container.
- For Spark, deploy a protected discovery script and verify executor logs list only assigned GPUs.
- Dashboard allocated versus usable GPUs per partition, and GPU utilisation per container.