A YARN cluster is rarely uniform. Some nodes have GPUs, some have fast local SSDs, some run an older operating system, and some services must never share a node with each other. YARN offers three separate tools to express where a container may run: node labels (partitions), node attributes, and placement constraints with allocation tags. They overlap just enough to confuse, and using the wrong one produces applications that sit in ACCEPTED forever or quietly land on the wrong hardware.
This article treats them as one system. It explains which question each tool answers, builds a worked cluster that uses all three, follows a request through the ResourceManager, and lists the failure modes and the commands that diagnose them. The partition mechanics on their own are covered in depth in YARN node labels architecture; here the focus is on combining them. YARN behaviour described here follows the Apache Hadoop documentation for node labels, node attributes and placement constraints; framework property names come from the Spark and MapReduce documentation, and API details should be checked against the javadoc of your Hadoop version.
Three tools, three questions
| Tool | Question it answers | Tied to queues and capacity? | Per node |
|---|---|---|---|
| Node label (partition) | Which sub-cluster may this container use, and whose capacity pays for it? | Yes: queues get per-label capacity | Exactly one |
| Node attribute | Which nodes have a property this container needs? | No | Many |
| Placement constraint | Where should this container be relative to other containers? | No | Evaluated at allocation time |
The rule of thumb follows from the middle column. Use a partition when the hardware is scarce or expensive enough that you want to give queues guaranteed shares of it, and keep other work off it. Use an attribute when nodes differ in a way that some applications care about but that should not carve up capacity, such as a library version or kernel. Use a constraint when the requirement is about co-location: spread replicas for fault tolerance, pack chatty tasks together, or cap how many of something run on one node.
Partitions: the rules that bite
- A node belongs to exactly one partition. Nodes you never map stay in the DEFAULT partition, whose name is the empty string. The cluster is therefore cut into disjoint sub-clusters.
- Partitions are exclusive by default: only requests asking for that label get its containers. A non-exclusive partition lends idle resources to requests for the DEFAULT partition, and takes them back, by preemption if necessary, when label requesters need them.
- A queue can use a partition only if it is listed in
accessible-node-labels. Capacity on each label is configured per queue, and for every parent the label capacities of its direct children must sum to 100. - Capacity on a label is separate from capacity on DEFAULT. A queue with 50% of the cluster and 0% of
gpucannot run a single container on GPU nodes. - The labels feature is configured through the Capacity Scheduler; the documented queue properties are Capacity Scheduler properties.
Worked example: a cluster with GPU and SSD nodes
The cluster has 60 NodeManagers: 40 general nodes, 8 GPU nodes and 12 nodes with large local SSDs. Two teams share it. ml trains models and needs GPUs; analytics runs Spark SQL that benefits from SSD shuffle space and occasionally uses GPUs for inference.
Design decisions. GPU nodes become an exclusive partition: they are expensive, and letting a long Spark job squat on them would block training. SSD nodes become a non-exclusive partition: when analytics is quiet, their CPU and memory should serve general work rather than sit idle. CUDA version and OS image become attributes, because they matter to only some jobs and should not split capacity. Parameter-server spreading becomes a placement constraint.
# yarn-site.xml on the ResourceManager
# yarn.node-labels.enabled = true
# yarn.node-labels.fs-store.root-dir = hdfs://nn:8020/yarn/node-labels
# yarn.node-labels.configuration-type = centralized
# Declare the partitions once. exclusive defaults to true.
yarn rmadmin -addToClusterNodeLabels "gpu(exclusive=true),ssd(exclusive=false)"
# Map nodes. A node holds exactly one partition; unmapped nodes stay in DEFAULT.
yarn rmadmin -replaceLabelsOnNode "gpu01=gpu gpu02=gpu ssd01=ssd ssd02=ssd" -failOnUnknownNodes
# Check the result
yarn cluster --list-node-labels
yarn node -list -showDetailsNow queue capacity on each label. Root must be able to access both labels with 100% of each; below root, the children's shares of every label sum to 100.
<!-- capacity-scheduler.xml: root has two children, ml and analytics -->
<property><name>yarn.scheduler.capacity.root.accessible-node-labels</name><value>gpu,ssd</value></property>
<property><name>yarn.scheduler.capacity.root.accessible-node-labels.gpu.capacity</name><value>100</value></property>
<property><name>yarn.scheduler.capacity.root.accessible-node-labels.ssd.capacity</name><value>100</value></property>
<property><name>yarn.scheduler.capacity.root.ml.accessible-node-labels</name><value>gpu,ssd</value></property>
<property><name>yarn.scheduler.capacity.root.ml.accessible-node-labels.gpu.capacity</name><value>80</value></property>
<property><name>yarn.scheduler.capacity.root.ml.accessible-node-labels.gpu.maximum-capacity</name><value>100</value></property>
<property><name>yarn.scheduler.capacity.root.ml.accessible-node-labels.ssd.capacity</name><value>30</value></property>
<property><name>yarn.scheduler.capacity.root.analytics.accessible-node-labels</name><value>gpu,ssd</value></property>
<property><name>yarn.scheduler.capacity.root.analytics.accessible-node-labels.gpu.capacity</name><value>20</value></property>
<property><name>yarn.scheduler.capacity.root.analytics.accessible-node-labels.gpu.maximum-capacity</name><value>40</value></property>
<property><name>yarn.scheduler.capacity.root.analytics.accessible-node-labels.ssd.capacity</name><value>70</value></property>
<property><name>yarn.scheduler.capacity.root.analytics.default-node-label-expression</name><value>ssd</value></property>Work the arithmetic. Suppose each GPU node offers 4 GPUs, 256 GB and 48 vcores, so the gpu partition has 32 GPUs. ml is guaranteed 80% of it, roughly 25 GPUs, and may grow to the whole partition when analytics is idle. analytics is guaranteed about 6 GPUs and capped at 40%, about 12. On the ssd partition analytics is guaranteed 70% and ml 30%. Because analytics has default-node-label-expression=ssd, its jobs land on SSD nodes unless they ask otherwise. Because ssd is non-exclusive, when neither team asks for ssd, a DEFAULT-partition job from either queue can borrow those nodes, and it is preempted when a labelled request needs them back. That lending is why preemption must be enabled and its tuning understood; see YARN preemption.
Node attributes: properties without capacity
Attributes, added in Hadoop 3.2, describe nodes without partitioning them. A node can carry many attributes; they need no cluster-wide declaration and are not attached to queues. Attributes set centrally through the ResourceManager carry the rm.yarn.io prefix, and attributes reported by the NodeManager itself in distributed mode carry nm.yarn.io. Only string values are supported, and the only operators are equals and not-equals, combinable with AND and OR.
# Centralized attributes (rm.yarn.io prefix). Values are strings.
yarn nodeattributes -add "gpu01:cuda=12,os=ubuntu22 gpu02:cuda=11,os=ubuntu22"
yarn nodeattributes -nodestoattributes
yarn nodeattributes -attributestonodesIn distributed mode, each NodeManager runs a provider, configured with yarn.nodemanager.node-attributes.provider, that can execute a script on an interval and report what it finds. This suits facts that change with the node image, such as installed driver versions, because the node reports the truth instead of an operator remembering to update a mapping. Partitions have a similar distributed mode: with yarn.node-labels.configuration-type=distributed and yarn.nodemanager.node-labels.provider=script, the script prints a line starting with NODE_PARTITION: and the NodeManager reports that label.
Placement constraints and allocation tags
Constraints are expressed against allocation tags, strings an application attaches to its containers, such as ps, worker or hbase-rs. A constraint has a scope, NODE or RACK, and one of three forms: IN (affinity: place where the target tag exists), NOTIN (anti-affinity) and CARDINALITY (between a minimum and maximum matching containers in the scope). Targets can also be node partitions or node attributes. Tags are scoped by namespace: self (the default, this application only), not-self, all, app-id/<id> and app-tag/<tag>, so one application can avoid nodes running another's tagged containers.
The ResourceManager processes constraints only when yarn.resourcemanager.placement-constraints.handler is set. The default, disabled, rejects requests that carry constraints. placement-processor places constraint requests before the scheduler sees them, supports affinity, anti-affinity and cardinality, and works with both the Capacity and Fair Schedulers. scheduler integrates with the Capacity Scheduler directly and respects queue and task priorities, but currently supports only anti-affinity.
The quickest way to experiment is the distributed shell's -placement_spec. The documented example zk(3),NOTIN,NODE,zk:hbase(5),IN,RACK,zk:spark(7),CARDINALITY,NODE,hbase,1,3 asks for three zk containers on different nodes, five hbase containers in racks that hold a zk container, and seven spark containers on nodes holding between one and three hbase containers. Attribute expressions such as python!=3:java=1.8 use the same equality operators.
A real ApplicationMaster builds constraints with the PlacementConstraints API and submits them as SchedulingRequest objects:
// In the ApplicationMaster (Hadoop 3.1+ API; check the javadoc of your version).
import static org.apache.hadoop.yarn.api.resource.PlacementConstraints.*;
import static org.apache.hadoop.yarn.api.resource.PlacementConstraints.PlacementTargets.*;
// Parameter servers: at most one per node, so one node failure loses one shard.
PlacementConstraint psSpread = targetNotIn(NODE, allocationTag("ps")).build();
// Workers: on GPU nodes running CUDA 12, and in the same rack as a parameter server.
PlacementConstraint workerPlace = and(
targetIn(NODE, nodeAttribute("cuda", "12")),
targetIn(RACK, allocationTag("ps"))).build();
SchedulingRequest ps = SchedulingRequest.newBuilder()
.allocationRequestId(1L)
.priority(Priority.newInstance(1))
.allocationTags(Collections.singleton("ps"))
.executionType(ExecutionTypeRequest.newInstance(ExecutionType.GUARANTEED))
.resourceSizing(ResourceSizing.newInstance(4, Resource.newInstance(8192, 4)))
.placementConstraintExpression(psSpread)
.build();
amrmClient.addSchedulingRequests(Arrays.asList(ps /*, workers */));
Following a request through the ResourceManager
- The client submits the application to a queue with an optional label expression on the submission context. If none is given, the queue's
default-node-label-expressionapplies. - The scheduler checks the queue can access that label. If not, submission fails or, worse, the application is admitted but its requests can never be satisfied.
- The ApplicationMaster container is placed using the AM label expression. An AM that needs a partition where the queue has no capacity stays in ACCEPTED.
- The AM sends ordinary resource requests with a label expression, or scheduling requests with tags and constraints.
- With the placement-processor handler, constraint requests are placed first, honouring partition and attribute targets, then committed through the scheduler, which still enforces queue capacity on the chosen partition. Requests whose constraints cannot be met are rejected back to the AM after retries rather than waiting forever.
- Allocated containers carry their tags, so later constraints from this or other applications can see them.
The AM's role in this loop is described in YARN ApplicationMaster.
Framework settings
Most users never write an AM; they set framework properties. Spark uses spark.yarn.am.nodeLabelExpression and spark.yarn.executor.nodeLabelExpression, which lets you keep the driver's AM on cheap general nodes while executors go to gpu. MapReduce uses mapreduce.job.node-label-expression, with mapreduce.job.am.node-label-expression and per-phase map and reduce variants. Placing the AM on the default partition is usually right: it is small, long-lived, and should not consume scarce partition capacity.
Failure modes and diagnostics
| Symptom | Likely cause | How to confirm |
|---|---|---|
| Application stuck in ACCEPTED | AM label has no capacity in the queue, or label typo | RM UI application diagnostics; yarn queue -status <queue> |
| Containers never allocated, AM healthy | Executor label not in queue's accessible labels, or label capacity 0 | Check per-label capacity in the scheduler page |
| GPU nodes busy with non-GPU work | Partition created non-exclusive, or not mapped | yarn node -list -showDetails |
| Borrowed containers killed repeatedly | Non-exclusive lending plus preemption when owners return | Preemption events in RM log; expected behaviour |
| Constraint requests rejected immediately | Handler left at disabled, or affinity requested with the scheduler handler | Check the handler property |
| Placement skewed after node replacement | New node not mapped, sits in DEFAULT | Use distributed providers or automate mapping in provisioning |
Two operational habits prevent most of these. Make label and attribute mapping part of node provisioning rather than a manual step, and alert on applications that stay in ACCEPTED longer than a few minutes, because label misconfiguration fails silently.
Trade-offs
Every partition fragments the cluster. Twelve SSD nodes in an exclusive partition are twelve nodes the rest of the cluster cannot use, and each extra label multiplies the queue configuration you must keep summing to 100. Prefer few partitions, make them non-exclusive unless the hardware is genuinely scarce, and push everything else into attributes. Constraints add scheduling cost and can be unsatisfiable; strict anti-affinity on a small cluster turns a node failure into an unschedulable job. Use cardinality ranges rather than absolute rules where the application tolerates it. For queue design across partitions, see the Capacity Scheduler.
What to do next
- List your hardware differences and classify each one: capacity-worthy (partition), descriptive (attribute) or relational (constraint).
- Create partitions with explicit
exclusive=flags and map nodes from provisioning, using-failOnUnknownNodes. - Write per-label capacities for every queue level and check that children sum to 100 for each label.
- Keep ApplicationMasters on the default partition with the AM label properties.
- Enable a constraints handler deliberately and test it with the distributed shell before relying on it.
- Alert on long ACCEPTED states and on DEFAULT-partition nodes that should be labelled.