A shared Hadoop cluster has an old problem: every job has to run with whatever is installed on the worker nodes. One team needs Python 3.11 with a particular pyarrow, another needs a native library built against a newer glibc, a third needs CUDA user-space libraries. Installing all of these on every NodeManager host is slow, needs coordination between teams, and turns every upgrade into a negotiation.
YARN's Docker runtime addresses the packaging problem without replacing YARN. The ResourceManager still schedules, queues still decide who gets capacity, and the NodeManager still enforces memory and CPU. The only change is that the process inside a container starts inside a Docker image the application chose, instead of directly on the host. This article explains how that launch works, how to enable it safely, how users and identities map into the image, and what breaks in practice. Names and defaults below follow the Apache Hadoop 3.5.0 Docker containers documentation; check the page for your own release before copying values.
How a Docker container is launched
Docker support is a runtime inside the LinuxContainerExecutor. It is not a separate executor and it is not a separate scheduler. The launch goes like this:
- The ApplicationMaster asks for containers as usual. In the container launch context it sets environment variables, most importantly
YARN_CONTAINER_RUNTIME_TYPE=dockerandYARN_CONTAINER_RUNTIME_DOCKER_IMAGE. - The NodeManager localizes resources (jars, archives, config files) into the application's local directories, just as it does for any container. See YARN containers in depth for localization and the kill sequence.
- The Docker runtime turns the request into a docker command: the image, the user, the network, and bind mounts for the localized directories, log directories and any mounts the application asked for. It writes this to a command file.
- The
container-executorbinary, which is setuid root and must be owned by root:hadoop with 6050 permissions, reads the command file. It rejects anything thatcontainer-executor.cfgdoes not allow, and only then calls the Docker daemon. - Inside the container, the process is either YARN's generated launch script or the image's own ENTRYPOINT, depending on one environment variable covered below. Memory and CPU limits still come from YARN; the cgroups article explains how they interact with Docker's own limits.
The important design point is that there are two layers of configuration. yarn-site.xml controls what the NodeManager will request. container-executor.cfg, owned by root, controls what the root binary will do. The second one is the real security boundary, because a compromised NodeManager account can write any command file it likes, but it cannot change the root-owned policy file.
Enabling the runtime safely
Enabling the runtime needs both files on every NodeManager, Docker installed on every host, and the LinuxContainerExecutor. Start from the most restrictive settings and loosen only what a real workload needs:
<!-- yarn-site.xml -->
<property><name>yarn.nodemanager.container-executor.class</name>
<value>org.apache.hadoop.yarn.server.nodemanager.LinuxContainerExecutor</value></property>
<property><name>yarn.nodemanager.runtime.linux.allowed-runtimes</name>
<value>default,docker</value></property>
<property><name>yarn.nodemanager.runtime.linux.docker.allowed-container-networks</name>
<value>host,none,bridge</value></property>
<property><name>yarn.nodemanager.runtime.linux.docker.default-container-network</name>
<value>host</value></property>
<property><name>yarn.nodemanager.runtime.linux.docker.privileged-containers.allowed</name>
<value>false</value></property>
<property><name>yarn.nodemanager.runtime.linux.docker.capabilities</name>
<value>CHOWN,DAC_OVERRIDE,FSETID,FOWNER,SETGID,SETUID,KILL</value></property># container-executor.cfg (root-owned)
yarn.nodemanager.linux-container-executor.group=hadoop
[docker]
module.enabled=true
docker.binary=/usr/bin/docker
docker.trusted.registries=registry.corp.example
docker.allowed.networks=host,none,bridge
docker.allowed.ro-mounts=/etc/passwd,/etc/group,/etc/hadoop/conf,/opt/hadoop
docker.allowed.rw-mounts=/var/hadoop/yarn/local-dir,/var/hadoop/yarn/log-dir
docker.allowed.capabilities=CHOWN,DAC_OVERRIDE,FSETID,FOWNER,SETGID,SETUID,KILL
docker.privileged-containers.enabled=false
docker.no-new-privileges.enabled=trueNotes on these settings. The NodeManager's local and log directories have to be in the read-write mount list, because that is how localized files reach the container and how logs get out. The default capability list in yarn.nodemanager.runtime.linux.docker.capabilities is broad (it includes NET_RAW, MKNOD and SYS_CHROOT), and the root-side docker.allowed.capabilities list allows none by default. Narrow both lists together and keep the requested list a subset of the allowed one, as in the example above; otherwise expect launches to be refused. docker.trusted.registries is more than a source allowlist. The documentation says images from untrusted registries do not get device mounts or access to host Hadoop configuration, and cannot write to volumes unless an administrator allows it, so a job that needs your mounted Hadoop config has to use an image from a trusted registry.
User identity inside the image
Most first attempts fail on user identity. YARN launches the container with docker run --user <uid>:<gid>, using the application owner's numeric UID (the documentation describes non-secure mode as launching as nobody). The process inside the image therefore runs as a number that may have no entry in the image's /etc/passwd. Shells show I have no name!, Java reports odd values for user.name, and HDFS clients that look up the current user can fail.
There are three ways to fix it, each with trade-offs:
- Bind-mount the host's user database, read-only: add
/etc/passwd:/etc/passwd:ro,/etc/group:/etc/group:roto the mounts and todocker.allowed.ro-mounts. This is simple and matches the host exactly, but it replaces the image's own users, which can break images that rely on them. - Use SSSD or LDAP inside the image, so it resolves the same directory the hosts use. This is closest to how a secured cluster works, at the cost of a heavier image and a dependency on the directory service.
- Static users baked into the image, with UIDs that match the cluster. This works only if UIDs are managed centrally and never drift.
On a Kerberos-secured cluster, the container runs as the submitting user's UID, so the image must resolve real user accounts, not just nobody. The container also needs credentials to talk to HDFS: delegation tokens are localized like any other resource, so a job that works on bare metal usually works in Docker once the user resolves and the Hadoop client configuration is mounted. Background is in Hadoop Kerberos architecture.
One more variable decides what runs inside the container: YARN_CONTAINER_RUNTIME_DOCKER_RUN_OVERRIDE_DISABLE. Left unset, the container runs in YARN mode: YARN's generated launch script replaces the image's command, and log redirection and environment setup work as they do on bare metal. This is what the documented MapReduce and Spark examples use. Set to true, it runs in Docker mode: the image's ENTRYPOINT runs, and YARN's launch command is passed to it as CMD parameters. The documentation's ENTRYPOINT section and its Docker registry examples use true this way, although its environment-variable table words the setting the other way round. Test your image in both modes before relying on either.
Worked example: PySpark on a newer Python
Worked example. A team needs PySpark on Python 3.11 with pyarrow on a cluster whose hosts have an older Python. Build an image that contains Java and the Python stack, and mount Hadoop and Spark from the host so their versions always match the cluster:
# Dockerfile
FROM registry.corp.example/base/jre17:2026.09
RUN apt-get update && apt-get install -y --no-install-recommends python3.11 python3-pip \
&& python3.11 -m pip install --no-cache-dir pyarrow==17.* pandas==2.2.* \
&& rm -rf /var/lib/apt/lists/*
ENV PYSPARK_PYTHON=/usr/bin/python3.11IMG=registry.corp.example/team-a/pyspark-311:1.4
MOUNTS=/etc/passwd:/etc/passwd:ro,/etc/group:/etc/group:ro,/etc/hadoop/conf:/etc/hadoop/conf:ro,/opt/hadoop:/opt/hadoop:ro
spark-submit --master yarn --deploy-mode cluster \
--conf spark.yarn.appMasterEnv.YARN_CONTAINER_RUNTIME_TYPE=docker \
--conf spark.yarn.appMasterEnv.YARN_CONTAINER_RUNTIME_DOCKER_IMAGE=$IMG \
--conf spark.yarn.appMasterEnv.YARN_CONTAINER_RUNTIME_DOCKER_MOUNTS=$MOUNTS \
--conf spark.yarn.appMasterEnv.YARN_CONTAINER_RUNTIME_DOCKER_CLIENT_CONFIG=hdfs:///user/team-a/.docker/config.json \
--conf spark.executorEnv.YARN_CONTAINER_RUNTIME_TYPE=docker \
--conf spark.executorEnv.YARN_CONTAINER_RUNTIME_DOCKER_IMAGE=$IMG \
--conf spark.executorEnv.YARN_CONTAINER_RUNTIME_DOCKER_MOUNTS=$MOUNTS \
--conf spark.executorEnv.YARN_CONTAINER_RUNTIME_DOCKER_CLIENT_CONFIG=hdfs:///user/team-a/.docker/config.json \
job.pyNote that the ApplicationMaster (the Spark driver in cluster mode) and the executors are configured separately. If you forget the appMasterEnv half, the driver runs on the host Python and the executors run in the image, and pickled data passed between them fails with version errors that look unrelated. The client configuration on HDFS holds the registry credentials. MapReduce works the same way through mapreduce.map.env, mapreduce.reduce.env and yarn.app.mapreduce.am.env.
Networking: the default host network is the right choice for Spark and MapReduce, because executors and the driver connect to each other on ports chosen at runtime, and bridge networking would need explicit port mappings (YARN_CONTAINER_RUNTIME_DOCKER_PORTS_MAPPING). Use bridge or none for services that should be isolated.
Failure modes
Most problems with the Docker runtime fall into these patterns.
- The first launch on every node is slow. The image is pulled on each NodeManager the first time it is used, so a large image delays every first task. Keep images small, use a shared base layer, and pre-pull hot images onto nodes.
yarn.nodemanager.runtime.linux.docker.image-update(default false) controls whether the NodeManager re-pulls an image tag it already has; mutable tags plus no re-pull means different nodes can run different code under the same tag, so use immutable tags. - The container exits immediately. Usually the launch mode is wrong for the image (see the override variable above), the daemon uses the systemd cgroup driver (the documentation states only cgroupfs is supported, and the launch fails with a cgroup-parent error), JAVA_HOME inside the image points to the wrong place, or a mount was silently dropped because it was not in the allowed list. Set
YARN_CONTAINER_RUNTIME_DOCKER_DELAYED_REMOVAL=trueon the job (it needsyarn.nodemanager.runtime.linux.docker.delayed-removal.allowedon the node) so the stopped container stays around fordocker inspectanddocker logs. - The launch is rejected. container-executor refused something, such as an untrusted registry, a mount outside the allowed list, a network that is not allowed or a privileged request. The NodeManager log has the reason. This is the policy working as intended, so fix the request rather than widening the policy.
- The container is killed for memory. YARN's limit includes everything in the container. A Python process with a large heap outside the JVM still counts, so raise
spark.executor.memoryOverheadrather than the heap. - Processes outlive the stop. On stop, the container gets
yarn.nodemanager.runtime.linux.docker.stop.grace-period(default 10 seconds) before it is killed. Images that ignore SIGTERM always use the full grace period, so preemption and decommission take longer. - The Docker daemon itself fails. If dockerd hangs, every Docker container on the node fails to launch. NodeManager health checks (NodeManager in depth) should include a cheap docker command so a node with a hung daemon is marked unhealthy and stops getting containers.
Debugging on a node
When a job misbehaves, start from the YARN side and work inward. The application and container IDs from the ResourceManager UI or yarn logs tell you which node to look at. On that node, the NodeManager log shows the runtime decision, the docker command the runtime asked for, and any rejection from container-executor. In current releases the runtime names each Docker container after its YARN container ID, which makes the two views easy to join; confirm with docker ps on your own nodes:
# Which YARN containers are running in Docker on this node
docker ps --format '{{.Names}} {{.Image}} {{.Status}}' | grep '^container_'
# Why did the runtime or container-executor refuse a launch?
grep -n 'container_1727900000000_0042_01_000003' /var/log/hadoop-yarn/*nodemanager*.log | tail -40
# Inspect a stopped container kept by delayed removal
docker inspect container_1727900000000_0042_01_000003 --format '{{json .Mounts}}'
docker logs container_1727900000000_0042_01_000003 | tail -50
# Aggregated stdout and stderr, same as for any YARN container
yarn logs -applicationId application_1727900000000_0042 -containerId container_1727900000000_0042_01_000003Check the mounts list in the inspect output first. A mount that the job requested but the policy did not allow is the most common reason a container works on a developer laptop and fails on the cluster. Then compare the user the container ran as with what the image can resolve; docker exec with id shows both the number and whether it has a name.
Trade-offs
Docker against packed environments. Spark can ship a packed conda or virtualenv archive with --archives and needs no host changes at all. That is enough for Python-only dependencies. Docker is worth its operational cost when you need system libraries, a different OS userland, or GPU user-space stacks.
Mounting the host against baking Hadoop into the image. Mounting Hadoop and its config keeps the client in step with the cluster and keeps images small, but it ties images to host paths. Baking Hadoop in makes images self-contained but means they must be rebuilt for every cluster upgrade.
Security against convenience. Every allowed mount, capability or privileged ACL expands what a job owner can do as root-adjacent code on your hosts. Keep the root-side policy narrow and give exceptions to specific registries, not to everyone.
YARN against Kubernetes. If most workloads already need images, ask whether the cluster should move to Kubernetes. YARN's Docker runtime is the right answer when HDFS locality, YARN queues and existing jobs keep you on YARN.
What to do next
- Confirm every NodeManager uses the LinuxContainerExecutor and that container-executor is root:hadoop with mode 6050.
- Write a minimal [docker] section: one trusted registry, explicit read-only and read-write mounts, a narrowed capability list, no privileged containers, no-new-privileges on.
- Choose and document one identity strategy (bind-mounted passwd and group, SSSD, or static users) and test it with a job running as a real user.
- Publish a base image with Java and your cluster's conventions, using immutable tags.
- Run a sample Spark job with both appMasterEnv and executorEnv set, and record which override setting your image needs.
- Add a docker health check to the NodeManager health script.
- Enable delayed removal on one node pool for debugging, and turn it off when you are done.