Out of the box, Hadoop trusts whatever username the client claims. With the default hadoop.security.authentication=simple, a client's identity is its operating-system user, and on most distributions setting the HADOOP_USER_NAME environment variable is enough to become hdfs, the superuser. File permissions, Ranger policies and queue ACLs are all decisions based on that identity, so on an unauthenticated cluster they slow down honest users and stop nobody else.

Securing Hadoop takes five layers applied in order: authenticate with Kerberos, authorise with HDFS permissions and Ranger, encrypt on the wire, encrypt at rest with the KMS, and audit everything. This page is the umbrella checklist. It gives the configuration that matters for each layer, a rollout sequence that will not take production down, the failures you will actually hit, and links to the deeper pages for each component.

The five layers and their order

Defence in depth for a Hadoop clusterUsers and jobskinit / keytabPerimeterfirewall, edge nodes, KnoxKerberos KDCwho are youTGTAuthorizationRanger policies + HDFS perms/ACLs + YARN queue ACLsservice ticketEncryption in transitRPC privacy, SASL data transfer, HTTPSEncryption at restKMS + HDFS encryption zonesAuditingHDFS audit log + Ranger auditallow/deny eventsKMSkey ACLs, hdfs blacklistedEDEK decryptSIEM / alertingSolr search, HDFS retentionEach layer assumes the one above it can fail: authorization without authentication is meaningless,encryption without key ACLs only moves the secret, and nothing is provable without audit.
The five layers. Authentication comes first because every later decision depends on knowing who the caller is.

Two rules shape the programme. First, order matters: authorization rules written before Kerberos is enforced give a false sense of safety, and encryption keys protected only by unauthenticated ACLs are not protected at all. Second, every layer must fail closed. A Ranger plugin that cannot reach Ranger Admin keeps enforcing its cached policies, and a client that cannot get a Kerberos ticket gets an error rather than falling back to simple authentication. Check each of these behaviours in staging, not just in documentation.

Authentication with Kerberos

Kerberos gives every daemon and every user a cryptographic identity. In core-site.xml, the two switches are authentication and service-level authorization. Each daemon then gets its own principal, using the _HOST placeholder so one config file works on every node:

<!-- core-site.xml -->
<property><name>hadoop.security.authentication</name><value>kerberos</value></property>
<property><name>hadoop.security.authorization</name><value>true</value></property>
<property><name>hadoop.security.auth_to_local</name><value>
  RULE:[2:$1@$0](nn@CORP.EXAMPLE.COM)s/.*/hdfs/
  RULE:[2:$1@$0](dn@CORP.EXAMPLE.COM)s/.*/hdfs/
  RULE:[2:$1@$0](rm@CORP.EXAMPLE.COM)s/.*/yarn/
  RULE:[2:$1@$0](nm@CORP.EXAMPLE.COM)s/.*/yarn/
  RULE:[2:$1@$0](jhs@CORP.EXAMPLE.COM)s/.*/mapred/
  DEFAULT
</value></property>

<!-- hdfs-site.xml -->
<property><name>dfs.block.access.token.enable</name><value>true</value></property>
<property><name>dfs.namenode.kerberos.principal</name><value>nn/_HOST@CORP.EXAMPLE.COM</value></property>
<property><name>dfs.namenode.keytab.file</name><value>/etc/security/keytab/nn.service.keytab</value></property>
<property><name>dfs.web.authentication.kerberos.principal</name><value>HTTP/_HOST@CORP.EXAMPLE.COM</value></property>

The auth_to_local rules map principals to short Unix names. This is where subtle privilege bugs hide. A loose regex can map a user principal onto hdfs. Test every rule with hadoop kerbname nn/host1.corp.example.com@CORP.EXAMPLE.COM before deploying. Run HDFS daemons as hdfs, YARN daemons as yarn and the JobHistory server as mapred, all in group hadoop, so a compromise of one service account does not hand over the others.

Keytabs are passwords on disk. Make them mode 0400, owned by the service user, and never shared across services. Jobs do not carry keytabs to worker nodes. Instead, the client obtains delegation tokens at submit time. By default an HDFS token must be renewed daily and expires after seven days (dfs.namenode.delegation.token.renew-interval and dfs.namenode.delegation.token.max-lifetime), so a streaming job that runs for weeks needs a keytab login and periodic re-login:

Configuration conf = new Configuration();
conf.set("hadoop.security.authentication", "kerberos");
UserGroupInformation.setConfiguration(conf);
UserGroupInformation.loginUserFromKeytab(
    "etl-svc@CORP.EXAMPLE.COM", "/etc/security/keytabs/etl-svc.keytab");
// in a long-running loop, before each unit of work:
UserGroupInformation.getLoginUser().checkTGTAndReloginFromKeytab();

DataNodes need one extra decision. Without SASL on the data transfer protocol, a secure DataNode must bind privileged ports (1004 and 1006 by default), starting as root through jsvc and then dropping privileges. The cleaner modern option is to set dfs.data.transfer.protection, use non-privileged ports, set dfs.http.policy=HTTPS_ONLY and leave HDFS_DATANODE_SECURE_USER unset. The deep dive on principals, realms and KDC layout is in Hadoop and Kerberos.

Authorization with HDFS ACLs, Ranger and YARN

Once identities can be trusted, decide what each identity can do. There are three tools, and they overlap, so pick one source of truth for each resource.

  • HDFS permissions and ACLs are POSIX-style mode bits plus named-user and named-group entries (with dfs.namenode.acls.enabled=true). They are enforced inside the NameNode and work even when Ranger is down. See HDFS permissions and ACLs.
  • Ranger centralises policies for HDFS, Hive, HBase, YARN, Kafka and more, with deny rules, tag-based policies, row filtering and column masking for SQL engines. Plugins run inside each service and cache policies locally. See Ranger for Hadoop and Ranger policies in depth.
  • YARN queue ACLs (yarn.acl.enable=true, yarn.admin.acl, and per-queue submit and administer ACLs) control who may run work where. Unrestricted submission is unrestricted compute on every node.

The usual design is: lock sensitive HDFS paths down with restrictive mode bits (for example /data/finance owned by hdfs:hadoop at 0700), and let Ranger grant access on top. With the HDFS plugin, a request that no Ranger policy allows falls back to HDFS permissions, so a forgotten 0755 directory silently overrides your Ranger design. Audit mode bits on every governed path, or disable the fallback in the plugin configuration if your Ranger version supports that and you are sure every path has a policy.

On the compute side, switch YARN to the LinuxContainerExecutor. Containers then run as the submitting user rather than as yarn, so a job cannot read another user's local files or credentials:

# yarn-site.xml
yarn.nodemanager.container-executor.class = org.apache.hadoop.yarn.server.nodemanager.LinuxContainerExecutor

# container-executor.cfg  (root-owned; the binary is setuid root)
yarn.nodemanager.linux-container-executor.group=hadoop
banned.users=hdfs,yarn,mapred,bin
min.user.id=1000
allowed.system.users=

Encryption in transit

Kerberos authenticates the connection, but on its own it does not encrypt the data. Three channels need attention:

ChannelSettingValues
Hadoop RPC (clients, NameNode, ResourceManager)hadoop.rpc.protectionauthentication (default), integrity, privacy
DataNode block transferdfs.data.transfer.protection or dfs.encrypt.data.transfer=trueprivacy; cipher suite AES/CTR/NoPadding
Web UIs and WebHDFSdfs.http.policyHTTPS_ONLY, with a keystore in ssl-server.xml
MapReduce shufflemapreduce.shuffle.ssl.enabledtrue, sharing the HTTPS keystores

Set dfs.encrypt.data.transfer.cipher.suites to AES/CTR/NoPadding. Without it, block encryption falls back to much slower legacy ciphers. The overhead depends heavily on CPU (AES-NI) and workload, so measure it on your own hardware with a TeraSort or DFSIO run before and after, rather than trusting published figures. Roll RPC protection out carefully: clients and servers negotiate it, so move servers to accept a list such as authentication,privacy first, then switch clients, then drop the weaker value.

Encryption at rest with the KMS

Disk encryption at the OS level protects against stolen drives, not against a user who can read HDFS. HDFS transparent encryption encrypts file contents end-to-end. The client encrypts and decrypts, and the DataNodes only ever store ciphertext. Each file has its own data encryption key (DEK). The NameNode stores that key only in encrypted form (the EDEK), wrapped by a zone key held in the Hadoop KMS:

hadoop key create sales_k -size 256          # zone key lives in the KMS
hdfs dfs -mkdir /data/sales                  # zone root must be empty
hdfs crypto -createZone -keyName sales_k -path /data/sales
hdfs crypto -listZones

The security of the whole scheme comes down to KMS ACLs. The NameNode needs to generate EDEKs, but the hdfs superuser should never be able to decrypt them. Otherwise an HDFS admin can read every encrypted file, and the encryption only adds latency:

<!-- kms-acls.xml -->
<property><name>hadoop.kms.blacklist.DECRYPT_EEK</name><value>hdfs</value></property>
<property><name>key.acl.sales_k.GENERATE_EEK</name><value>hdfs</value></property>
<property><name>key.acl.sales_k.READ</name><value>hdfs</value></property>
<property><name>key.acl.sales_k.DECRYPT_EEK</name><value>etl-svc sales-analysts</value></property>

Know the operational consequences. Files cannot be renamed across zone boundaries, so moving data into a zone is a copy. Trash needs a per-zone .Trash. The KMS sits on the read path, so it must be highly available behind a load balancer. Back up the keystore that sits behind the KMS separately from the cluster: lose the zone key and the data is gone. Details are in HDFS encryption zones.

Auditing and detection

Auditing is what turns your controls into evidence. There are two streams. The NameNode writes one line per file-system operation to its audit logger, and Ranger plugins send allow and deny events to Solr for search and to HDFS for long retention. A NameNode audit line looks like this:

2026-10-08 09:14:02,311 INFO FSNamesystem.audit: allowed=false ugi=analyst1@CORP.EXAMPLE.COM (auth:KERBEROS) ip=/10.20.4.17 cmd=open src=/data/finance/q3.parquet dst=null perm=null proto=rpc

Ship both streams off the cluster: an attacker with superuser rights can edit local logs. Then alert on patterns, not just individual lines. A small detector:

import re, collections

LINE = re.compile(r"allowed=(\w+) ugi=(\S+).*? cmd=(\w+) src=(\S+)")
SENSITIVE = ("/data/finance", "/data/sales")

def scan(lines, read_burst=5000):
    denied, reads, alerts = collections.Counter(), collections.defaultdict(set), []
    for ln in lines:
        m = LINE.search(ln)
        if not m:
            continue
        allowed, ugi, cmd, src = m.groups()
        if allowed == "false":
            denied[ugi] += 1
        if cmd == "open" and src.startswith(SENSITIVE):
            reads[ugi].add(src)
        if cmd in ("setPermission", "setOwner", "setAcl") and src.startswith(SENSITIVE):
            alerts.append(f"permission change on {src} by {ugi}")
        if "(auth:SIMPLE)" in ln:
            alerts.append(f"simple auth seen for {ugi}: Kerberos not enforced somewhere")
    alerts += [f"{u}: {n} denials" for u, n in denied.items() if n > 100]
    alerts += [f"{u}: read {len(s)} sensitive files" for u, s in reads.items() if len(s) > read_burst]
    return alerts

Run it per hour window. Tune the thresholds against a few weeks of baseline, because one busy ETL account can dwarf all other users. Pair the audit trail with lineage from Apache Atlas so an alert on a table can be traced to the jobs that wrote it.

Worked example: a phased rollout

Here is a realistic sequence for a 40-node production cluster that currently runs with simple authentication. Each phase is reversible and is rehearsed on a staging cluster first.

  1. Week 1, groundwork. Make forward and reverse DNS consistent, because _HOST resolution depends on it, and get NTP healthy everywhere. Stand up a replicated KDC or join Active Directory. Create principals and keytabs per service per host.
  2. Week 2, Kerberos. Enable it in one maintenance window: all daemons restart with the new config. Then fix the clients that break. These are usually cron jobs that never ran kinit and service accounts with no keytab.
  3. Week 3, authorization. Install Ranger plugins in audit-only mode by granting broad policies and recording who touches what. After two weeks of audit data, write least-privilege policies from it, then tighten HDFS mode bits on governed paths.
  4. Week 5, wire encryption. Move RPC to accept both levels, then privacy, then enable SASL data transfer with AES and HTTPS-only UIs. Benchmark at each step.
  5. Week 6, encryption at rest. Deploy KMS in HA with blacklisted hdfs. Create zones for regulated datasets and copy data in with distcp -update -skipcrccheck. The stored ciphertext never matches the source checksums, so verify by hashing the content as read back through a client.
  6. Ongoing, audit. Forward logs to the SIEM, alert as above, and review Ranger policies quarterly.

Failure modes

  • Clock skew. Kerberos rejects tickets once clocks drift past the KDC tolerance (commonly five minutes) with "Clock skew too great". Monitor NTP offset as a first-class metric.
  • Keytab rotation. Re-exporting a keytab bumps the key version number (kvno), so every running daemon holding the old key fails at its next authentication. Rotate one service at a time and restart it immediately.
  • Token expiry. Jobs that run longer than the delegation-token maximum lifetime die after a week with confusing token errors. Give long-running services keytabs and re-login logic.
  • KMS outage. Reads and writes inside zones fail while files outside zones keep working, which looks like a partial HDFS outage. Run at least two KMS instances and alert on its latency.
  • Ranger Admin outage. Plugins keep enforcing cached policies, so nothing breaks, but policy changes stop propagating, and a revoked user keeps access until Admin returns.
  • auth_to_local drift. A new realm or a trust with another domain adds principals your rules never anticipated. A DEFAULT fallback that maps them to an existing local user is a privilege-escalation path.

Trade-offs

ControlMain costWhen you might defer it
KerberosKDC operations, client friction, keytab handlingNever on a shared or networked cluster
RangerAnother HA service, policy sprawlSingle-team clusters can live on HDFS ACLs for a while
RPC and data privacyCPU and throughputOnly on isolated networks with documented risk acceptance
Encryption zonesKMS on the read path, no cross-zone renameData with no regulatory or contractual requirement
Audit shippingStorage and SIEM licensingNever: it is the evidence for every other row

What to do next

  1. Prove the problem: on a staging cluster, run HADOOP_USER_NAME=hdfs hdfs dfsadmin -report as an ordinary user. This command is superuser-only, so if it succeeds, anyone can impersonate hdfs and your cluster has no authentication.
  2. Fix DNS and NTP, stand up an HA KDC, and enable Kerberos with service-specific principals and tested auth_to_local rules.
  3. Switch YARN to LinuxContainerExecutor and enable queue ACLs.
  4. Run Ranger in audit-first mode, then write least-privilege policies and lock HDFS mode bits under them.
  5. Turn on RPC privacy, SASL data transfer with AES/CTR and HTTPS-only, benchmarking each step.
  6. Create encryption zones for regulated data with KMS ACLs that blacklist hdfs from DECRYPT_EEK.
  7. Ship NameNode and Ranger audit logs off-cluster and put the pattern alerts above into production.
Key takeaway: A Hadoop cluster without Kerberos trusts whatever name the client claims, so authenticate first. Then authorise with one clear source of truth per resource, encrypt RPC and block transfer with AES, protect data at rest with encryption zones whose KMS ACLs keep the hdfs superuser away from the keys, and ship audit logs off-cluster with alerts on patterns. Roll out in phases, and monitor clock skew, keytabs, tokens and the KMS as the things that will break.