Kerberos and Ranger answer two different questions. Kerberos answers "who is this" with cryptographic proof and no passwords on the wire. Ranger answers "may this person do this to that table, column or row" from central policies, and records the answer. In a secured Hive deployment every query depends on both, and on a chain of translations between them: a Kerberos principal becomes a short user name, the user name becomes a set of groups, and the groups are what Ranger policies usually name. Most production incidents in secured Hive are a break somewhere in that chain, and the error messages rarely say where.

This article follows one identity end to end, shows the configuration for each hop, gives a rollout order that does not take a running cluster down, and ends with a runbook organised by the error you actually see. HiveServer2 authentication modes and transport options are covered in Hive security, and the authorization models and Ranger policy semantics in Hive authorization; this page is about joining them.

Advertisement

The identity chain for one query

Alice runs a SELECT through Beeline. Ten things have to go right:

  1. Alice runs kinit and gets a ticket-granting ticket for alice@EXAMPLE.COM from the KDC.
  2. Her JDBC driver asks the KDC for a service ticket for hive/hs2-1.example.com@EXAMPLE.COM.
  3. The driver opens a SASL GSSAPI connection to HiveServer2 and presents the ticket.
  4. HiveServer2 decrypts it with the key in its keytab and now knows the caller is alice@EXAMPLE.COM.
  5. Hadoop's auth_to_local rules map that principal to the short name alice.
  6. Group mapping resolves alice to her groups, for example analysts and emea.
  7. The Ranger Hive plugin inside HiveServer2 evaluates its cached policies for alice and her groups against the tables and columns in the query, and applies any row filter or mask.
  8. If allowed, HiveServer2 compiles the query and runs it on Tez. With doAs off, tasks run as the hive service user.
  9. HiveServer2 talks to the metastore over SASL as its own hive principal.
  10. Tasks read HDFS using delegation tokens obtained by HiveServer2, so no task needs a keytab.
One query, one identity: from kinit to a Ranger decision to an HDFS read1. kinitTGT for alice@REALM2. KDCservice ticket for hive3. Beeline JDBCSASL GSSAPI to HS24. HiveServer2authenticates alice5. auth_to_localalice@REALM to alice6. Group lookupOS, SSSD or LDAP7. Ranger plugincached policiesRanger Adminpolicy sourcepoll8. Compile and runTez as user hiveallow9. MetastoreSASL, hive principal10. HDFSdelegation tokensRanger auditSolr or HDFSSteps 1-5 are Kerberos and Hadoop naming; 6-7 are Ranger and group resolution; 8-10 run as the hiveservice with doAs off. A mismatch at any hop shows up as a denial or an opaque GSS error.
Every arrow is a place where a name, a key or a group can disagree with what the next component expects.

Service principals and keytabs

Each HiveServer2 and metastore host needs a service principal of the form hive/fully.qualified.host@REALM and a keytab holding its key. The hostname part must match the name clients use to reach the host, after any canonicalisation the client's Kerberos library performs, which is why reverse DNS and the krb5.conf canonicalisation settings matter. Create and check them like this:

# On the KDC: one service principal per HS2 / HMS host, keys exported once
kadmin.local -q "addprinc -randkey hive/hs2-1.example.com@EXAMPLE.COM"
kadmin.local -q "ktadd -k /tmp/hive.service.keytab hive/hs2-1.example.com@EXAMPLE.COM"
#   NB: ktadd re-randomises the key and bumps the kvno; running it again
#   invalidates every keytab already deployed for that principal.

# On the host: check the keytab before Hive ever touches it
klist -kte /etc/security/keytabs/hive.service.keytab
kinit -kt /etc/security/keytabs/hive.service.keytab hive/hs2-1.example.com@EXAMPLE.COM && klist

Keep keytabs readable only by the service user, mode 400, and treat them like private keys: anyone who copies one can be that service. Enforce AES encryption types in krb5.conf and on the principals, and remove legacy RC4 and DES types, which are weak and which newer KDCs and JDKs reject anyway.

Advertisement

Configuring HiveServer2 and the metastore

The server side needs Kerberos authentication, the keytab, wire protection and the Ranger authorizer. The _HOST token is replaced by each host's own name at startup, so one configuration file serves every node:

<!-- hive-site.xml (HiveServer2 and Metastore) -->
<property><name>hive.server2.authentication</name><value>KERBEROS</value></property>
<property><name>hive.server2.authentication.kerberos.principal</name><value>hive/_HOST@EXAMPLE.COM</value></property>
<property><name>hive.server2.authentication.kerberos.keytab</name><value>/etc/security/keytabs/hive.service.keytab</value></property>
<property><name>hive.server2.thrift.sasl.qop</name><value>auth-conf</value></property>
<property><name>hive.server2.enable.doAs</name><value>false</value></property>

<property><name>hive.metastore.sasl.enabled</name><value>true</value></property>
<property><name>hive.metastore.kerberos.principal</name><value>hive/_HOST@EXAMPLE.COM</value></property>
<property><name>hive.metastore.kerberos.keytab.file</name><value>/etc/security/keytabs/hive.service.keytab</value></property>

<property><name>hive.security.authorization.enabled</name><value>true</value></property>
<property><name>hive.security.authorization.manager</name>
  <value>org.apache.ranger.authorization.hive.authorizer.RangerHiveAuthorizerFactory</value></property>

Setting the SASL QOP to auth-conf encrypts the JDBC session itself rather than only authenticating it. If you run HiveServer2 in HTTP transport mode instead, authentication uses SPNEGO with an HTTP/_HOST principal and wire protection comes from TLS. The metastore must have SASL enabled as well; an unauthenticated metastore lets anyone with network access rewrite table locations and bypass everything else on this page. The Ranger plugin also reads its own file, ranger-hive-security.xml, which names the Ranger service, the policy REST URL, the local policy cache directory and the poll interval.

The client side

Users and jobs authenticate with kinit (people) or a keytab (scheduled jobs), then connect with a URL that names the server's principal, not their own:

kinit alice@EXAMPLE.COM
beeline -u "jdbc:hive2://hs2-1.example.com:10000/default;principal=hive/_HOST@EXAMPLE.COM;saslQop=auth-conf"

# With ZooKeeper discovery the principal is still the SERVER's, not the user's
beeline -u "jdbc:hive2://zk1:2181,zk2:2181,zk3:2181/;serviceDiscoveryMode=zooKeeper;zooKeeperNamespace=hiveserver2;principal=hive/_HOST@EXAMPLE.COM"

The driver substitutes _HOST with the host it connects to, which is what makes ZooKeeper discovery work across several HiveServer2 instances. A common mistake is putting the user's principal in the URL; the result is a GSS error that looks like a KDC problem. Scheduled jobs should use their own service principals and keytabs, never a person's, so access can be audited and revoked per job.

From principal to user name: auth_to_local

Ranger policies and HDFS permissions name short users, not principals, so the mapping rule decides who alice is. Rules are evaluated in order and the first match wins:

<!-- core-site.xml: map principals to short names; first matching rule wins -->
<property>
  <name>hadoop.security.auth_to_local</name>
  <value>
    RULE:[2:$1@$0](hive@EXAMPLE.COM)s/.*/hive/
    RULE:[1:$1@$0](.*@EXAMPLE.COM)s/@.*//
    RULE:[1:$1@$0](.*@CORP.EXAMPLE.COM)s/@.*//
    DEFAULT
  </value>
</property>

# Test a mapping and the groups Hadoop resolves for the result
hadoop org.apache.hadoop.security.HadoopKerberosName alice@CORP.EXAMPLE.COM
hdfs groups alice

The first rule maps two-component hive service principals from every host to the single user hive. The next two strip the realm for user principals from the cluster realm and a trusted corporate realm. Without an explicit rule for a trusted realm, users from a cross-realm trust either fail to map or map to names that match no policy. Test every rule change with the HadoopKerberosName class before deploying, and remember that HiveServer2, the metastore and HDFS all read core-site.xml, so they must be restarted together after a change.

Groups: where most denials actually come from

Policies usually grant to groups. Two independent systems resolve groups. Hadoop's group mapping (operating-system lookup through SSSD, or direct LDAP) is what HiveServer2 passes to the Ranger plugin at access time. Ranger usersync copies users and groups from LDAP or Active Directory into Ranger Admin so administrators can pick them in the policy UI. If the two disagree, the policy UI shows alice in analysts while HiveServer2 resolves her to nothing, and the query is denied by a policy that looks correct.

Make both point at the same directory with the same filters, compare their output for a sample of users after every change, and note that Hadoop caches group lookups for a configurable period, so a new group membership can take minutes to apply. Ranger on Hadoop covers usersync, policy priorities and deny exceptions in more depth.

Ranger in the request path

The plugin evaluates cached policies in process, so Ranger Admin is not on the query path. It polls for changes, by default every 30 seconds, and keeps the last good copy on local disk, so HiveServer2 continues to enforce the previous policy set if Ranger Admin is down. That is good for availability, and it also means a revoked grant takes effect only after the next successful poll on every HiveServer2 instance. Watch policy download times per instance and alert on stale caches.

Every decision is written to the audit store, usually Solr for search and HDFS for retention, with the user, resource, action, result and the policy ID. When a user reports a denial, the audit entry tells you which user name and resource were evaluated, which is often the fastest way to find a mapping problem.

doAs, delegation tokens and the data layer

With doAs off, queries run as hive, and Ranger is the only gate; HDFS permissions on the warehouse should then allow only the hive user, or anyone could read the files directly with an HDFS client and bypass Ranger. With doAs on, tasks run as the end user and HDFS permissions apply too, but column masks and row filters cannot be enforced, because the user could read the raw files. For Ranger-based fine-grained access, doAs off with a locked-down warehouse is the standard choice.

Tasks authenticate to HDFS with delegation tokens. In Hadoop's defaults a token must be renewed every 24 hours and expires after seven days regardless of renewal, so very long jobs and long-lived sessions eventually fail with token errors. The HiveServer2 architecture page explains how sessions and operations hold those resources.

Rolling it out on a live cluster

  1. Stand up the KDC or join the Active Directory realm, fix forward and reverse DNS, and synchronise clocks with NTP on every host.
  2. Create principals and keytabs, and validate each with klist and kinit on its own host.
  3. Kerberise HDFS and YARN first, then the metastore, then HiveServer2, testing each with a known user.
  4. Write and test auth_to_local rules and group mapping; compare Hadoop's groups with usersync's for real users.
  5. Install the Ranger plugin with a policy that mirrors existing access, verify through audit logs, and only then tighten.
  6. Move jobs to their own keytabs, lock warehouse directories to the hive user and switch doAs off.
  7. Turn on auth-conf, then remove any remaining unauthenticated endpoints.

Debugging runbook

Error or symptomUsual causeCheck
Clock skew too greatHost clocks differ by more than the KDC allows, five minutes by defaultNTP status on client, server and KDC
Server not found in Kerberos databasePrincipal hostname differs from the name the client resolvedReverse DNS, krb5.conf canonicalisation, the principal in the URL
Checksum failed or pre-authentication failedKeytab key version is stale after a ktaddklist -kte versus the KDC kvno; redeploy keytab
Failed to find any Kerberos tgtNo kinit, or ticket cache not visible to the processklist as that user; KRB5CCNAME; job keytab path
No rules applied to principalauth_to_local has no rule for that realmHadoopKerberosName with the exact principal
Permission denied with a correct-looking policyGroups resolved differently from usersynchdfs groups alice; Ranger audit user and groups
Token expired or not found in cacheDelegation token past its renewal or maximum lifetimeJob duration versus token lifetimes; session age

For anything opaque, turn on Kerberos tracing: KRB5_TRACE=/dev/stderr for MIT command-line tools, and -Dsun.security.krb5.debug=true in the JVM options for Beeline or HiveServer2. The trace shows the principal requested, the realm, the encryption types and the exact KDC response.

Trade-offs

Kerberos with Ranger is operationally heavy: a KDC to run, keytabs to rotate, DNS and clocks that must be right everywhere, and two group systems to keep in step. In return you get strong authentication without passwords on the wire, column and row control, and a central audit trail. LDAP authentication with TLS is simpler for SQL users but does nothing for service-to-service traffic or HDFS access. On cloud platforms, a central catalog with its own access control may replace both, but if Hive and Impala share a cluster, one Ranger policy set covering both engines, as in Impala and Ranger, is a strong reason to keep this design.

Key takeaway: <p><strong>What to do next.</strong> Secured Hive is a chain of translations from ticket to principal to user to groups to policy. Test each link on its own, and most incidents become a lookup in the runbook.</p><ol><li>Draw your cluster's identity chain and name the configuration that controls each hop.</li><li>Validate every keytab with klist -kte and kinit on its host, and record each key version.</li><li>Test auth_to_local with HadoopKerberosName for every realm that users come from.</li><li>Compare Hadoop group mapping with Ranger usersync for a sample of real users.</li><li>Lock the warehouse to the hive user, keep doAs off and set the SASL QOP to auth-conf.</li><li>Alert on stale Ranger policy caches, clock drift and keytab or token expiry before users notice.</li></ol>