A secured Impala cluster answers three questions for every query, and each is handled by a different system. Who are you? Kerberos, or LDAP for some clients. What name and groups does that identity map to? Hadoop's auth_to_local rules and group mapping. Is that user allowed to do this? Apache Ranger. A fourth fact surprises many teams: whatever the answers, Impala reads and writes the data as its own service user, not as the end user.

Ranger policy semantics, masking and row filters for Impala are covered in Impala + Ranger Integration, and the auth_to_local and group details for Hive in Hive Security with Kerberos and Ranger. This article follows one identity through Impala specifically: the principals each daemon needs, what changes when a load balancer sits in front, how impala-shell and BI tools authenticate, how delegation works for Hue, where Ranger enters, and a runbook for the errors you will actually see. Flag names were checked against the Apache Impala documentation; anything not confirmed there is left out.

One identity, end to end

clientkinit aliceKDCTGT, service ticketAS / TGSticket for LB SPNload balancerTCP pass-throughcoordinator impalad--principal = LB SPN--be_principal = host SPNshort name + groupsRanger pluginhive service policiesRanger Adminpolicy downloadfragmentsexecutor impaladsKerberos between daemonsHDFS / storageaccessed as impalacatalogdsame Ranger flagsstatestoredown principalAlice is authenticated once, authorized by Ranger as alice, and the data is read as impala.
One query's identity path. The client obtains a ticket for the load balancer's service principal; the coordinator accepts it, maps the principal to a short name and groups, asks the Ranger plugin, and the executors read storage as the impala service user.

Kerberos never sends a password over the network. Alice runs kinit, proves knowledge of her key to the KDC and receives a ticket-granting ticket. When impala-shell connects, it asks the KDC for a service ticket for a specific service principal name, such as impala/impala-lb.corp.example.com@CORP.EXAMPLE.COM, and presents it. The server decrypts it with its key from a keytab file. If that works, the server knows the caller is alice@CORP.EXAMPLE.COM. Kerberos stops there; everything else is Impala's job.

The principle behind every configuration step below is that the service principal the client asks for must exist in the KDC and its key must be in the keytab of the process that answers. Most Impala Kerberos outages are a mismatch on one of those two points. The general Hadoop background, including ticket lifetimes and delegation tokens, is in Hadoop Kerberos architecture.

Principals and keytabs for every daemon

Each Impala daemon type, impalad, catalogd and statestored, authenticates with Kerberos to the others, and impalad also authenticates clients. Each needs a principal and a keytab, set with startup flags. The conventional principal is impala/<fqdn>@REALM, where the host part must be the fully qualified name that forward and reverse DNS agree on.

# /etc/default/impala (or your cluster manager's equivalent), per host
IMPALA_SERVER_ARGS=" \
  --principal=impala/node07.corp.example.com@CORP.EXAMPLE.COM \
  --keytab_file=/etc/impala/conf/impala.keytab \
  --load_auth_to_local_rules=true"

# Check the keytab holds that principal, and note the key version number (KVNO)
klist -kt /etc/impala/conf/impala.keytab

Impala renews its own credentials from the keytab; the renewal frequency is controlled by --kerberos_reinit_interval, and the defaults are fine for most clusters. Two infrastructure prerequisites cause more failures than any flag: clocks must be synchronised, because MIT Kerberos rejects tickets outside its allowed clock skew, which defaults to five minutes, and the keytab must be readable only by the impala user. A keytab is a password on disk; anyone who can read it can act as Impala.

Putting a load balancer in front

Clients should not connect to a single coordinator. A proxy such as HAProxy spreads coordinator work and hides failed hosts, but it creates a Kerberos problem. The client connects to impala-lb.corp.example.com and asks for a ticket for that host's service principal. The coordinator that receives the connection must hold the key for the load balancer's principal, not only its own. Impala solves this with two flags. --principal is set to the proxy's principal on every coordinator behind it, and is what clients authenticate against. --be_principal is set to the host's own principal and is used for internal daemon-to-daemon traffic, so it differs on every host. The keytab given in --keytab_file must contain both keys, which means merging them:

# on each coordinator, as root, with both keytabs present
ktutil
  rkt /tmp/impala-lb.keytab        # key for impala/impala-lb.corp.example.com
  rkt /etc/impala/conf/impala.keytab   # key for impala/node07.corp.example.com
  wkt /etc/impala/conf/impala-merged.keytab
  quit

IMPALA_SERVER_ARGS=" \
  --principal=impala/impala-lb.corp.example.com@CORP.EXAMPLE.COM \
  --be_principal=impala/node07.corp.example.com@CORP.EXAMPLE.COM \
  --keytab_file=/etc/impala/conf/impala-merged.keytab"

The Impala proxy documentation suggests different balancing for different client types. Connections from the legacy impala-shell Beeswax port 21000 can use balance leastconn. The HiveServer2 port used by Hue, JDBC and ODBC, exposed on the proxy at 21051 and forwarded to 21050, should use balance source so a client's session stays on one coordinator, because session state such as query handles lives in that coordinator. The same documentation uses one-hour client and server timeouts; short proxy defaults cut off long-running queries and idle BI sessions. Run the proxy in TCP mode so it passes the Kerberos exchange through untouched.

listen impala-hs2
    bind 0.0.0.0:21051
    mode tcp
    balance source
    timeout client 3600s
    timeout server 3600s
    server node07 node07.corp.example.com:21050 check
    server node08 node08.corp.example.com:21050 check

Clients: impala-shell, JDBC, ODBC and BI tools

For impala-shell, -k enables Kerberos. -i host[:port] picks the daemon, with 21050 as the default port in current versions. -s sets the service name if it is not the default, impala. -b (or --kerberos_host_fqdn) overrides the host name the shell expects in the server's principal, which is what you need when you connect through a DNS alias or an IP address that is not the name in the SPN. Add --ssl and --ca_cert in any production cluster: use TLS for encryption in transit rather than relying on Kerberos for it.

kinit alice@CORP.EXAMPLE.COM
impala-shell -k --ssl --ca_cert=/etc/pki/corp-ca.pem -i impala-lb.corp.example.com:21051
klist          # should now list a ticket for impala/impala-lb.corp.example.com@CORP.EXAMPLE.COM

JDBC and ODBC drivers expose the same three ingredients, realm, host FQDN and service name, under driver-specific property names; take those from your driver's documentation rather than copying a URL from another vendor. BI servers that cannot handle Kerberos for each end user are usually better served by LDAP authentication, -l in impala-shell, over TLS, or by delegation as described next.

From principal to Ranger user

Ranger policies name users such as alice, not principals such as alice@CORP.EXAMPLE.COM. The translation is done by auth_to_local rules in Hadoop's core-site.xml. Impala only applies those rules when --load_auth_to_local_rules=true is set on impalad and catalogd; it is off by default. Turn it on, so that Impala, HDFS and Hive agree on the short name for every principal. Otherwise a user from a second realm, or a service principal, can map to a different name in Impala than elsewhere, and policies that work in Hive will not match in Impala. Test a rule with Hadoop's own resolver before relying on it:

<!-- core-site.xml: strip the realm for users from the corporate realm -->
<property>
  <name>hadoop.security.auth_to_local</name>
  <value>
    RULE:[1:$1@$0](.*@CORP\.EXAMPLE\.COM)s/@.*//
    DEFAULT
  </value>
</property>

$ hadoop org.apache.hadoop.security.HadoopKerberosName alice@CORP.EXAMPLE.COM
Name: alice@CORP.EXAMPLE.COM to alice

Groups come next. The impalad resolves the short name to groups using the Hadoop group mapping configured on that host, typically operating system groups backed by SSSD or a direct LDAP mapping. If alice is missing a group on one coordinator, she will be denied on that host only, which looks random behind a load balancer. Check id alice on every coordinator when a group policy behaves inconsistently.

Delegation for Hue and other services

Hue authenticates to Impala once, as the hue service principal, and then runs queries on behalf of many users. Impala supports this through delegation: the authenticated user may name a different effective user, and Ranger evaluates the effective user. You list who may delegate with --authorized_proxy_user_config or --authorized_proxy_group_config, using the form authenticated_user=delegated_user1,delegated_user2 with semicolons between entries. The client requests a user with the HiveServer2 session property impala.doas.user (or DelegationUID), or a doAs parameter on HTTP connections. Impala requires Ranger to be enabled for delegation.

# allow hue to act for anyone in the analysts group only
--authorized_proxy_group_config=hue=analysts

The documentation's example, hue=*, lets Hue act as any user, including administrators. Prefer a group list, and treat the Hue host and its keytab as being as sensitive as a Ranger admin account, because whoever controls them can query as anyone on the list.

Ranger in the chain

Authorization is enabled on every impalad and on catalogd with the same flags: -server_name, identical on all of them, -ranger_service_type=hive, -ranger_app_id and -authorization_provider=ranger. Impala shares the Hive service's policies, so one policy set governs both engines. In a Kerberized cluster the plugin also authenticates to Ranger Admin to download policies, and the service definition in Ranger must allow the impala user to download them; if downloads fail, check the plugin status in Ranger Admin and the impalad log. Policy changes made in Ranger Admin arrive on the polling interval set by ranger.plugin.hive.policy.pollIntervalMs; REFRESH AUTHORIZATION forces an immediate reload. Policy structure and evaluation order are covered in Ranger Policy Deep Dive.

Data access happens as impala

The Impala documentation is explicit that, regardless of how the user authenticated, Impala creates directories and files owned by its own service user. Executors read files as impala too. Three consequences follow. First, Ranger inside Impala is the enforcement point for Impala queries; HDFS permissions or Ranger HDFS policies must grant the impala user access to every table location, or queries fail even for allowed users. Second, anyone who can read the files directly, through HDFS commands, Spark or a shared object-store credential, bypasses Impala's policies entirely, so storage access must be locked down separately. Third, audit trails must come from Ranger's Impala audit, not from HDFS audit logs, which will show impala for every read.

Debugging runbook

SymptomLikely causeCheck
Server not found in Kerberos databaseClient asked for an SPN that does not exist, often via an IP or aliasklist after the attempt; use the FQDN in the SPN or -b
Clock skew too greatHost clocks differ by more than the allowed skewchrony or NTP status on client, KDC and daemons
Checksum failed or decrypt integrity check failedKeytab KVNO is stale after a key rotationklist -kt keytab vs kvno principal
Works on one coordinator, fails via the proxyMerged keytab missing the LB key, or --principal not set on that hostklist -kt on every coordinator
Daemons cannot talk to each other--be_principal wrong or missing from the keytabimpalad log at startup
AuthorizationException for an allowed userShort name or groups differ from what the policy namesHadoopKerberosName and id on that coordinator
Policy change has no effectPlugin cache not yet refreshed, or download deniedREFRESH AUTHORIZATION; plugin status in Ranger Admin

For client-side tracing with MIT Kerberos, set KRB5_TRACE=/dev/stderr before kinit or impala-shell to see each ticket request and which SPN was asked for. That one line usually settles the first four rows of the table.

Trade-offs

  • Kerberos versus LDAP. Kerberos gives single sign-on and no passwords on the wire, but desktop BI tools struggle with it. LDAP over TLS is simpler for those tools and puts a password through Impala; many clusters run both.
  • Sticky versus balanced. balance source keeps sessions intact but can pile a large BI server's traffic onto one coordinator. Add coordinators or split BI traffic onto a dedicated proxy listener.
  • Wildcard delegation is convenient and turns one service keytab into everyone's credentials.
  • Service-user data access keeps Impala's daemons simple, since no per-user credentials are needed to read files, but means storage must be closed to every other path. How the daemons share that work is described in Impala Architecture.

What to do next

  1. List every impalad, catalogd and statestored host and verify with klist -kt that each keytab holds the expected principal and current KVNO.
  2. If a proxy is in front, confirm each coordinator sets --principal to the proxy SPN, --be_principal to its own, and uses a merged keytab; test a kinit plus impala-shell -k through the proxy.
  3. Set --load_auth_to_local_rules=true on impalad and catalogd and test your rules with HadoopKerberosName.
  4. Run id for a sample of users on every coordinator and compare group lists.
  5. Replace wildcard proxy-user entries with group-scoped ones.
  6. Confirm the impala user can download policies in Ranger Admin and document when to use REFRESH AUTHORIZATION.
  7. Audit storage permissions so that only the impala service user, and deliberate exceptions, can read table locations directly.
  8. Enable TLS for clients, and keep KRB5_TRACE in the on-call runbook.
Key takeaway: In a secured Impala cluster Kerberos proves who the client is, auth_to_local and group mapping turn that into a user and groups, and Ranger decides, while the data itself is read as the impala service user. Most failures come from a service principal the client asks for that is missing from the answering daemon's keytab, especially behind a load balancer, or from names and groups that differ between hosts.