A secured Impala cluster answers three questions for every query, and each is handled by a different system. Who are you? Kerberos, or LDAP for some clients. What name and groups does that identity map to? Hadoop's auth_to_local rules and group mapping. Is that user allowed to do this? Apache Ranger. A fourth fact surprises many teams: whatever the answers, Impala reads and writes the data as its own service user, not as the end user.
Ranger policy semantics, masking and row filters for Impala are covered in Impala + Ranger Integration, and the auth_to_local and group details for Hive in Hive Security with Kerberos and Ranger. This article follows one identity through Impala specifically: the principals each daemon needs, what changes when a load balancer sits in front, how impala-shell and BI tools authenticate, how delegation works for Hue, where Ranger enters, and a runbook for the errors you will actually see. Flag names were checked against the Apache Impala documentation; anything not confirmed there is left out.
One identity, end to end
Kerberos never sends a password over the network. Alice runs kinit, proves knowledge of her key to the KDC and receives a ticket-granting ticket. When impala-shell connects, it asks the KDC for a service ticket for a specific service principal name, such as impala/impala-lb.corp.example.com@CORP.EXAMPLE.COM, and presents it. The server decrypts it with its key from a keytab file. If that works, the server knows the caller is alice@CORP.EXAMPLE.COM. Kerberos stops there; everything else is Impala's job.
The principle behind every configuration step below is that the service principal the client asks for must exist in the KDC and its key must be in the keytab of the process that answers. Most Impala Kerberos outages are a mismatch on one of those two points. The general Hadoop background, including ticket lifetimes and delegation tokens, is in Hadoop Kerberos architecture.
Principals and keytabs for every daemon
Each Impala daemon type, impalad, catalogd and statestored, authenticates with Kerberos to the others, and impalad also authenticates clients. Each needs a principal and a keytab, set with startup flags. The conventional principal is impala/<fqdn>@REALM, where the host part must be the fully qualified name that forward and reverse DNS agree on.
# /etc/default/impala (or your cluster manager's equivalent), per host
IMPALA_SERVER_ARGS=" \
--principal=impala/node07.corp.example.com@CORP.EXAMPLE.COM \
--keytab_file=/etc/impala/conf/impala.keytab \
--load_auth_to_local_rules=true"
# Check the keytab holds that principal, and note the key version number (KVNO)
klist -kt /etc/impala/conf/impala.keytabImpala renews its own credentials from the keytab; the renewal frequency is controlled by --kerberos_reinit_interval, and the defaults are fine for most clusters. Two infrastructure prerequisites cause more failures than any flag: clocks must be synchronised, because MIT Kerberos rejects tickets outside its allowed clock skew, which defaults to five minutes, and the keytab must be readable only by the impala user. A keytab is a password on disk; anyone who can read it can act as Impala.
Putting a load balancer in front
Clients should not connect to a single coordinator. A proxy such as HAProxy spreads coordinator work and hides failed hosts, but it creates a Kerberos problem. The client connects to impala-lb.corp.example.com and asks for a ticket for that host's service principal. The coordinator that receives the connection must hold the key for the load balancer's principal, not only its own. Impala solves this with two flags. --principal is set to the proxy's principal on every coordinator behind it, and is what clients authenticate against. --be_principal is set to the host's own principal and is used for internal daemon-to-daemon traffic, so it differs on every host. The keytab given in --keytab_file must contain both keys, which means merging them:
# on each coordinator, as root, with both keytabs present
ktutil
rkt /tmp/impala-lb.keytab # key for impala/impala-lb.corp.example.com
rkt /etc/impala/conf/impala.keytab # key for impala/node07.corp.example.com
wkt /etc/impala/conf/impala-merged.keytab
quit
IMPALA_SERVER_ARGS=" \
--principal=impala/impala-lb.corp.example.com@CORP.EXAMPLE.COM \
--be_principal=impala/node07.corp.example.com@CORP.EXAMPLE.COM \
--keytab_file=/etc/impala/conf/impala-merged.keytab"The Impala proxy documentation suggests different balancing for different client types. Connections from the legacy impala-shell Beeswax port 21000 can use balance leastconn. The HiveServer2 port used by Hue, JDBC and ODBC, exposed on the proxy at 21051 and forwarded to 21050, should use balance source so a client's session stays on one coordinator, because session state such as query handles lives in that coordinator. The same documentation uses one-hour client and server timeouts; short proxy defaults cut off long-running queries and idle BI sessions. Run the proxy in TCP mode so it passes the Kerberos exchange through untouched.
listen impala-hs2
bind 0.0.0.0:21051
mode tcp
balance source
timeout client 3600s
timeout server 3600s
server node07 node07.corp.example.com:21050 check
server node08 node08.corp.example.com:21050 check
Clients: impala-shell, JDBC, ODBC and BI tools
For impala-shell, -k enables Kerberos. -i host[:port] picks the daemon, with 21050 as the default port in current versions. -s sets the service name if it is not the default, impala. -b (or --kerberos_host_fqdn) overrides the host name the shell expects in the server's principal, which is what you need when you connect through a DNS alias or an IP address that is not the name in the SPN. Add --ssl and --ca_cert in any production cluster: use TLS for encryption in transit rather than relying on Kerberos for it.
kinit alice@CORP.EXAMPLE.COM
impala-shell -k --ssl --ca_cert=/etc/pki/corp-ca.pem -i impala-lb.corp.example.com:21051
klist # should now list a ticket for impala/impala-lb.corp.example.com@CORP.EXAMPLE.COMJDBC and ODBC drivers expose the same three ingredients, realm, host FQDN and service name, under driver-specific property names; take those from your driver's documentation rather than copying a URL from another vendor. BI servers that cannot handle Kerberos for each end user are usually better served by LDAP authentication, -l in impala-shell, over TLS, or by delegation as described next.
From principal to Ranger user
Ranger policies name users such as alice, not principals such as alice@CORP.EXAMPLE.COM. The translation is done by auth_to_local rules in Hadoop's core-site.xml. Impala only applies those rules when --load_auth_to_local_rules=true is set on impalad and catalogd; it is off by default. Turn it on, so that Impala, HDFS and Hive agree on the short name for every principal. Otherwise a user from a second realm, or a service principal, can map to a different name in Impala than elsewhere, and policies that work in Hive will not match in Impala. Test a rule with Hadoop's own resolver before relying on it:
<!-- core-site.xml: strip the realm for users from the corporate realm -->
<property>
<name>hadoop.security.auth_to_local</name>
<value>
RULE:[1:$1@$0](.*@CORP\.EXAMPLE\.COM)s/@.*//
DEFAULT
</value>
</property>
$ hadoop org.apache.hadoop.security.HadoopKerberosName alice@CORP.EXAMPLE.COM
Name: alice@CORP.EXAMPLE.COM to aliceGroups come next. The impalad resolves the short name to groups using the Hadoop group mapping configured on that host, typically operating system groups backed by SSSD or a direct LDAP mapping. If alice is missing a group on one coordinator, she will be denied on that host only, which looks random behind a load balancer. Check id alice on every coordinator when a group policy behaves inconsistently.
Delegation for Hue and other services
Hue authenticates to Impala once, as the hue service principal, and then runs queries on behalf of many users. Impala supports this through delegation: the authenticated user may name a different effective user, and Ranger evaluates the effective user. You list who may delegate with --authorized_proxy_user_config or --authorized_proxy_group_config, using the form authenticated_user=delegated_user1,delegated_user2 with semicolons between entries. The client requests a user with the HiveServer2 session property impala.doas.user (or DelegationUID), or a doAs parameter on HTTP connections. Impala requires Ranger to be enabled for delegation.
# allow hue to act for anyone in the analysts group only
--authorized_proxy_group_config=hue=analystsThe documentation's example, hue=*, lets Hue act as any user, including administrators. Prefer a group list, and treat the Hue host and its keytab as being as sensitive as a Ranger admin account, because whoever controls them can query as anyone on the list.
Ranger in the chain
Authorization is enabled on every impalad and on catalogd with the same flags: -server_name, identical on all of them, -ranger_service_type=hive, -ranger_app_id and -authorization_provider=ranger. Impala shares the Hive service's policies, so one policy set governs both engines. In a Kerberized cluster the plugin also authenticates to Ranger Admin to download policies, and the service definition in Ranger must allow the impala user to download them; if downloads fail, check the plugin status in Ranger Admin and the impalad log. Policy changes made in Ranger Admin arrive on the polling interval set by ranger.plugin.hive.policy.pollIntervalMs; REFRESH AUTHORIZATION forces an immediate reload. Policy structure and evaluation order are covered in Ranger Policy Deep Dive.
Data access happens as impala
The Impala documentation is explicit that, regardless of how the user authenticated, Impala creates directories and files owned by its own service user. Executors read files as impala too. Three consequences follow. First, Ranger inside Impala is the enforcement point for Impala queries; HDFS permissions or Ranger HDFS policies must grant the impala user access to every table location, or queries fail even for allowed users. Second, anyone who can read the files directly, through HDFS commands, Spark or a shared object-store credential, bypasses Impala's policies entirely, so storage access must be locked down separately. Third, audit trails must come from Ranger's Impala audit, not from HDFS audit logs, which will show impala for every read.
Debugging runbook
| Symptom | Likely cause | Check |
|---|---|---|
| Server not found in Kerberos database | Client asked for an SPN that does not exist, often via an IP or alias | klist after the attempt; use the FQDN in the SPN or -b |
| Clock skew too great | Host clocks differ by more than the allowed skew | chrony or NTP status on client, KDC and daemons |
| Checksum failed or decrypt integrity check failed | Keytab KVNO is stale after a key rotation | klist -kt keytab vs kvno principal |
| Works on one coordinator, fails via the proxy | Merged keytab missing the LB key, or --principal not set on that host | klist -kt on every coordinator |
| Daemons cannot talk to each other | --be_principal wrong or missing from the keytab | impalad log at startup |
| AuthorizationException for an allowed user | Short name or groups differ from what the policy names | HadoopKerberosName and id on that coordinator |
| Policy change has no effect | Plugin cache not yet refreshed, or download denied | REFRESH AUTHORIZATION; plugin status in Ranger Admin |
For client-side tracing with MIT Kerberos, set KRB5_TRACE=/dev/stderr before kinit or impala-shell to see each ticket request and which SPN was asked for. That one line usually settles the first four rows of the table.
Trade-offs
- Kerberos versus LDAP. Kerberos gives single sign-on and no passwords on the wire, but desktop BI tools struggle with it. LDAP over TLS is simpler for those tools and puts a password through Impala; many clusters run both.
- Sticky versus balanced. balance source keeps sessions intact but can pile a large BI server's traffic onto one coordinator. Add coordinators or split BI traffic onto a dedicated proxy listener.
- Wildcard delegation is convenient and turns one service keytab into everyone's credentials.
- Service-user data access keeps Impala's daemons simple, since no per-user credentials are needed to read files, but means storage must be closed to every other path. How the daemons share that work is described in Impala Architecture.
What to do next
- List every impalad, catalogd and statestored host and verify with klist -kt that each keytab holds the expected principal and current KVNO.
- If a proxy is in front, confirm each coordinator sets --principal to the proxy SPN, --be_principal to its own, and uses a merged keytab; test a kinit plus impala-shell -k through the proxy.
- Set --load_auth_to_local_rules=true on impalad and catalogd and test your rules with HadoopKerberosName.
- Run id for a sample of users on every coordinator and compare group lists.
- Replace wildcard proxy-user entries with group-scoped ones.
- Confirm the impala user can download policies in Ranger Admin and document when to use REFRESH AUTHORIZATION.
- Audit storage permissions so that only the impala service user, and deliberate exceptions, can read table locations directly.
- Enable TLS for clients, and keep KRB5_TRACE in the on-call runbook.