Impala avoids per-query startup costs by using long-running daemons, a C++ backend and metadata cached on every coordinator. Security has to fit that design. Impala reads data files as its own service user, so the file system cannot tell one analyst from another. Every per-user decision about who can see which database, table, column or row has to happen inside Impala, and since Impala 4.0 the only supported authorization provider is Apache Ranger; Sentry support was removed.
This article explains where Ranger plugs into Impala, what each daemon does, how Impala shares policies with Hive, how column masking and row filters are enforced, and what goes wrong in production. It assumes you know what Ranger is; for the Hive side of the same policies, and the general model of authentication before authorization, read Hive authorization first.
The architecture: who checks what
Of Impala's three daemon types, described in Impala architecture, Ranger touches two: impalad and catalogd.
The coordinator is where access decisions are made. Its Java frontend parses and analyzes the statement, and analysis produces a list of privilege requests: SELECT on these columns of this table, INSERT on that table, CREATE in that database. The embedded Ranger plugin evaluates the list against its cached policies. If any request is denied, the query fails with an authorization error before a plan is produced, so no fragment ever reaches an executor. Executors do not consult Ranger at all; they run what the coordinator sends them.
catalogd also loads the plugin, because statements such as GRANT, REVOKE and ownership changes flow through it, and it needs the same server name and service configuration to translate them into Ranger policy updates. Getting catalogd's configuration out of step with the impalads is one of the classic mistakes covered below.
Turning it on
Ranger authorization is enabled with startup flags that must be set identically on every impalad and on catalogd:
# /etc/default/impala or the equivalent in your cluster manager
IMPALA_SERVER_ARGS="${IMPALA_SERVER_ARGS} \
-server_name=server1 \
-authorization_provider=ranger \
-ranger_service_type=hive \
-ranger_app_id=impala"
IMPALA_CATALOG_ARGS="${IMPALA_CATALOG_ARGS} \
-server_name=server1 \
-authorization_provider=ranger \
-ranger_service_type=hive \
-ranger_app_id=impala"Each flag has a job. -authorization_provider=ranger selects Ranger. -ranger_service_type=hive tells the plugin to use Ranger's Hive service definition, which is how Impala and Hive end up enforcing the same policies. -ranger_app_id labels this component in audit records, so you can tell an Impala access from a HiveServer2 access to the same table. -server_name is the name used for server-level privileges and must match across the cluster. The plugin's own settings, such as the Ranger Admin URL, the Ranger service name and the policy cache directory, live in ranger-hive-security.xml on the Impala hosts. After changing flags, restart catalogd and every impalad; a partially restarted cluster enforces two different regimes.
One policy set for two engines
Because Impala uses the Hive service type, a single Ranger service holds policies for both HiveServer2 and Impala. That is the main operational win: you define 'the finance group can SELECT from finance.*' once and both engines honour it. It also means differences between the engines become policy surprises rather than configuration errors.
The resource hierarchy is the one you would expect: server, then database, table and column, plus URIs for file locations. Impala's privilege set is SELECT, INSERT, CREATE, ALTER, DROP, REFRESH and ALL. REFRESH is the Impala-specific one: it governs REFRESH, INVALIDATE METADATA and related metadata statements, which matter because Impala caches metadata aggressively (see the Impala catalog). Give it to the pipeline accounts that land data, not to every analyst, since a flood of global invalidations can hurt the whole cluster.
| Statement | Privilege it needs | Note |
|---|---|---|
| SELECT, WITH, EXPLAIN | SELECT on the columns referenced | Column-level policies apply |
| INSERT, INSERT OVERWRITE | INSERT on the table | Plus SELECT on any source tables |
| CREATE TABLE ... LOCATION | CREATE on the database, ALL on the URI | URI must be fully qualified, e.g. hdfs:// or s3a:// |
| ALTER TABLE | ALTER on the table | Ownership can grant this implicitly |
| DROP TABLE | DROP on the table | Owners are covered by owner policies |
| REFRESH, INVALIDATE METADATA t | REFRESH on the table | Global INVALIDATE needs server scope |
| COMPUTE STATS | ALTER and SELECT on the table | Check your version's privilege table |
Object ownership is on by default: the creator of a database, table or view becomes its owner, and owner-based policies in Ranger can grant owners full control without a separate policy per object. Ownership moves with ALTER TABLE ... SET OWNER. Decide deliberately whether you want this; in regulated schemas many teams keep objects owned by a service role so that individuals never accumulate rights by creating tables.
Metadata statements are filtered too. SHOW DATABASES and SHOW TABLES only list objects on which the user holds some privilege, so a missing table in a listing is often an authorization symptom, not a catalog bug.
The request path, step by step
Step two is where most real incidents begin. Ranger policies usually name groups, but the enforcing decision uses the groups that the coordinator resolves for the user, through Hadoop's group mapping configured on that host. Ranger usersync populates the Ranger Admin UI from LDAP or Unix, which is what administrators see when they write policies. If the coordinator's group lookup and usersync disagree, a policy that looks right in the UI will not match at query time. The Kerberos principal is also shortened to a user name using the auth_to_local rules; a rule that differs between Hive hosts and Impala hosts produces a user who is allowed in one engine and denied in the other.
Impersonation is the other subtlety. Hue and other gateways connect as a service user and ask Impala to run the query as the end user. impalad only honours that when the service user is listed in --authorized_proxy_user_config, and the authorization check then uses the delegated user. Keep that list short and explicit; a wildcard there lets anyone who can authenticate as the proxy user act as anyone else.
Column masking and row filtering
Ranger supports two policy types beyond plain access: data masking policies attached to columns and row-level filter policies attached to tables. Impala 4.0 added support for Ranger row-filtering policies, and Impala enforces column masking with built-in mask types, including MASK, MASK_SHOW_LAST_4, MASK_HASH and MASK_NULL. One limitation is documented explicitly: Impala does not run Hive GenericUDFs, so masking expressions that call Hive's mask UDFs will not work in Impala even though they work in HiveServer2. If you share one service between both engines, stick to the built-in mask types or expressions both engines can evaluate.
Enforcement works by rewriting. When the analyzer sees a table reference that has a mask or filter for the current user, it replaces that reference with an inline view that applies the mask expressions to the affected columns and adds the filter as a WHERE predicate. The planner then optimizes the rewritten query as usual, which is why a row filter on a partition column can still prune partitions. Executors only ever receive the rewritten plan, so the unmasked values are never materialized for that user.
Test the consequences. A filter that references another table becomes a subquery and can change performance dramatically, and a hashed column changes join and GROUP BY results. Keep the matrix of groups and masks small and documented.
Worked example: a sales table with masks and a region filter
Consider sales.orders with columns order_id, region, customer_email, card_number and amount. Requirements: the analysts group can query the table; analysts in the EU team see only EU rows; nobody except payments sees full card numbers; emails are hashed for analysts. The three Ranger policies below express that through the Ranger Admin REST API. Policy type 0 is access, 1 is data mask and 2 is row filter; the service name is whatever your Hive service is called in Ranger.
# 1. Access policy: SELECT on the table for analysts and payments
curl -u admin -H 'Content-Type: application/json' \
-X POST https://ranger.example.com:6182/service/public/v2/api/policy -d '{
"service": "cm_hive", "name": "sales_orders_read", "policyType": 0,
"resources": {"database": {"values": ["sales"]},
"table": {"values": ["orders"]}, "column": {"values": ["*"]}},
"policyItems": [{"groups": ["analysts", "payments"],
"accesses": [{"type": "select", "isAllowed": true}]}]}'
# 2. Mask policy: last four digits of card_number for analysts
{ "service": "cm_hive", "name": "orders_card_mask", "policyType": 1,
"resources": {"database": {"values": ["sales"]}, "table": {"values": ["orders"]},
"column": {"values": ["card_number"]}},
"dataMaskPolicyItems": [{"groups": ["analysts"],
"accesses": [{"type": "select", "isAllowed": true}],
"dataMaskInfo": {"dataMaskType": "MASK_SHOW_LAST_4"}}]}
# 3. Row filter: EU analysts see EU rows only
{ "service": "cm_hive", "name": "orders_eu_rows", "policyType": 2,
"resources": {"database": {"values": ["sales"]}, "table": {"values": ["orders"]}},
"rowFilterPolicyItems": [{"groups": ["analysts_eu"],
"accesses": [{"type": "select", "isAllowed": true}],
"rowFilterInfo": {"filterExpr": "region = 'EU'"}}]}Add an email mask with MASK_HASH the same way. Then verify from Impala as a member of each group, rather than trusting the UI:
-- as an analysts_eu member
REFRESH AUTHORIZATION; -- pull the new policies now instead of waiting for the poll
SELECT region, count(*) FROM sales.orders GROUP BY region; -- expect only EU
SELECT card_number FROM sales.orders LIMIT 3; -- expect masked values
EXPLAIN SELECT * FROM sales.orders; -- look for the region predicateThe EXPLAIN output is the useful proof: the row filter should appear as a predicate on the scan, and if region is a partition column you should see partition pruning. Write these queries into an automated test that runs as a dedicated test principal in each group after every policy change.
Policy caching and refresh
The plugin does not call Ranger Admin per query. It downloads the service's policies, keeps them in memory and in a local cache directory, and polls for changes on the interval set by ranger.plugin.hive.policy.pollIntervalMs in ranger-hive-security.xml. Evaluation is therefore a local, in-process operation and adds little latency. The price is propagation delay: a revoked grant keeps working until the next poll on every coordinator. REFRESH AUTHORIZATION forces an immediate refresh, and INVALIDATE METADATA also triggers one. Privilege changes made from inside Impala with GRANT and REVOKE do not need a separate invalidation.
The cache is also your resilience story. If Ranger Admin is down, coordinators keep enforcing the last policies they downloaded, and a restarted daemon reloads them from the cache directory. A new coordinator with an empty cache and no reachable Ranger Admin has no policies and denies everything. That fail-closed behaviour is correct but will look like an outage, so monitor Ranger Admin like a production dependency and keep the cache directory on persistent disk.
Failure modes
| Symptom | Likely cause | What to check |
|---|---|---|
| Allowed in Hive, denied in Impala | Different group mapping or auth_to_local rules on Impala hosts | Resolve groups for the user on the coordinator host |
| Revoked user still queries | Policy poll not yet run on some coordinators | REFRESH AUTHORIZATION; review poll interval |
| GRANT succeeds but has no effect | catalogd and impalads use different server_name or service | Compare flags on every daemon |
| Masked column errors only in Impala | Mask expression uses a Hive UDF | Use built-in mask types |
| Everything denied after a rebuild | Empty policy cache and Ranger Admin unreachable | Admin health, network path, cache directory |
| Queries slow after a filter change | Row filter with a subquery or non-sargable expression | EXPLAIN the rewritten plan |
| Users read data outside Impala | Files readable directly on HDFS or S3 | Lock storage to service users; Ranger HDFS policies |
The last row is the one auditors find. Impala authorization protects the SQL path only. If analysts can also read the table's files with a Spark job, the HDFS CLI or an S3 client, masks and row filters do nothing. Restrict warehouse paths to the service users, and use Ranger's HDFS or cloud storage policies for the remaining direct access. Hive security covers these side doors in more depth.
Operating it
Audit volume grows with query volume: every access decision generates an event, and a busy Impala cluster can produce millions a day. Size the audit store accordingly, and filter noisy service accounts through Ranger's audit filters rather than turning audit off. Track policy count too; hundreds of fine-grained per-user policies are slower to evaluate and far harder to review than a few group-based ones.
Treat policies as code: export them, keep them in version control, apply changes through the REST API and run the verification queries afterwards, and again after every Impala upgrade. Keep statestore and catalog health on the same dashboard as Ranger Admin, since missing metadata and missing policies look alike to users.
What to do next
- Confirm that every impalad and catalogd runs with identical
-server_name,-authorization_provider,-ranger_service_typeand-ranger_app_idvalues. - Pick a test user per group and compare group resolution on an Impala coordinator, a HiveServer2 host and Ranger usersync.
- Replace any mask expressions that call Hive UDFs with built-in mask types.
- Write verification queries, including EXPLAIN, for each mask and row filter, and run them after every policy change and upgrade.
- Restrict warehouse storage so only service principals can read table files directly.
- Export Ranger policies to version control and apply changes through the REST API.
- Alert on Ranger Admin availability and keep the plugin cache directory on persistent storage.