Impala avoids per-query startup costs by using long-running daemons, a C++ backend and metadata cached on every coordinator. Security has to fit that design. Impala reads data files as its own service user, so the file system cannot tell one analyst from another. Every per-user decision about who can see which database, table, column or row has to happen inside Impala, and since Impala 4.0 the only supported authorization provider is Apache Ranger; Sentry support was removed.

This article explains where Ranger plugs into Impala, what each daemon does, how Impala shares policies with Hive, how column masking and row filters are enforced, and what goes wrong in production. It assumes you know what Ranger is; for the Hive side of the same policies, and the general model of authentication before authorization, read Hive authorization first.

Advertisement

The architecture: who checks what

Where Ranger sits in an Impala clusterClientimpala-shell, JDBC, HueCoordinator impaladFrontendparse, analyze, planRanger pluginpolicy cacheRanger Adminhive service policiesAudit storeSolr / HDFScatalogdmetadata, GRANT/REVOKE, ownersHive Metastoretables, ownersExecutor impaladsscan HDFS / S3 / KuduStorage read as the impala service user: files must not be readable by end users directlyquerypollauditmetadatafragments
The Ranger plugin runs inside the coordinator's Java frontend and inside catalogd. It polls Ranger Admin for policies and keeps a local cache; audit events go to the configured audit store.

Of Impala's three daemon types, described in Impala architecture, Ranger touches two: impalad and catalogd.

The coordinator is where access decisions are made. Its Java frontend parses and analyzes the statement, and analysis produces a list of privilege requests: SELECT on these columns of this table, INSERT on that table, CREATE in that database. The embedded Ranger plugin evaluates the list against its cached policies. If any request is denied, the query fails with an authorization error before a plan is produced, so no fragment ever reaches an executor. Executors do not consult Ranger at all; they run what the coordinator sends them.

catalogd also loads the plugin, because statements such as GRANT, REVOKE and ownership changes flow through it, and it needs the same server name and service configuration to translate them into Ranger policy updates. Getting catalogd's configuration out of step with the impalads is one of the classic mistakes covered below.

Turning it on

Ranger authorization is enabled with startup flags that must be set identically on every impalad and on catalogd:

# /etc/default/impala or the equivalent in your cluster manager
IMPALA_SERVER_ARGS="${IMPALA_SERVER_ARGS} \
  -server_name=server1 \
  -authorization_provider=ranger \
  -ranger_service_type=hive \
  -ranger_app_id=impala"

IMPALA_CATALOG_ARGS="${IMPALA_CATALOG_ARGS} \
  -server_name=server1 \
  -authorization_provider=ranger \
  -ranger_service_type=hive \
  -ranger_app_id=impala"

Each flag has a job. -authorization_provider=ranger selects Ranger. -ranger_service_type=hive tells the plugin to use Ranger's Hive service definition, which is how Impala and Hive end up enforcing the same policies. -ranger_app_id labels this component in audit records, so you can tell an Impala access from a HiveServer2 access to the same table. -server_name is the name used for server-level privileges and must match across the cluster. The plugin's own settings, such as the Ranger Admin URL, the Ranger service name and the policy cache directory, live in ranger-hive-security.xml on the Impala hosts. After changing flags, restart catalogd and every impalad; a partially restarted cluster enforces two different regimes.

Advertisement

One policy set for two engines

Because Impala uses the Hive service type, a single Ranger service holds policies for both HiveServer2 and Impala. That is the main operational win: you define 'the finance group can SELECT from finance.*' once and both engines honour it. It also means differences between the engines become policy surprises rather than configuration errors.

The resource hierarchy is the one you would expect: server, then database, table and column, plus URIs for file locations. Impala's privilege set is SELECT, INSERT, CREATE, ALTER, DROP, REFRESH and ALL. REFRESH is the Impala-specific one: it governs REFRESH, INVALIDATE METADATA and related metadata statements, which matter because Impala caches metadata aggressively (see the Impala catalog). Give it to the pipeline accounts that land data, not to every analyst, since a flood of global invalidations can hurt the whole cluster.

StatementPrivilege it needsNote
SELECT, WITH, EXPLAINSELECT on the columns referencedColumn-level policies apply
INSERT, INSERT OVERWRITEINSERT on the tablePlus SELECT on any source tables
CREATE TABLE ... LOCATIONCREATE on the database, ALL on the URIURI must be fully qualified, e.g. hdfs:// or s3a://
ALTER TABLEALTER on the tableOwnership can grant this implicitly
DROP TABLEDROP on the tableOwners are covered by owner policies
REFRESH, INVALIDATE METADATA tREFRESH on the tableGlobal INVALIDATE needs server scope
COMPUTE STATSALTER and SELECT on the tableCheck your version's privilege table

Object ownership is on by default: the creator of a database, table or view becomes its owner, and owner-based policies in Ranger can grant owners full control without a separate policy per object. Ownership moves with ALTER TABLE ... SET OWNER. Decide deliberately whether you want this; in regulated schemas many teams keep objects owned by a service role so that individuals never accumulate rights by creating tables.

Metadata statements are filtered too. SHOW DATABASES and SHOW TABLES only list objects on which the user holds some privilege, so a missing table in a listing is often an authorization symptom, not a catalog bug.

The request path, step by step

One SELECT: authorization happens once, during analysis, before any fragment runs1. AuthenticateKerberos / LDAP2. Resolve usershort name + groups3. Analyze SQLcollect privilege requests4. Ranger checkaccess, mask, row filter5. Plan + runrewritten queryDenied: AuthorizationExceptionaudit record written, nothing executesMasks and row filters are applied by rewriting the table reference, so executors never see unmasked columns.
Authentication establishes a principal, the coordinator maps it to a short user name and groups, analysis collects privilege requests, Ranger decides, and only then is the plan built.

Step two is where most real incidents begin. Ranger policies usually name groups, but the enforcing decision uses the groups that the coordinator resolves for the user, through Hadoop's group mapping configured on that host. Ranger usersync populates the Ranger Admin UI from LDAP or Unix, which is what administrators see when they write policies. If the coordinator's group lookup and usersync disagree, a policy that looks right in the UI will not match at query time. The Kerberos principal is also shortened to a user name using the auth_to_local rules; a rule that differs between Hive hosts and Impala hosts produces a user who is allowed in one engine and denied in the other.

Impersonation is the other subtlety. Hue and other gateways connect as a service user and ask Impala to run the query as the end user. impalad only honours that when the service user is listed in --authorized_proxy_user_config, and the authorization check then uses the delegated user. Keep that list short and explicit; a wildcard there lets anyone who can authenticate as the proxy user act as anyone else.

Column masking and row filtering

Ranger supports two policy types beyond plain access: data masking policies attached to columns and row-level filter policies attached to tables. Impala 4.0 added support for Ranger row-filtering policies, and Impala enforces column masking with built-in mask types, including MASK, MASK_SHOW_LAST_4, MASK_HASH and MASK_NULL. One limitation is documented explicitly: Impala does not run Hive GenericUDFs, so masking expressions that call Hive's mask UDFs will not work in Impala even though they work in HiveServer2. If you share one service between both engines, stick to the built-in mask types or expressions both engines can evaluate.

Enforcement works by rewriting. When the analyzer sees a table reference that has a mask or filter for the current user, it replaces that reference with an inline view that applies the mask expressions to the affected columns and adds the filter as a WHERE predicate. The planner then optimizes the rewritten query as usual, which is why a row filter on a partition column can still prune partitions. Executors only ever receive the rewritten plan, so the unmasked values are never materialized for that user.

Test the consequences. A filter that references another table becomes a subquery and can change performance dramatically, and a hashed column changes join and GROUP BY results. Keep the matrix of groups and masks small and documented.

Worked example: a sales table with masks and a region filter

Consider sales.orders with columns order_id, region, customer_email, card_number and amount. Requirements: the analysts group can query the table; analysts in the EU team see only EU rows; nobody except payments sees full card numbers; emails are hashed for analysts. The three Ranger policies below express that through the Ranger Admin REST API. Policy type 0 is access, 1 is data mask and 2 is row filter; the service name is whatever your Hive service is called in Ranger.

# 1. Access policy: SELECT on the table for analysts and payments
curl -u admin -H 'Content-Type: application/json' \
  -X POST https://ranger.example.com:6182/service/public/v2/api/policy -d '{
  "service": "cm_hive", "name": "sales_orders_read", "policyType": 0,
  "resources": {"database": {"values": ["sales"]},
                "table": {"values": ["orders"]}, "column": {"values": ["*"]}},
  "policyItems": [{"groups": ["analysts", "payments"],
                   "accesses": [{"type": "select", "isAllowed": true}]}]}'

# 2. Mask policy: last four digits of card_number for analysts
{ "service": "cm_hive", "name": "orders_card_mask", "policyType": 1,
  "resources": {"database": {"values": ["sales"]}, "table": {"values": ["orders"]},
                "column": {"values": ["card_number"]}},
  "dataMaskPolicyItems": [{"groups": ["analysts"],
      "accesses": [{"type": "select", "isAllowed": true}],
      "dataMaskInfo": {"dataMaskType": "MASK_SHOW_LAST_4"}}]}

# 3. Row filter: EU analysts see EU rows only
{ "service": "cm_hive", "name": "orders_eu_rows", "policyType": 2,
  "resources": {"database": {"values": ["sales"]}, "table": {"values": ["orders"]}},
  "rowFilterPolicyItems": [{"groups": ["analysts_eu"],
      "accesses": [{"type": "select", "isAllowed": true}],
      "rowFilterInfo": {"filterExpr": "region = 'EU'"}}]}

Add an email mask with MASK_HASH the same way. Then verify from Impala as a member of each group, rather than trusting the UI:

-- as an analysts_eu member
REFRESH AUTHORIZATION;                   -- pull the new policies now instead of waiting for the poll
SELECT region, count(*) FROM sales.orders GROUP BY region;   -- expect only EU
SELECT card_number FROM sales.orders LIMIT 3;                -- expect masked values
EXPLAIN SELECT * FROM sales.orders;                          -- look for the region predicate

The EXPLAIN output is the useful proof: the row filter should appear as a predicate on the scan, and if region is a partition column you should see partition pruning. Write these queries into an automated test that runs as a dedicated test principal in each group after every policy change.

Policy caching and refresh

The plugin does not call Ranger Admin per query. It downloads the service's policies, keeps them in memory and in a local cache directory, and polls for changes on the interval set by ranger.plugin.hive.policy.pollIntervalMs in ranger-hive-security.xml. Evaluation is therefore a local, in-process operation and adds little latency. The price is propagation delay: a revoked grant keeps working until the next poll on every coordinator. REFRESH AUTHORIZATION forces an immediate refresh, and INVALIDATE METADATA also triggers one. Privilege changes made from inside Impala with GRANT and REVOKE do not need a separate invalidation.

The cache is also your resilience story. If Ranger Admin is down, coordinators keep enforcing the last policies they downloaded, and a restarted daemon reloads them from the cache directory. A new coordinator with an empty cache and no reachable Ranger Admin has no policies and denies everything. That fail-closed behaviour is correct but will look like an outage, so monitor Ranger Admin like a production dependency and keep the cache directory on persistent disk.

Failure modes

SymptomLikely causeWhat to check
Allowed in Hive, denied in ImpalaDifferent group mapping or auth_to_local rules on Impala hostsResolve groups for the user on the coordinator host
Revoked user still queriesPolicy poll not yet run on some coordinatorsREFRESH AUTHORIZATION; review poll interval
GRANT succeeds but has no effectcatalogd and impalads use different server_name or serviceCompare flags on every daemon
Masked column errors only in ImpalaMask expression uses a Hive UDFUse built-in mask types
Everything denied after a rebuildEmpty policy cache and Ranger Admin unreachableAdmin health, network path, cache directory
Queries slow after a filter changeRow filter with a subquery or non-sargable expressionEXPLAIN the rewritten plan
Users read data outside ImpalaFiles readable directly on HDFS or S3Lock storage to service users; Ranger HDFS policies

The last row is the one auditors find. Impala authorization protects the SQL path only. If analysts can also read the table's files with a Spark job, the HDFS CLI or an S3 client, masks and row filters do nothing. Restrict warehouse paths to the service users, and use Ranger's HDFS or cloud storage policies for the remaining direct access. Hive security covers these side doors in more depth.

Operating it

Audit volume grows with query volume: every access decision generates an event, and a busy Impala cluster can produce millions a day. Size the audit store accordingly, and filter noisy service accounts through Ranger's audit filters rather than turning audit off. Track policy count too; hundreds of fine-grained per-user policies are slower to evaluate and far harder to review than a few group-based ones.

Treat policies as code: export them, keep them in version control, apply changes through the REST API and run the verification queries afterwards, and again after every Impala upgrade. Keep statestore and catalog health on the same dashboard as Ranger Admin, since missing metadata and missing policies look alike to users.

What to do next

  1. Confirm that every impalad and catalogd runs with identical -server_name, -authorization_provider, -ranger_service_type and -ranger_app_id values.
  2. Pick a test user per group and compare group resolution on an Impala coordinator, a HiveServer2 host and Ranger usersync.
  3. Replace any mask expressions that call Hive UDFs with built-in mask types.
  4. Write verification queries, including EXPLAIN, for each mask and row filter, and run them after every policy change and upgrade.
  5. Restrict warehouse storage so only service principals can read table files directly.
  6. Export Ranger policies to version control and apply changes through the REST API.
  7. Alert on Ranger Admin availability and keep the plugin cache directory on persistent storage.
Key takeaway: Impala enforces Ranger policies inside the coordinator's frontend during analysis, with catalogd handling GRANT, REVOKE and ownership, so a denied query never reaches an executor. Because Impala uses Ranger's Hive service type, one policy set covers both engines; the catch is that user and group resolution, masking UDFs and supported statements differ between them. Masks and row filters work by rewriting the query, which is secure and optimizable but needs testing with EXPLAIN. Policies are cached and polled, so plan for propagation delay and for a fail-closed start when Ranger Admin is unreachable. And remember that Ranger in Impala protects only the SQL path; lock down the storage underneath it.