Most of what you will ever do to an Impala cluster by hand goes through two tools. impala-shell is the command-line client: it opens a session on a coordinator, runs SQL, and prints the summary and profile of what it ran. The Web UI is a small HTTP server built into each Impala daemon: it lists running and recent queries, sessions, backends, memory, metrics, logs and startup flags, and every page can also return JSON.

Knowing the tools well pays off in incidents and in automation. This page covers how to connect correctly (the default port changed, and many old guides still show the old one), how to script impala-shell safely, what each Web UI page is for, how to scrape it, why the query you are looking for has vanished from the UI, and how to keep the UI from becoming a security hole. Reading plans and profiles is covered by Impala query plans and Impala troubleshooting; here the focus is the tooling around them.

Advertisement

How the pieces fit

Every impalad can act as a coordinator, and each one runs its own web server. The statestore and catalog daemons run web servers too. impala-shell talks to a coordinator over HiveServer2; a browser or curl talks to any daemon's web port. The two views complement each other: the shell shows one query from the client's side, while the coordinator's Web UI shows every query that coordinator is running or recently ran. A query is visible only on the coordinator that ran it, so when a load balancer spreads sessions over several coordinators, you need to know which one served the query before you look.

Two windows onto the same daemonsimpala-shellSQL, PROFILE, SUMMARYbrowser / curlpages, ?json, Prometheusimpalad (coordinator)hs2 21050, hs2-http 28000, beeswax 21000web UI on 25000HS2 sessionHTTP 25000impalad (executors)web UI 25000 eachstatestoredweb UI 25010catalogdweb UI 25020The shell shows one query from the client sidethe Web UI shows every query, session, backend, metric and flag from the daemon sideCompleted querieskept in memory, bounded by query_log_sizeSave what you needPROFILE to a file, or archive profiles
impala-shell opens an HS2 session to one coordinator. Each daemon's own web server exposes its pages and their JSON form. Completed queries live in a bounded in-memory log, so save profiles you need to keep.

Connecting: protocols, ports and authentication

impala-shell's --protocol option defaults to hs2, HiveServer2 over binary TCP, which a coordinator serves on port 21050. Older releases defaulted to the Beeswax protocol on port 21000, and many tutorials still show -i host:21000; if you pass that port with the default protocol, the connection fails. The third option, hs2-http, carries HS2 over HTTP on port 28000, which suits proxies and gateways that only pass HTTP.

PortFlagUsed by
21050--hs2_portimpala-shell (default hs2 protocol), JDBC and ODBC
28000--hs2_http_portimpala-shell with --protocol=hs2-http, HTTP clients
21000--beeswax_portlegacy Beeswax clients
25000 / 25010 / 25020--webserver_portWeb UI of impalad / statestored / catalogd
# HS2 over binary TCP (the default protocol) to a coordinator
impala-shell -i coord1.example.com:21050

# HS2 over HTTP, e.g. through a proxy that only passes HTTP, with TLS
impala-shell -i impala-lb.example.com:28000 --protocol=hs2-http --ssl --ca_cert=/etc/pki/ca.pem

# Kerberos (run kinit first)
impala-shell -k -i coord1.example.com:21050 --ssl

# LDAP: -l enables it, -u names the user; the password is prompted for
impala-shell -l -u analyst1 -i coord1.example.com:21050 --ssl

Pass -i the host and port of a coordinator or of a load balancer in front of several. HS2 sessions are stateful, so the load balancer must keep a session on one coordinator for its lifetime. -k uses the Kerberos ticket from kinit. -l turns on LDAP and -u names the user; for unattended use, --ldap_password_cmd runs a command that prints the password, so it never appears on the command line. With TLS, --ssl encrypts the connection and --ca_cert points at the CA certificate, or at the server certificate for a self-signed setup. Sending LDAP passwords without TLS sends them in clear text.

Advertisement

Working interactively

In an interactive session, statements end with a semicolon. A few shell commands do most of the work:

  • SET with no arguments lists query options and their current values; SET mem_limit=4g; changes one for the rest of the session.
  • SUMMARY prints the execution summary of the last query: per operator, the number of hosts, average and maximum time, rows produced and estimated, and peak memory.
  • PROFILE prints the full runtime profile of the last query: the plan, the timeline and every counter for every fragment instance.
  • EXPLAIN before a statement shows the plan without running it.
  • Ctrl-C cancels the running query rather than closing the shell.

Two startup options make long queries less opaque: --live_progress shows a progress bar as scan ranges complete, and --live_summary redraws the execution summary while the query runs, so you can watch which operator is slow. -E (--vertical, Impala 4.2 and later) prints each row as a block of name-value lines, which helps with wide rows. -p (--show_profiles) prints the profile after every query, which is useful when capturing a session for later analysis.

Scripting impala-shell

For jobs and exports, run the shell non-interactively. -q runs a single statement and -f runs a file of semicolon-separated statements. -B switches from the boxed table to delimited output, --output_delimiter sets the delimiter (tab by default), --print_header adds the column names, and -o writes results to a file. -d sets the starting database. --var defines substitution variables that SQL references as ${var:name}, and -Q (--query_option) sets query options such as the memory limit or the admission pool for this run.

#!/usr/bin/env bash
set -euo pipefail
DAY="$1"

impala-shell -i coord1.example.com:21050 -k --ssl \
  -d sales \
  --var=day="$DAY" \
  -Q mem_limit=8g -Q request_pool=etl \
  -B --output_delimiter=',' --print_header \
  -o "/data/exports/orders_${DAY}.csv" \
  -q "SELECT region, count(*) AS orders, sum(amount) AS revenue
      FROM orders WHERE order_date = '\${var:day}' GROUP BY region"

Two kinds of quoting are at work. The SQL sits inside a double-quoted bash string, so bash would expand ${var:day} itself, as a substring of an unset variable named var, and under set -u the script would stop with an unbound-variable error. The backslash in \${var:day} stops bash and hands impala-shell the literal reference; putting the SQL in a file run with -f avoids the problem entirely. The single quotes around it are SQL quoting: substitution replaces text, so string values need them. Validate outside values before passing them to --var, since substitution is not parameter binding.

By default a script stops at the first failing statement. -c continues past failures, which is convenient for best-effort maintenance scripts and dangerous for anything where later statements assume earlier ones succeeded. Setting a request pool and memory limit per job, as above, lets admission control queue the job rather than letting it starve interactive users.

Shell commands also work in files, so a script can capture its own evidence:

-- nightly_check.sql, run with: impala-shell -i coord1.example.com:21050 -f nightly_check.sql -o nightly_check.out
SET mem_limit=4g;
SELECT count(*) FROM orders WHERE order_date = to_date(now());
SUMMARY;
PROFILE;

The impalarc file

Options you type every time belong in a configuration file. impala-shell reads ~/.impalarc (and a global /etc/impalarc), or the file named by --config_file. The [impala] section takes long option names with underscores; the [impala.query_options] section takes query options. Options on the command line override the file.

# ~/.impalarc
[impala]
impalad=coord1.example.com:21050
kerberos=true
ssl=true
default_db=sales
live_progress=true

[impala.query_options]
mem_limit=8g

Keep secrets out of this file. Use Kerberos, or ldap_password_cmd pointing at a secrets tool, rather than a stored password.

The Web UI, page by page

Each daemon serves its own set of pages. The ones you will use most on an impalad:

PageWhat it showsUse it for
/queriesrunning queries and recently completed ones, with state, duration, rows and a details linkfinding a slow or stuck query and its plan, summary and profile
/sessionsopen client sessionsspotting idle sessions holding resources, or a client leaking sessions
/backendsthe executors this coordinator knows aboutchecking that every node is registered and alive
/admissionadmission control pools: running and queued queries, memory reserved and limitsexplaining why a query is queued or rejected
/memzprocess memory: limits, current consumption and the memory tracker treeseeing which queries or caches hold memory
/metricsevery daemon metricpoint checks; scrape the Prometheus form for trends
/varzstartup flags and their valuesconfirming what a daemon is really running with
/threadzthread groups and threadsdiagnosing hangs and thread exhaustion
/logsthe tail of the daemon logquick checks without shell access
/catalogdatabases and tables this daemon has cachedchecking whether metadata is loaded or stale
/rpczRPC services and their queuesdiagnosing slow or rejected internal RPCs

The statestore's UI adds /subscribers, which lists every daemon subscribed to it with its last heartbeat, and /topics, which shows the topics it distributes, such as cluster membership and catalog updates. A daemon missing from /subscribers is not receiving membership updates. The catalog daemon's UI centres on /catalog, its view of every database and table, and the same /memz, /metrics, /logs and /varz pages. Impala's architecture explains what each daemon does, which makes its pages easier to read.

From /queries, a query's details link opens its plan, query text, execution summary, full profile and memory use, and an in-flight query can be cancelled from the list. The summary there is the same as the shell's SUMMARY; the memory view is the quickest way to see which fragment is holding memory, which pairs well with memory limits.

Scraping: ?json and Prometheus

Append ?json to a page and it returns the same information as JSON, which makes ad-hoc tooling easy. Each daemon also serves /metrics_prometheus in the Prometheus text format, which is the right input for dashboards and alerts.

COORD=http://coord1.example.com:25000

# Every page has a JSON twin: append ?json
curl -s "$COORD/queries?json"   | python3 -m json.tool | head -50
curl -s "$COORD/backends?json"  | python3 -m json.tool | head -30
curl -s "$COORD/admission?json" | python3 -m json.tool | head -40

# Prometheus text format for a metrics pipeline
curl -s "$COORD/metrics_prometheus" | grep -i -E 'admission|mem' | head

# What did this daemon actually start with?
curl -s "$COORD/varz?json" | python3 -m json.tool | grep -E 'query_log_size|webserver'

The JSON layout is the page's internal model and can change between releases, so inspect it on your version before you write a parser, and pin your scripts to the fields you checked. For alerting, prefer the Prometheus metrics, which are designed for collection. Useful first alerts: queued queries per pool, admission rejections, process memory against its limit, and the number of live backends against the expected count.

Retention: where completed queries go

A coordinator keeps completed queries, including their profiles, in an in-memory log. Two impalad startup flags bound it: --query_log_size, a count of queries, and --query_log_size_in_bytes, a memory bound. Defaults have differed between releases and distributions, so read the values from /varz instead of trusting a guide. On a busy coordinator the log turns over quickly, and a restart empties it. The question "what ran at 3 a.m.?" usually cannot be answered from the Web UI the next morning.

Raising the limits keeps more history at the cost of coordinator memory, since profiles of complex queries are large. For lasting history, archive profiles instead: have jobs save their own with PROFILE or -p as above, or use the query archive your platform provides. The habit that pays off most is saving the profile of any query you may need to explain, at the moment it finishes.

Securing the Web UI

The Web UI shows query text, which can contain sensitive values, startup flags, logs and memory details, and it can cancel queries. Treat the web ports as an administrative interface. If you do not use it, turn it off with --enable_webserver=false. If you do, require authentication, for example Kerberos SPNEGO with --webserver_require_spnego=true, serve it over TLS, and limit the ports to operator networks at the firewall. The full set of webserver_ options for TLS and other authentication methods depends on your release; check the daemon's /varz page or its --help output for the exact names before you change them.

Failure modes

  • Wrong port for the protocol: connecting to 21000 with the default hs2 protocol, or to 21050 with Beeswax, fails with a confusing connection error.
  • Looking on the wrong coordinator: behind a load balancer, the query is on a different impalad's /queries page.
  • Broken substitution: bash expands an unescaped ${var:day} inside double quotes before impala-shell sees it, and a reference without SQL quotes, or fed untrusted input, produces broken or injected SQL.
  • -c hiding failures: a script continues after a failed step and loads wrong data.
  • Profiles lost: the query fell out of the in-memory log or the coordinator restarted before anyone saved the profile.
  • An open Web UI: anyone on the network can read query text and flags or cancel queries.
  • Parsing JSON across upgrades: a scraper breaks when a page's layout changes.

Trade-offs

The shell gives exact, scriptable access to one session; the Web UI gives a broad live view but keeps little history. Larger query logs keep more profiles at the cost of coordinator memory; archiving costs storage and pipeline work but survives restarts. Turning the Web UI off removes a risk and a diagnostic tool at once, so most teams keep it on behind authentication and a firewall.

What to do next

  1. Update scripts and docs that connect to port 21000 to use 21050 with the default protocol, or set the protocol explicitly.
  2. Create an impalarc with your coordinator, authentication and default database, and keep secrets out of it.
  3. Make every scheduled impala-shell job set a request pool and memory limit, quote its variables and fail on error.
  4. Read query_log_size and query_log_size_in_bytes from /varz on each coordinator, and decide how you will archive profiles.
  5. Scrape /metrics_prometheus and alert on queued queries, rejections, process memory and live backends.
  6. Put the Web UI behind authentication and TLS, restrict its ports, or disable it where nobody uses it.
Key takeaway: impala-shell defaults to the HS2 protocol on port 21050, not the old Beeswax port 21000, and supports Kerberos, LDAP and TLS. Script it with -q or -f, -B, -o, --var and -Q, and keep defaults in an impalarc. Each daemon's Web UI on ports 25000, 25010 and 25020 shows queries, sessions, backends, admission, memory, metrics and flags, with a JSON form of every page and a Prometheus endpoint. Completed queries live in a bounded in-memory log, so save profiles you need, and secure or disable the UI.