Most of what you will ever do to an Impala cluster by hand goes through two tools. impala-shell is the command-line client: it opens a session on a coordinator, runs SQL, and prints the summary and profile of what it ran. The Web UI is a small HTTP server built into each Impala daemon: it lists running and recent queries, sessions, backends, memory, metrics, logs and startup flags, and every page can also return JSON.
Knowing the tools well pays off in incidents and in automation. This page covers how to connect correctly (the default port changed, and many old guides still show the old one), how to script impala-shell safely, what each Web UI page is for, how to scrape it, why the query you are looking for has vanished from the UI, and how to keep the UI from becoming a security hole. Reading plans and profiles is covered by Impala query plans and Impala troubleshooting; here the focus is the tooling around them.
How the pieces fit
Every impalad can act as a coordinator, and each one runs its own web server. The statestore and catalog daemons run web servers too. impala-shell talks to a coordinator over HiveServer2; a browser or curl talks to any daemon's web port. The two views complement each other: the shell shows one query from the client's side, while the coordinator's Web UI shows every query that coordinator is running or recently ran. A query is visible only on the coordinator that ran it, so when a load balancer spreads sessions over several coordinators, you need to know which one served the query before you look.
Connecting: protocols, ports and authentication
impala-shell's --protocol option defaults to hs2, HiveServer2 over binary TCP, which a coordinator serves on port 21050. Older releases defaulted to the Beeswax protocol on port 21000, and many tutorials still show -i host:21000; if you pass that port with the default protocol, the connection fails. The third option, hs2-http, carries HS2 over HTTP on port 28000, which suits proxies and gateways that only pass HTTP.
| Port | Flag | Used by |
|---|---|---|
| 21050 | --hs2_port | impala-shell (default hs2 protocol), JDBC and ODBC |
| 28000 | --hs2_http_port | impala-shell with --protocol=hs2-http, HTTP clients |
| 21000 | --beeswax_port | legacy Beeswax clients |
| 25000 / 25010 / 25020 | --webserver_port | Web UI of impalad / statestored / catalogd |
# HS2 over binary TCP (the default protocol) to a coordinator
impala-shell -i coord1.example.com:21050
# HS2 over HTTP, e.g. through a proxy that only passes HTTP, with TLS
impala-shell -i impala-lb.example.com:28000 --protocol=hs2-http --ssl --ca_cert=/etc/pki/ca.pem
# Kerberos (run kinit first)
impala-shell -k -i coord1.example.com:21050 --ssl
# LDAP: -l enables it, -u names the user; the password is prompted for
impala-shell -l -u analyst1 -i coord1.example.com:21050 --sslPass -i the host and port of a coordinator or of a load balancer in front of several. HS2 sessions are stateful, so the load balancer must keep a session on one coordinator for its lifetime. -k uses the Kerberos ticket from kinit. -l turns on LDAP and -u names the user; for unattended use, --ldap_password_cmd runs a command that prints the password, so it never appears on the command line. With TLS, --ssl encrypts the connection and --ca_cert points at the CA certificate, or at the server certificate for a self-signed setup. Sending LDAP passwords without TLS sends them in clear text.
Working interactively
In an interactive session, statements end with a semicolon. A few shell commands do most of the work:
SETwith no arguments lists query options and their current values;SET mem_limit=4g;changes one for the rest of the session.SUMMARYprints the execution summary of the last query: per operator, the number of hosts, average and maximum time, rows produced and estimated, and peak memory.PROFILEprints the full runtime profile of the last query: the plan, the timeline and every counter for every fragment instance.EXPLAINbefore a statement shows the plan without running it.- Ctrl-C cancels the running query rather than closing the shell.
Two startup options make long queries less opaque: --live_progress shows a progress bar as scan ranges complete, and --live_summary redraws the execution summary while the query runs, so you can watch which operator is slow. -E (--vertical, Impala 4.2 and later) prints each row as a block of name-value lines, which helps with wide rows. -p (--show_profiles) prints the profile after every query, which is useful when capturing a session for later analysis.
Scripting impala-shell
For jobs and exports, run the shell non-interactively. -q runs a single statement and -f runs a file of semicolon-separated statements. -B switches from the boxed table to delimited output, --output_delimiter sets the delimiter (tab by default), --print_header adds the column names, and -o writes results to a file. -d sets the starting database. --var defines substitution variables that SQL references as ${var:name}, and -Q (--query_option) sets query options such as the memory limit or the admission pool for this run.
#!/usr/bin/env bash
set -euo pipefail
DAY="$1"
impala-shell -i coord1.example.com:21050 -k --ssl \
-d sales \
--var=day="$DAY" \
-Q mem_limit=8g -Q request_pool=etl \
-B --output_delimiter=',' --print_header \
-o "/data/exports/orders_${DAY}.csv" \
-q "SELECT region, count(*) AS orders, sum(amount) AS revenue
FROM orders WHERE order_date = '\${var:day}' GROUP BY region"Two kinds of quoting are at work. The SQL sits inside a double-quoted bash string, so bash would expand ${var:day} itself, as a substring of an unset variable named var, and under set -u the script would stop with an unbound-variable error. The backslash in \${var:day} stops bash and hands impala-shell the literal reference; putting the SQL in a file run with -f avoids the problem entirely. The single quotes around it are SQL quoting: substitution replaces text, so string values need them. Validate outside values before passing them to --var, since substitution is not parameter binding.
By default a script stops at the first failing statement. -c continues past failures, which is convenient for best-effort maintenance scripts and dangerous for anything where later statements assume earlier ones succeeded. Setting a request pool and memory limit per job, as above, lets admission control queue the job rather than letting it starve interactive users.
Shell commands also work in files, so a script can capture its own evidence:
-- nightly_check.sql, run with: impala-shell -i coord1.example.com:21050 -f nightly_check.sql -o nightly_check.out
SET mem_limit=4g;
SELECT count(*) FROM orders WHERE order_date = to_date(now());
SUMMARY;
PROFILE;
The impalarc file
Options you type every time belong in a configuration file. impala-shell reads ~/.impalarc (and a global /etc/impalarc), or the file named by --config_file. The [impala] section takes long option names with underscores; the [impala.query_options] section takes query options. Options on the command line override the file.
# ~/.impalarc
[impala]
impalad=coord1.example.com:21050
kerberos=true
ssl=true
default_db=sales
live_progress=true
[impala.query_options]
mem_limit=8gKeep secrets out of this file. Use Kerberos, or ldap_password_cmd pointing at a secrets tool, rather than a stored password.
The Web UI, page by page
Each daemon serves its own set of pages. The ones you will use most on an impalad:
| Page | What it shows | Use it for |
|---|---|---|
/queries | running queries and recently completed ones, with state, duration, rows and a details link | finding a slow or stuck query and its plan, summary and profile |
/sessions | open client sessions | spotting idle sessions holding resources, or a client leaking sessions |
/backends | the executors this coordinator knows about | checking that every node is registered and alive |
/admission | admission control pools: running and queued queries, memory reserved and limits | explaining why a query is queued or rejected |
/memz | process memory: limits, current consumption and the memory tracker tree | seeing which queries or caches hold memory |
/metrics | every daemon metric | point checks; scrape the Prometheus form for trends |
/varz | startup flags and their values | confirming what a daemon is really running with |
/threadz | thread groups and threads | diagnosing hangs and thread exhaustion |
/logs | the tail of the daemon log | quick checks without shell access |
/catalog | databases and tables this daemon has cached | checking whether metadata is loaded or stale |
/rpcz | RPC services and their queues | diagnosing slow or rejected internal RPCs |
The statestore's UI adds /subscribers, which lists every daemon subscribed to it with its last heartbeat, and /topics, which shows the topics it distributes, such as cluster membership and catalog updates. A daemon missing from /subscribers is not receiving membership updates. The catalog daemon's UI centres on /catalog, its view of every database and table, and the same /memz, /metrics, /logs and /varz pages. Impala's architecture explains what each daemon does, which makes its pages easier to read.
From /queries, a query's details link opens its plan, query text, execution summary, full profile and memory use, and an in-flight query can be cancelled from the list. The summary there is the same as the shell's SUMMARY; the memory view is the quickest way to see which fragment is holding memory, which pairs well with memory limits.
Scraping: ?json and Prometheus
Append ?json to a page and it returns the same information as JSON, which makes ad-hoc tooling easy. Each daemon also serves /metrics_prometheus in the Prometheus text format, which is the right input for dashboards and alerts.
COORD=http://coord1.example.com:25000
# Every page has a JSON twin: append ?json
curl -s "$COORD/queries?json" | python3 -m json.tool | head -50
curl -s "$COORD/backends?json" | python3 -m json.tool | head -30
curl -s "$COORD/admission?json" | python3 -m json.tool | head -40
# Prometheus text format for a metrics pipeline
curl -s "$COORD/metrics_prometheus" | grep -i -E 'admission|mem' | head
# What did this daemon actually start with?
curl -s "$COORD/varz?json" | python3 -m json.tool | grep -E 'query_log_size|webserver'The JSON layout is the page's internal model and can change between releases, so inspect it on your version before you write a parser, and pin your scripts to the fields you checked. For alerting, prefer the Prometheus metrics, which are designed for collection. Useful first alerts: queued queries per pool, admission rejections, process memory against its limit, and the number of live backends against the expected count.
Retention: where completed queries go
A coordinator keeps completed queries, including their profiles, in an in-memory log. Two impalad startup flags bound it: --query_log_size, a count of queries, and --query_log_size_in_bytes, a memory bound. Defaults have differed between releases and distributions, so read the values from /varz instead of trusting a guide. On a busy coordinator the log turns over quickly, and a restart empties it. The question "what ran at 3 a.m.?" usually cannot be answered from the Web UI the next morning.
Raising the limits keeps more history at the cost of coordinator memory, since profiles of complex queries are large. For lasting history, archive profiles instead: have jobs save their own with PROFILE or -p as above, or use the query archive your platform provides. The habit that pays off most is saving the profile of any query you may need to explain, at the moment it finishes.
Securing the Web UI
The Web UI shows query text, which can contain sensitive values, startup flags, logs and memory details, and it can cancel queries. Treat the web ports as an administrative interface. If you do not use it, turn it off with --enable_webserver=false. If you do, require authentication, for example Kerberos SPNEGO with --webserver_require_spnego=true, serve it over TLS, and limit the ports to operator networks at the firewall. The full set of webserver_ options for TLS and other authentication methods depends on your release; check the daemon's /varz page or its --help output for the exact names before you change them.
Failure modes
- Wrong port for the protocol: connecting to 21000 with the default hs2 protocol, or to 21050 with Beeswax, fails with a confusing connection error.
- Looking on the wrong coordinator: behind a load balancer, the query is on a different impalad's
/queriespage. - Broken substitution: bash expands an unescaped
${var:day}inside double quotes before impala-shell sees it, and a reference without SQL quotes, or fed untrusted input, produces broken or injected SQL. -chiding failures: a script continues after a failed step and loads wrong data.- Profiles lost: the query fell out of the in-memory log or the coordinator restarted before anyone saved the profile.
- An open Web UI: anyone on the network can read query text and flags or cancel queries.
- Parsing JSON across upgrades: a scraper breaks when a page's layout changes.
Trade-offs
The shell gives exact, scriptable access to one session; the Web UI gives a broad live view but keeps little history. Larger query logs keep more profiles at the cost of coordinator memory; archiving costs storage and pipeline work but survives restarts. Turning the Web UI off removes a risk and a diagnostic tool at once, so most teams keep it on behind authentication and a firewall.
What to do next
- Update scripts and docs that connect to port 21000 to use 21050 with the default protocol, or set the protocol explicitly.
- Create an impalarc with your coordinator, authentication and default database, and keep secrets out of it.
- Make every scheduled impala-shell job set a request pool and memory limit, quote its variables and fail on error.
- Read query_log_size and query_log_size_in_bytes from /varz on each coordinator, and decide how you will archive profiles.
- Scrape /metrics_prometheus and alert on queued queries, rejections, process memory and live backends.
- Put the Web UI behind authentication and TLS, restrict its ports, or disable it where nobody uses it.