Hive + Impala (SQL on Hadoop)

Hive + Impala (SQL on Hadoop)

Deep technical articles on this topic.

150Articles
150Topics covered
Articles in this category

All 108 articles, sorted alphabetically

ARTICLE · 001

Cloudera Data Warehouse (CDW), in depth: Database Catalogs, Virtual Warehouses, sizing, auto-scaling and auto-suspend in production

How Cloudera Data Warehouse runs Hive, Impala and Trino as containerised Virtual Warehouses over a shared Database Catalog and object storage: the thr…

Read article →
ARTICLE · 002

Hive 3 ACID Improvements, in depth: split-update deltas, per-table write ids, unbucketed and insert-only tables, managed-by-default and the upgrade path

What changed in Hive 3 transactional tables compared with Hive 1 and 2 ACID and why it matters: the ACID v2 split-update layout with delete deltas, pe…

Read article →
ARTICLE · 003

Hive 4 Key Features, in depth: the 4.0 to 4.2 release line, what was removed, Iceberg, compaction pools and rebalance, compiler changes, security, and what to adopt first

A practical map of Apache Hive 4: release dates and Java requirements from 4.0.0 to 4.2.1, removal of Hive on Spark and deprecation of MapReduce and t…

Read article →
ARTICLE · 004

Hive ACID architecture

Deep-dive on Hive ACID transactions: write IDs and snapshot isolation, delta and delete-delta layout with hidden row IDs, MERGE mechanics, lock manage…

Read article →
ARTICLE · 005

Hive Authorization, in depth: where checks run, storage-based and SQL-standard models, Ranger policies, doAs and the paths that bypass HiveServer2

How authorization works in Apache Hive: authentication as a prerequisite, where HiveServer2 checks privileges, the storage-based, SQL-standard and Ran…

Read article →
ARTICLE · 006

Beeline vs Hive CLI, in depth: thick client versus JDBC client, why the CLI bypasses security, JDBC URLs that work, scripting flags, and migrating old scripts

How the legacy Hive CLI and Beeline differ in architecture and behaviour: in-process compilation versus HiveServer2 over JDBC, why the CLI bypasses Hi…

Read article →
ARTICLE · 007

Hive Bucketing: CLUSTERED BY, Bucket Map Join and SMB Join

How Hive bucketing works: CLUSTERED BY hashes rows into fixed bucket files, enabling bucket map joins and sort-merge-bucket (SMB) joins without shuffl…

Read article →
ARTICLE · 008

Hive Bucketing for Efficient Joins, in depth: bucket map joins, sort-merge-bucket joins, bucketing versions and how to prove the plan

How Hive turns a bucketed table layout into a cheaper join: the bucket-id formula, bucket map join and sort-merge-bucket join on Tez, the settings and…

Read article →
ARTICLE · 009

Hive Table Buckets, in depth: designing, writing, sizing and migrating bucketed tables

A table-design guide to Hive buckets: how buckets sit on disk, choosing the key and bucket count against data volume, the write path, Hive 3 ACID tabl…

Read article →
ARTICLE · 010

Hive CBO architecture

Deep-dive on Hive's cost-based optimizer: Calcite RelNode pipeline, predicate pushdown and partition pruning, NDV-based cardinality estimation, j…

Read article →
ARTICLE · 011

Hive ACID compaction architecture

Deep-dive on Hive ACID compaction: because updates and deletes are written as delta directories on immutable storage, reads amplify and space leaks un…

Read article →
ARTICLE · 012

Hive Complex Types, in depth: ARRAY, MAP, STRUCT, LATERAL VIEW and modelling nested data

How Hive's complex types work in practice: ARRAY, MAP, STRUCT and UNIONTYPE DDL, constructors and access syntax, flattening with LATERAL VIEW and…

Read article →
ARTICLE · 013

Hive Compression Options, in depth: storage, shuffle and output codecs, splittability and safe codec changes

Where Hive and Impala compress data and how to choose: storage codecs in ORC and Parquet, output compression for text and SequenceFile, intermediate s…

Read article →
ARTICLE · 014

Hive Deprecated Features: MapReduce Engine, Indexes, Old SerDes

Which Hive features are deprecated and what to use instead: the MapReduce engine (move to Tez), Hive indexes, legacy SerDes and codecs, and migration.

Read article →
ARTICLE · 015

Hive dynamic partitioning

Deep-dive on Hive dynamic partitioning: static vs dynamic keys and trailing-column binding, DISTRIBUTE BY for file-count control, max-partition guardr…

Read article →
ARTICLE · 016

Hive Future, in depth: what Hive 4 changed, why the metastore outlives the engine, and how to pick a direction for a Hive estate in 2026

A decision guide to Apache Hive's future: the 4.x release line and end-of-life dates, Hive as engine, metastore and table format with diverging f…

Read article →
ARTICLE · 017

HiveServer2 architecture

Deep-dive on HiveServer2: Thrift transports and session/operation managers, async execution over Tez session pools and LLAP, result fetch and spooling…

Read article →
ARTICLE · 018

Apache Iceberg -- the open table format for the lakehouse

Deep-dive on Apache Iceberg: layered metadata (catalog/manifest/data files), immutable snapshots, ACID via atomic pointer swaps, time travel, hidden p…

Read article →
ARTICLE · 019

Hive + Iceberg Tables, in depth: the storage handler, catalogs, the commit path, row-level DML, time travel and maintenance from HiveQL

How Apache Hive 4 runs Iceberg tables: the storage handler and what the metastore still stores, creating tables with partition transforms, HiveCatalog…

Read article →
ARTICLE · 020

Hive Join Strategy Selection, in depth: how the Tez compiler chooses between map join, bucket map join, dynamically partitioned hash join, SMB and merge join

The physical join selection procedure in Hive on Tez, read from ConvertJoinMapJoin: the order of attempts, how the big table is chosen, the bucket and…

Read article →
ARTICLE · 021

Hive Joins, in depth: join semantics, shuffle versus MapJoin, reading the plan and the bugs that return wrong rows

How Hive joins work, from the rows they return to how they run: ON versus WHERE in outer joins, NULL keys and null-safe equality, semi and anti joins,…

Read article →
ARTICLE · 022

Hive with JSON and Avro data, in depth: SerDes, malformed records, Avro schema evolution and the path to columnar tables

How Hive reads JSON and Avro: the InputFormat and SerDe pipeline, JsonSerDe tables and the raw-string pattern for dirty data, STORED AS AVRO with avro…

Read article →
ARTICLE · 023

Hive JVM Tuning, in depth: Tez containers vs heap, sort and map-join memory, the Tez AM, HiveServer2 and Metastore heaps, GC flags by JDK and reading OOM kills

Tuning the JVMs of a Hive-on-Tez deployment: container size versus -Xmx and off-heap memory, sizing sort buffers and map-join thresholds, a worked 128…

Read article →
ARTICLE · 024

Hive LLAP architecture

Deep-dive on Hive LLAP: HiveServer2 planning, Tez AM coordination, daemon executor slots, async IO elevator, off-heap columnar cache with SSD tier, Zo…

Read article →
ARTICLE · 025

Hive LLAP Performance Deep, in depth: sizing the daemon, executors and IO threads, making the cache hit, reading LLAP IO counters and benchmarking honestly

A practical guide to Hive LLAP performance: where interactive query time goes, a worked container memory budget (heap per executor, off-heap cache, he…

Read article →
ARTICLE · 026

Hive materialized views architecture

Deep-dive on Hive materialized views: precomputing join/aggregate results, cost-based-optimizer automatic query rewrite, the freshness gate, increment…

Read article →
ARTICLE · 027

Hive Metastore architecture

Deep-dive on the Hive Metastore: Thrift API over MySQL/Postgres, partition pruning paths, column statistics for CBOs, ACID transaction state, event no…

Read article →
ARTICLE · 028

Hive Metastore Backup, in depth: consistent dumps, point-in-time recovery, and keeping ACID write IDs and Iceberg pointers consistent with the data

How to back up and restore the Hive Metastore database properly: what it stores and what it does not, consistent logical dumps and point-in-time recov…

Read article →
ARTICLE · 029

Hive Metastore HA, in depth: client failover, service discovery, housekeeping leader election and a highly available backing database

How to make the Hive Metastore highly available: why HMS servers are stateless and the database is the real single point of failure, client-side URI s…

Read article →
ARTICLE · 030

Hive Monitoring

Hive monitoring at scale: JMX metrics via Prometheus, query logging to HDFS/Elasticsearch, alerting on HS2 saturation and metastore bottlenecks, track…

Read article →
ARTICLE · 031

Hive on Spark, in depth: a tuning guide for clusters still running it, from executor sizing and session budgets to map joins, skew and small files

Tuning Hive on Spark on Hive 2.x and 3.x clusters: carving YARN hosts into executors with a worked example, driver sizing, dynamic allocation and per-…

Read article →
ARTICLE · 032

Hive on Tez Deep Dive, in depth: from the EXPLAIN plan to the settings that control splits, reducers, shuffle and memory

A practical deep dive into Hive on Tez: reading EXPLAIN edge types, how split grouping sets mapper counts, reducer estimation and runtime auto paralle…

Read article →
ARTICLE · 033

Hive Cost-Based Optimizer

How Hive's Calcite-based cost optimizer picks join orders and plan alternatives: hive.cbo.enable and the stats keys, the join-order search space,…

Read article →
ARTICLE · 034

ORC format architecture

Deep-dive on ORC internals: file/stripe/stream layout, RLEv2 and dictionary encodings, row-group indexes with seek positions, min/max stats and bloom …

Read article →
ARTICLE · 035

What Is Apache Hive? Architecture, Metastore and HiveQL on Hadoop

What Apache Hive is and how it works: HiveQL over HDFS, the metastore, HiveServer2, Tez and LLAP execution, and how it compares with Impala and Spark.

Read article →
ARTICLE · 036

Parquet Format

The internal structure of Parquet: row groups, column chunks, page indexes, dictionary encoding, and cross-engine compatibility.

Read article →
ARTICLE · 037

Hive Partition Evolution, in depth: changing partitioning in classic Hive tables and in Iceberg tables without rewriting history

How to change a table's partitioning over time in Hive: what classic Hive lets you evolve per partition (columns with CASCADE or RESTRICT, file f…

Read article →
ARTICLE · 038

Hive Table Partitions, in depth: layout and pruning, choosing keys with real arithmetic, static and dynamic inserts, repair and retention

How Hive partitions work and how to design them: directory layout and metastore rows, what partition pruning needs to happen, choosing partition keys …

Read article →
ARTICLE · 039

Hive predicate pushdown architecture

Deep-dive on Hive predicate pushdown: how the compiler splits pushable from residual predicates, how SearchArguments reach ORC/Parquet readers, partit…

Read article →
ARTICLE · 040

Hive replication architecture

Deep-dive on Hive replication: event-driven incremental replication over the notification log, bootstrap and checkpointed cycles, DistCp data movement…

Read article →
ARTICLE · 041

Hive Security, in depth: authentication, wire protection, the metastore side door, secrets, hardening and audit

A layer-by-layer guide to securing Apache Hive: a threat model, HiveServer2 authentication modes and transports, SASL and TLS wire protection, locking…

Read article →
ARTICLE · 042

Hive Security with Kerberos and Ranger, in depth: following one identity from kinit to a policy decision, configuring each hop and debugging it when it breaks

How Kerberos authentication and Ranger authorization fit together in Hive: the identity chain for one query, service principals and keytabs, client JD…

Read article →
ARTICLE · 043

Hive SerDes, in depth: how Hive turns bytes into rows, the built-in SerDes, their properties and traps, and writing your own

How Hive SerDes work: the read and write paths, AbstractSerDe and ObjectInspectors, LazySimpleSerDe delimiters and NULLs, OpenCSVSerde, RegexSerDe, JS…

Read article →
ARTICLE · 044

Hive Skew Join Optimization: hive.optimize.skewjoin and MapJoin

How Hive handles data skew in joins: runtime vs compile-time skew join, the hive.skewjoin.key threshold, hot keys sent to a follow-up MapJoin, salting…

Read article →
ARTICLE · 045

Hive small-file problem

Deep-dive on the small-file problem: sources (streaming, dynamic partitioning, appends), the three-layer cost (metadata, task overhead, read amplifica…

Read article →
ARTICLE · 046

Hive on Spark Engine, in depth: how HiveQL became Spark jobs, the per-session remote driver, tuning, failure modes and migrating off it

How Hive on Spark worked in Hive 2.x and 3.x: SparkCompiler and SparkWork, the remote Spark driver per HiveServer2 session, how map and reduce work be…

Read article →
ARTICLE · 047

Hive Storage Handlers, in depth: HBase, Kafka and JDBC tables, the handler contract and when to copy instead

Hive storage handlers explained: the HiveStorageHandler and HiveMetaHook contract, native versus non-native tables, HBase column mapping and upsert se…

Read article →
ARTICLE · 048

Hive Streaming Ingestion, in depth: the V2 API, commit cadence and compaction

How Hive Streaming Data Ingest V2 works: transactional ORC tables, write IDs and delta directories, a complete Java HiveStreamingConnection client, co…

Read article →
ARTICLE · 049

Hive Subqueries + CTEs, in depth: where subqueries are allowed, how Hive rewrites them into joins, the NOT IN trap, and CTE inlining versus materialization

A working guide to subqueries and common table expressions in Apache Hive: FROM, WHERE, HAVING and SELECT subqueries and their version history, how ea…

Read article →
ARTICLE · 050

Hive Tables, in depth: managed vs external, transactional tables, CREATE TABLE clauses, partitions, schema evolution and DROP semantics

What a Hive table really is: a Metastore record plus a directory of files. Table kinds (managed, external, full ACID, insert-only, temporary), the CRE…

Read article →
ARTICLE · 051

Hive on Tez architecture

Deep-dive on Tez execution for Hive: vertices and typed edges vs MapReduce chains, the per-session application master, YARN container reuse, runtime a…

Read article →
ARTICLE · 052

Hive Tez Optimization, in depth: reading counters, split grouping, memory, reducers and pruning

A measurement-first tuning loop for Hive on Tez: reading plans and counters, cutting session startup, the split-grouping arithmetic, container, heap a…

Read article →
ARTICLE · 053

Tez UI + Application Timeline Server, in depth: the history pipeline, ATS v1 vs v1.5, the entity model, REST queries and diagnosing slow Hive DAGs

How the Tez UI gets its data from the YARN Application Timeline Server: history logging services, LevelDB versus EntityGroupFSTimelineStore, the Tez e…

Read article →
ARTICLE · 054

Hive ACID Transactions Deep Dive: Transaction IDs and Write IDs, Snapshot Visibility, Locks, Write-Set Conflicts and the Stuck Transaction That Stalls a Cluster

The Hive ACID transaction protocol in depth: how DbTxnManager opens transactions and allocates per-table write ids, how snapshot visibility uses a hig…

Read article →
ARTICLE · 055

Hive TRANSFORM with Python, in depth: the streaming protocol, reduce-side scripts, testing, security and when not to use it

How Hive's TRANSFORM clause streams rows through an external Python script: the tab-separated wire format and NULL marker, typed output with AS, …

Read article →
ARTICLE · 056

Hive Query Tuning Playbook, in depth: reading plans, statistics, partition pruning, file layout, join choice and skew

A symptom-to-fix playbook for slow Hive queries at the SQL and table-layout level: EXPLAIN variants and estimate-versus-actual checks, statistics and …

Read article →
ARTICLE · 057

Hive UDFs

How to write Hive UDFs, UDAFs, and UDTFs, when to use them, and the performance and safety implications.

Read article →
ARTICLE · 058

Hive Upgrade Path, in depth: planning 2.x to 3.x to 4.x, metastore schema, ACID, clients and rollback

How to upgrade Apache Hive safely: what changes between 2.x, 3.x and 4.x, in-place versus side-by-side strategies, metastore schema upgrades with sche…

Read article →
ARTICLE · 059

Hive vectorized execution

Deep-dive on Hive's vectorized engine: VectorizedRowBatch and ColumnVector anatomy, expression templates and monomorphic loops, selection-vector …

Read article →
ARTICLE · 060

Hive View Types

Hive views in practice: how virtual views expand into the plan, materialized view rewriting and rebuild, union views for hot-plus-cold tables, and Imp…

Read article →
ARTICLE · 061

Hive Window Functions: OVER, PARTITION BY, ROW_NUMBER, RANK

Hive window functions explained: the OVER clause with PARTITION BY and ORDER BY, ROWS vs RANGE frames, ROW_NUMBER, RANK, LAG and LEAD, and window vs s…

Read article →
ARTICLE · 062

HiveServer2 Clustering, in depth: ZooKeeper discovery, load balancers, active/passive, sizing and rolling restarts

How to run several HiveServer2 instances as a resilient pool: what state stays server-local, ZooKeeper dynamic service discovery and JDBC URLs, load b…

Read article →
ARTICLE · 063

Impala 4 Key Features: What 4.0 to 4.5 Changed and How to Adopt It

Apache Impala 4.x key features from 4.0 to 4.5, verified against the changelogs: MT_DOP for all operators, Iceberg DELETE, UPDATE and MERGE, ACID ORC,…

Read article →
ARTICLE · 064

Impala Administration, in depth: daemon roles, flags, HA, rolling restarts, metadata hygiene and security wiring

A day-2 operations guide for Apache Impala: what each daemon does and how to configure it, the ports and web UI endpoints, statestore and catalog high…

Read article →
ARTICLE · 065

Impala admission control architecture

Deep-dive on Impala admission control: resource pools, per-host memory estimates vs MEM_LIMIT, statestore-gossiped local decisions, the admissiond cen…

Read article →
ARTICLE · 066

Impala Analytic Functions, in depth: window semantics, Impala's restrictions, execution and performance

Apache Impala analytic (window) functions from first principles: OVER, PARTITION BY, ORDER BY and frames, the full function list, Impala's docume…

Read article →
ARTICLE · 067

Impala Architecture

The three daemon types in Impala, how they cooperate, and what each is responsible for.

Read article →
ARTICLE · 068

Impala Catalog and Statestore: catalogd, REFRESH vs INVALIDATE

How Impala catalogd and statestored keep impalad metadata in sync with the Hive Metastore: REFRESH vs INVALIDATE METADATA, SYNC_DDL and local catalog …

Read article →
ARTICLE · 069

Impala Cluster Sizing, in depth: executors from measured memory, scan rate and concurrency, cache and scratch disks, node shape and validation

How to size Impala executors bottom-up: which resource binds first, collecting per-host peak memory, bytes scanned and concurrency from profiles, deri…

Read article →
ARTICLE · 070

Impala code generation architecture

Deep-dive on Impala runtime code generation: why interpretation is slow at scale, cross-compiling the C++ runtime to LLVM bitcode, building per-query …

Read article →
ARTICLE · 071

Impala Complex Types, in depth: querying ARRAY, MAP and STRUCT, subplans, UNNEST and schema traps

How Impala queries nested Parquet and ORC data: supported formats and hard limits, collection joins with ITEM, POS, KEY and VALUE, outer joins and cor…

Read article →
ARTICLE · 072

Impala Configuration and Admin, in depth: the four configuration layers, resource pool files, query option clamps and drift-free changes

How Impala configuration fits together: startup flags and flag files, fair-scheduler.xml and llama-site.xml resource pools, pool default query options…

Read article →
ARTICLE · 073

Impala Coordinator + Executor Roles, in depth: what each role does for a query, dedicated coordinators, sizing, failure semantics and transparent retries

How Impala splits work between coordinator and executor roles: the is_coordinator and is_executor flags, a query traced through parse, plan, admission…

Read article →
ARTICLE · 074

Impala Cost Management, in depth: metering every query from the query log, a reservation-based cost model, budgets, alerts and an enforcement ladder

How to run cost management for Apache Impala as a loop: enable workload management and the sys.impala_query_log Iceberg table, charge queries by their…

Read article →
ARTICLE · 075

Impala Cost Optimization, in depth: reading less, reserving honestly and paying only for busy nodes

A practical guide to cutting the cost of Apache Impala: the cost equation, measuring with EXPLAIN, SUMMARY and PROFILE, partition pruning, Parquet lay…

Read article →
ARTICLE · 076

Impala data cache architecture

Deep-dive on the Impala data cache: per-daemon local NVMe caching of remote byte ranges keyed by file and offset, the hit/miss/insert read-through pat…

Read article →
ARTICLE · 077

Impala DML Support, in depth: INSERT and OVERWRITE on file tables, the staging write path, partition hints, and what UPDATE and DELETE need

What Impala DML supports on each storage type, how an INSERT into HDFS or S3 Parquet tables is written through _impala_insert_staging, static and dyna…

Read article →
ARTICLE · 078

Impala Future, in depth: where Apache Impala stands in 2026, what still keeps it fast, and how to decide between staying, coexisting with Trino and migrating

An evidence-based look at Apache Impala's future: the 4.5 release line and its Iceberg MERGE support, why resident C++ daemons still win high-con…

Read article →
ARTICLE · 079

Impala on HDFS + S3, in depth: locality versus remote reads, split sizing, rename-free writes, metadata cost and tables that span both

How Apache Impala reads and writes tables on HDFS and on Amazon S3: scan scheduling with and without locality, short-circuit reads, S3A configuration,…

Read article →
ARTICLE · 080

Impala + Iceberg Tables, in depth: how Impala plans, writes, deletes and maintains Iceberg tables, and how to run them alongside Spark

Running Apache Iceberg tables on Impala: the metadata tree and how the planner prunes with it, CREATE TABLE with partition transforms and catalogs, Pa…

Read article →
ARTICLE · 081

Impala In-Memory Execution, in depth: row batches, tuples, the pull model, streaming and blocking operators, and where the memory goes

What in-memory execution actually means in Impala: row batches and tuple layout, the Open/GetNext/Close pull model, streaming versus blocking operator…

Read article →
ARTICLE · 082

Impala Join Strategies, in depth: broadcast versus partitioned hash joins, join order, missing statistics, hints and the query options that steer them

How Apache Impala executes joins: hash join and nested loop join, broadcast versus partitioned distribution and the network and memory each costs, how…

Read article →
ARTICLE · 083

Impala and Kudu: UPSERT, UPDATE, DELETE on Kudu Tables

Using Apache Kudu with Impala: STORED AS KUDU tables, primary keys and range partitions, UPSERT, UPDATE and DELETE, pushdown in EXPLAIN, and Kudu vs I…

Read article →
ARTICLE · 084

Impala MEM_LIMIT: Memory Limits, Spilling, Admission Control

Tuning Impala memory: SET MEM_LIMIT per query, per-node daemon limits, EXPLAIN estimates, spill to disk vs OOM, buffer pool reservations, admission co…

Read article →
ARTICLE · 085

Impala Metadata Deep Dive: What the Catalog Caches, How It Goes Stale and How to Keep It Fresh

How Impala's metadata really works: what catalogd caches from the Hive Metastore and storage, how coordinators receive it, why data written by Hi…

Read article →
ARTICLE · 086

Impala Monitoring, in depth: the Prometheus endpoint and its naming quirks, signals per layer, alert rules and query-log SLOs

How to monitor Apache Impala in production: scraping /metrics_prometheus on impalad, statestored and catalogd, how Impala renames metrics and embeds p…

Read article →
ARTICLE · 087

Impala MPP Architecture, in depth: fragments and instances, scan-range assignment, exchanges, intra-node parallelism, skew and the serial tail

How Impala divides work across a cluster: shared-nothing execution over shared storage, plan fragments and instances, scan-range scheduling and locali…

Read article →
ARTICLE · 088

Impala Node Types, in depth: impalad roles, statestored, catalogd and admissiond, and what breaks when each one dies

A catalogue of Impala's node types: impalad coordinator and executor roles selected by is_coordinator and is_executor, executor groups, statestor…

Read article →
ARTICLE · 089

Impala ORC Support, in depth: what Impala can and cannot do with ORC, how the scanner reads it, schema resolution, full ACID tables and when to convert to Parquet

A practical guide to querying ORC tables from Apache Impala: the support matrix (reads yes, inserts no), how the ORC scanner and metadata path work, a…

Read article →
ARTICLE · 090

Impala Parquet Native Reader, in depth: the C++ scanner pipeline, five levels of skipping, late materialization, schema resolution and writing files it can skip

How Impala reads Parquet with its own C++ scanner: scan ranges and footers, row-group skipping by statistics, dictionaries and bloom filters, page-ind…

Read article →
ARTICLE · 091

Impala Partition Pruning, in depth: which predicates prune, reading EXPLAIN, dynamic pruning, stale metadata and key design

How Apache Impala prunes partitions: planner evaluation of partition predicates, predicate propagation, what does not prune (OR with data columns, dat…

Read article →
ARTICLE · 092

Impala query execution architecture

Deep-dive on Impala's MPP execution: coordinator planning, statestore and catalog daemons, admission control pools, pipelined plan fragments with…

Read article →
ARTICLE · 093

Impala Query Hints, in depth: join distribution, STRAIGHT_JOIN, INSERT shuffle and clustering, scan scheduling, and when not to hint

A practical guide to Apache Impala optimizer hints: syntax and placement, BROADCAST versus SHUFFLE joins, STRAIGHT_JOIN, SHUFFLE/NOSHUFFLE and CLUSTER…

Read article →
ARTICLE · 094

Impala Query Plans, in depth: reading EXPLAIN from the bottom up, fragments and exchanges, join distribution, estimates versus actuals, and fixing bad plans

A practical guide to Impala query plans: how the planner builds single-node and distributed plans, how to read EXPLAIN output node by node, EXPLAIN_LE…

Read article →
ARTICLE · 095

Impala Query Queue, in depth

How Impala's admission queue works: admit, queue or reject decisions, per-pool limits and timeouts, how per-host memory to admit is computed, poo…

Read article →
ARTICLE · 096

Impala + Ranger + Kerberos, in depth: principals per daemon, load balancers and merged keytabs, delegation and the identity Ranger finally sees

How a Kerberized Impala cluster authenticates and authorizes a query: principals and keytabs for impalad, catalogd and statestored, --principal and --…

Read article →
ARTICLE · 097

Impala + Ranger Integration, in depth: where the check runs, shared Hive policies, masking, row filters and operating it safely

How Apache Impala enforces Apache Ranger policies: the coordinator-side plugin and catalogd's role, startup flags, sharing the hive service defin…

Read article →
ARTICLE · 098

Impala result spooling architecture

Deep-dive on Impala result spooling: why coupling execution to client fetch speed pins cluster resources, how the coordinator buffers results in memor…

Read article →
ARTICLE · 099

Impala Runtime Filters: Pushing Join Knowledge Into the Scan

How Impala builds Bloom, min/max, and IN-list runtime filters from a join's build side and pushes them into fact-table scans to skip row groups a…

Read article →
ARTICLE · 100

Impala Scale, in depth: dedicated coordinators, metadata size, on-demand catalog, admission control across coordinators and sizing a large cluster

What breaks as an Apache Impala cluster grows in nodes, tables, partitions, files and concurrent queries, and how to fix each: dedicated coordinator a…

Read article →
ARTICLE · 101

Impala Shell and Web UI, in depth: connecting and scripting with impala-shell, the daemon web pages, JSON scraping, query retention and locking it down

A practical guide to Impala's two operator tools: impala-shell protocols, ports, Kerberos, LDAP and TLS; interactive commands; scripting with -q,…

Read article →
ARTICLE · 102

Impala spill-to-disk architecture

Deep-dive on Impala spill-to-disk: how partitioned hash joins, aggregations, and sorts spill a victim partition to scratch disk when their memory rese…

Read article →
ARTICLE · 103

Impala statestore architecture

Deep-dive on the Impala statestore: a lightweight, soft-state publish/subscribe broker that disseminates cluster membership and catalog metadata acros…

Read article →
ARTICLE · 104

Impala COMPUTE STATS: Table and Column Stats

How Impala COMPUTE STATS and incremental stats feed the cost-based planner: table vs column stats, reading them with SHOW TABLE STATS, and keeping the…

Read article →
ARTICLE · 105

Impala Text and CSV Support, in depth: delimited tables, the quoting gap, NULLs, headers, bad data, compressed text and converting to Parquet

How Impala reads and writes delimited text: the text scan path, ROW FORMAT DELIMITED and its single-character rule, why quoted CSV fields break, skip.…

Read article →
ARTICLE · 106

Impala Troubleshooting, in depth: triage by symptom, reading the query profile, and a playbook for memory, admission, metadata and slow queries

A diagnostic method for Apache Impala: triage failures by symptom, capture and read the ExecSummary and query profile, and fix the common classes of p…

Read article →
ARTICLE · 107

Impala Upgrade Path, in depth

Upgrading Apache Impala in depth: the compatibility boundaries around the daemons, 3.x and 4.0 breaking changes (Hive 3 metastore, Ranger only, AVX, L…

Read article →
ARTICLE · 108

Impala vs Hive: Latency, Fault Tolerance, Memory and Metadata

Impala vs Hive compared: long-lived Impala daemons vs Hive on Tez DAGs, latency, fault tolerance, memory spill, metadata, ACID writes, LLAP, when to u…

Read article →