Hive 4 is less a single release than a line: 4.0.0 in March 2024, followed by 4.0.1, 4.1.0 and 4.2.0, and the security-fix release 4.2.1 in August 2026. Across that line Hive removed an execution engine, deprecated another along with its old command-line client, made Apache Iceberg a first-class table format, reworked compaction, extended the metastore and HiveServer2 security model, and twice raised the minimum Java version.
This page is a map of those features organised by what they change for a team running Hive: what you must do, what you can stop doing, and what you can adopt when ready. Feature lists are from the Apache Hive release notes and documentation as of October 2026. The procedure for upgrading a 2.x or 3.x estate (metastore schema, ACID preparation, client cut-over, rollback) is covered in the Hive upgrade path and is not repeated here.
The 4.x release line and its Java requirements
Start with the release line itself, because the Java requirement is the constraint most likely to decide which 4.x version you can run.
| Release | Date | Works with | Headline |
|---|---|---|---|
| 4.0.0 | 29 March 2024 | Hadoop 3.3.6, Tez 0.10.3 | Iceberg integration, ACID and compaction rework, HS2 and HMS security, Hive on Spark removed |
| 4.0.1 | 2 October 2024 | Hadoop 3.3.6, Tez 0.10.4 | Fixes; release notes urge moving workloads to Tez |
| 4.1.0 | 31 July 2025 | Hadoop 3.4.1, Tez 0.10.5 | JDK 17 support; Hive Metastore as a standalone binary and Docker image; Iceberg storage-partitioned joins, partition-level column statistics and table compaction |
| 4.2.0 | 23 November 2025 | Hadoop 3.4.1, Tez 0.10.5 | JDK 21 is the minimum; Iceberg deletion vectors, auto compaction, ViewCatalog, column defaults, REST catalog client, Z-ordering, variant type |
| 4.2.1 | 24 August 2026 | Hadoop 3.4.1, Tez 0.10.5 | Fixes three security vulnerabilities (metastore direct-SQL partition paths, HS2 SAML bearer tokens, Avro SerDe schema URLs) |
Read the table as a planning tool. If your platform must stay on JDK 17, the 4.1 line is your ceiling; 4.2.x requires JDK 21 everywhere Hive runs, including HiveServer2, the metastore and any client that embeds Hive libraries. If you are on 4.2.0 and expose HiveServer2 with SAML, or allow user-supplied Avro schema URLs, 4.2.1 is not optional.
What was removed or deprecated
Three things went away or started going away in 4.0, and they generate most upgrade work.
- Hive on Spark is removed. A
SET hive.execution.engine=spark;in a script or a session default no longer works. Every such job moves to Tez, which is the default engine, or to Spark SQL reading the same metastore tables. - Hive on MapReduce is deprecated. It still exists, but it receives no performance work, and some newer features depend on Tez. Rebalance compaction, for example, only runs as query-based compaction. Treat MapReduce as a migration target date, not a fallback.
- The Hive CLI is deprecated. The old
hiveshell ran the compiler inside the client process, bypassing HiveServer2 authentication and authorisation. Use Beeline against HiveServer2:beeline -u jdbc:hive2://hs2.example.com:10000/default. Scripts that relied on local CLI behaviour, such as reading local files or skipping Ranger checks, need review.
A useful inventory before upgrading: grep scheduler definitions, shell scripts and connection strings for execution.engine, for invocations of the bare hive command, and for hive -e and hive -f. Each hit is a job that must change, and finding them after cut-over is far more expensive than before.
The shape of a Hive 4 deployment
The architecture is recognisably the Hive 3 one. Clients connect to HiveServer2, which compiles SQL with a Calcite-based cost optimiser and runs it on Tez, optionally with LLAP daemons, against tables whose metadata lives in the Hive Metastore. What changed is inside each box.
Iceberg as a first-class table format
The largest functional change is that Iceberg tables are native citizens. A table declared STORED BY ICEBERG is created, queried and modified through ordinary HiveQL, including row-level DELETE, UPDATE and MERGE, with vectorised reads and writes. 4.0 added snapshot management, branches and tags, partition-level operations, in-place migration of existing tables, materialised views over Iceberg tables, and Iceberg major compaction through OPTIMIZE TABLE.
Later releases filled gaps that mattered for performance and interoperability. 4.1 added storage-partitioned joins, partition-level column statistics and table compaction. 4.2 added deletion vectors, automatic compaction, a REST catalog client, Z-ordering and the variant type. The REST catalog client is important for mixed estates: it lets Hive share tables through the same catalog service that Spark, Trino or Flink use, instead of every engine pointing at the metastore.
-- Hive 4: an Iceberg table is declared like any other Hive table.
CREATE TABLE sales.orders_ice (
order_id BIGINT, customer_id BIGINT, amount DECIMAL(12,2), ts TIMESTAMP)
STORED BY ICEBERG;
INSERT INTO sales.orders_ice SELECT order_id, customer_id, amount, ts FROM sales.orders;
-- Rewrite small files and apply deletes.
OPTIMIZE TABLE sales.orders_ice REWRITE DATA;Catalog choice, the commit path, time travel, branches and the maintenance schedule are covered in Hive and Iceberg tables. The decision this page is concerned with is whether to adopt Iceberg at all, which is a trade-off discussed below.
ACID and compaction: rebalance and pools
Hive ACID tables did not go away; they got cheaper to run. Transaction ids are generated by sequence, readers can avoid waiting on lock acquisition ("zero-wait readers"), read-only queries skip transactional overhead, and two concurrency-control modes are available. The visible operational changes are in compaction, whose basics are in Hive compaction.
Rebalance compaction fixes skew between the implicit bucket files that Hive creates even for tables you never declared as bucketed. Skew builds up after large updates or uneven inserts, and makes a few tasks run much longer than the rest. Rebalance redistributes rows evenly, optionally changes the bucket count, and optionally sorts the data.
ALTER TABLE sales.orders COMPACT 'REBALANCE';
ALTER TABLE sales.orders COMPACT 'REBALANCE' CLUSTERED INTO 64 BUCKETS;
ALTER TABLE sales.orders COMPACT 'REBALANCE' ORDER BY customer_id DESC;The documented limits are strict. Rebalance is never started automatically. It applies only to full ACID tables that are implicitly bucketed: explicitly clustered tables and insert-only tables are not supported. It works within a partition, runs only as query-based compaction, not MapReduce, and takes an exclusive write lock on the table while it runs. Schedule it in a maintenance window and never inside an ingest pipeline.
Compaction pooling stops one busy table starving the rest. Tables, partitions or whole databases are assigned to a named pool with the hive.compactor.worker.pool property, and workers are dedicated to pools with hive.compactor.worker.<pool>.threads. A manual request can name a pool with a POOL clause on ALTER TABLE ... COMPACT, which overrides the property. Unlabelled requests go to the default pool, and so do labelled requests that wait longer than hive.compactor.worker.pool.timeout (0 disables that fallback). At least one worker thread cluster-wide must serve the default pool, or unlabelled requests are never processed.
-- Give the high-churn ingest tables their own workers.
ALTER TABLE sales.orders SET TBLPROPERTIES ('hive.compactor.worker.pool'='ingest');
-- hive-site.xml on compactor hosts:
-- hive.compactor.worker.threads = 6 (the maximum across all pools)
-- hive.compactor.worker.ingest.threads = 4 (carved out of the 6; 2 remain for the default pool)
Compiler changes and scheduled queries
Query-level changes are mostly invisible until a plan improves. 4.0 upgraded Calcite to 1.25, added anti-join planning for NOT EXISTS and similar patterns, split updates into a delete plus an insert, prunes branches that cannot match, and uses column histogram statistics for better selectivity estimates on skewed columns. Materialised views gained Iceberg sources and improved refresh; their mechanics are in Hive materialized views.
Scheduled queries let HiveServer2 run SQL on a cron schedule without an external scheduler. The common use is keeping a materialised view fresh:
CREATE SCHEDULED QUERY refresh_daily_revenue
EVERY 10 MINUTES
AS ALTER MATERIALIZED VIEW sales.daily_revenue REBUILD;
-- Run it now instead of waiting for the next slot.
ALTER SCHEDULED QUERY refresh_daily_revenue EXECUTE;If the view can be rebuilt incrementally and its sources have not changed, the rebuild is a no-op, so a short interval is cheap.
Keep scheduled queries for metastore-local housekeeping like this. Anything with cross-system dependencies, such as waiting for an upstream load, still belongs in your workflow scheduler, which can see the dependency. HPL/SQL, the procedural SQL dialect, was also improved and integrated, and ESRI-based geospatial UDFs ship with Hive.
Security and operations
Security and operability features matter most to platform teams.
- HiveServer2 authentication. SAML 2.0 and JWT modes were added, and Kerberos and LDAP can be enabled in parallel, so service accounts keep keytabs while people log in through LDAP or single sign-on.
- Graceful shutdown. HiveServer2 can drain running queries before stopping, which makes rolling restarts possible without failing user sessions.
- Metastore access. The metastore can serve Thrift over HTTP, accept JWT authentication and register in ZooKeeper for discovery, and it elects a leader dynamically for background housekeeping tasks. From 4.1 it ships as a standalone component with its own Docker image, which suits teams that only need a catalog for Spark or Trino.
- Replication. Bootstrap is faster, external tables can be replicated from snapshots, and replication has checkpoints and better metrics.
- Packaging. Official Docker images, Apache Ozone support and AArch64 builds.
The 4.2.1 security fixes are a reminder that the newest features carry the newest bugs. Track the Hive security announcements alongside your Hadoop and Tez versions.
Worked example: an adoption order
Consider a team running Hive 3 with Tez and a nightly ACID-based orders warehouse, several Hive-on-Spark jobs, and analysts using the Hive CLI. A sensible adoption order on Hive 4 is:
- Cut-over blockers first. Port the Hive-on-Spark jobs to Tez (most run unchanged once the engine setting is removed) and move analysts to Beeline. These are prerequisites, not features.
- Pick the Java line. JDK 17 shops land on 4.1.x; if JDK 21 is available everywhere, go to the latest 4.2.x and take its security fixes.
- Fix compaction contention. Put the high-churn
ordersandorder_itemstables in an ingest pool, leaving default-pool workers for everything else, then rebalance any table whose task durations show skew, inside a maintenance window. - Automate refreshes. Replace a cron job that rebuilds materialised views with scheduled queries.
- Pilot Iceberg on one new table. Choose a table that another engine also needs to read, create it
STORED BY ICEBERG, and run it beside the ACID tables for a quarter before migrating anything existing.
Each step is independent and reversible, which is the point: Hive 4 rewards adopting features one at a time against measured problems rather than all at once.
Failure modes
- A job still sets the Spark engine. It fails after the upgrade. Inventory before cut-over, as above.
- Wrong JDK on one component. A 4.2.x metastore on JDK 17 does not start. Pin the JDK in every image and host role.
- No default-pool workers. After introducing pools, unlabelled tables silently stop compacting and delta directories pile up. Alert on the age of the oldest initiated compaction.
- Rebalance during ingest. The exclusive write lock blocks writers until it finishes.
- Assuming Iceberg maintenance is free. Snapshots, orphan files and small files still need scheduled maintenance, even with auto compaction in 4.2.
- Two catalogs for one table. Pointing some engines at the metastore and others at a REST catalog for the same Iceberg table gives conflicting commits. Choose one owner per table.
Trade-offs: ACID or Iceberg, and which 4.x
The central choice in the 4.x line is ACID tables versus Iceberg tables. Stay on ACID when Hive (and Impala, through the metastore) are the only engines, your operational tooling is built around compaction, and nothing needs branches or time travel. Move to Iceberg when Spark, Trino or Flink must read and write the same tables, when partition evolution or snapshot rollback would save real work, or when you want the newer 4.1 and 4.2 performance features, most of which target Iceberg. Hive ACID transactions explains what you would be leaving behind.
The second choice is version. 4.2.x carries the most Iceberg work and the latest security fixes, but requires JDK 21. 4.1.x is the conservative choice for JDK 17 estates. Staying on 4.0.x is hard to justify for new deployments.
What to do next
- Inventory jobs for the Spark engine, bare
hiveCLI calls and Hadoop or Tez version pins. - Decide the Java line (17 or 21) and therefore the 4.x version, and record why.
- Move all interactive users and scripts to Beeline against HiveServer2.
- Enable compaction pools for your two or three highest-churn tables and keep default-pool workers.
- Find skewed ACID tables from task-duration spread and schedule rebalance compaction in a maintenance window.
- Replace external materialised-view refresh jobs with scheduled queries.
- Pilot one Iceberg table that a second engine needs, and pick a single catalog owner for it.
- Subscribe to Hive security announcements and plan to take fix releases such as 4.2.1 within weeks.