People have been predicting the death of Apache Hive for a decade, and they keep being half right. Fewer teams start new interactive analytics on Hive, and engines such as Spark, Trino and Impala run most of the queries that once went through HiveServer2. Yet nearly every one of those engines still finds its tables through the Hive Metastore, and plenty of nightly ETL still runs as HiveQL on Tez. The 4.x line has shipped regular releases since 2024.

The useful question is which parts of Hive you will still run in three years. This article splits Hive into its separable parts, sets out where each is heading from the project's own release notes, and gives you a framework and an inventory script to choose a direction. Step-by-step upgrade mechanics are covered in the Hive upgrade path, and the list of removed features in Hive deprecated features.

Advertisement

Where Hive stands today: the facts

ReleaseDateWhat matters for planning
4.0.02024-03-29Hive on Spark removed; Hive on MapReduce and the Hive CLI deprecated; first-class Iceberg tables with DML, branches and tags
4.0.12024-10-02Maintenance release; the project encourages moving remaining workloads to Tez
4.1.02025-07-31JDK 17 support; standalone Hive Metastore shipped as a binary and Docker image; Calcite 1.33
4.2.02025-11-23JDK 21 becomes the minimum Java version; more Iceberg and metastore work
4.2.12026-08-24Bug fix and security release

The end-of-life dates matter as much as the releases. The project declared the 1.x line end of life on 2024-04-11, 2.x on 2024-05-20 and 3.x on 2024-10-08. If you run Apache Hive 3.1 from upstream, you are on an unmaintained line and will receive no security fixes. Vendor distributions have their own support windows, so check your vendor's matrix.

Two facts from that table drive most planning decisions. First, Tez is now the only supported execution engine for HiveQL. Any job still setting hive.execution.engine=mr is on a deprecated path and any job set to spark will not run on 4.x at all. Second, the JDK floor jumped from 8 to 21 between 3.x and 4.2. Every UDF jar, SerDe and custom hook compiled for Java 8 is now code you have to rebuild and retest.

Hive is three things, and their futures diverge

One name, three components: each has a different futureQuery engineHiveServer2 + Tez (+ LLAP)Hive Metastore (HMS)Thrift API + Iceberg RESTTable formatsHive ACID ORC / Icebergshrinking share of computethe lakehouse catalogIceberg for new tablesHive (Tez)batch ETL, ACIDSparkETL, MLTrinointeractive SQLImpalaBI on CDPFlinkstreaming writesThriftThrift or RESTRESTHMSdatabases, tables, partitions, locks, statsRDBMS backendMySQL / PostgreSQL / OracleObject store / HDFSORC, Parquet, Iceberg metadataRanger / authpolicies on HMS objectsEngines are interchangeable; the catalog and the file layout are what you migrate slowly
Figure 1. Hive as three components. Many engines share one metastore, and the metastore and file layout are the slow-moving assets.

When someone says they run Hive, they usually mean some combination of three components that can be replaced independently.

The query engine is HiveServer2 compiling HiveQL through the Calcite cost-based optimizer into Tez DAGs, optionally with LLAP daemons for caching and low-latency execution. This is the part with the most competition from Spark, Trino, Impala and cloud warehouses.

The metastore (HMS) is a Thrift service in front of a relational database that stores databases, tables, partitions, column statistics, locks and transaction state. It is the part with the least competition in on-premises and hybrid estates, because every major engine can speak to it.

The table format is how data files and metadata are laid out: plain external tables, Hive ACID tables (ORC with base and delta directories managed by compaction), or Iceberg tables whose snapshots live in metadata files and whose pointer lives in the catalog. Hive 4 treats Iceberg as first-class: DML, copy-on-write and merge-on-read, branches and tags, and a migration command.

So plan each component separately. Treating Hive as one thing to keep or kill forces a far bigger migration than you need.

Advertisement

The metastore as the lakehouse catalog

The metastore is the part most likely to still be running in five years. Two developments define its future.

The first is the Iceberg REST catalog. The Iceberg project defines a REST protocol for catalogs so that engines do not need a catalog-specific client. HMS can serve that protocol: the Hive admin documentation lists metastore.catalog.servlet.port (required, default -1, meaning disabled) for the port the REST API listens on, and metastore.catalog.servlet.auth with values jwt (the default), oauth2, simple and none. Simple trusts an x-actor-username header and none disables authentication; the documentation flags both as unsuitable for production.

<!-- metastore-site.xml: serve Iceberg REST next to Thrift on the same HMS -->
<property>
  <name>metastore.catalog.servlet.port</name>
  <value>9001</value>
</property>
<property>
  <name>metastore.catalog.servlet.auth</name>
  <value>oauth2</value>   <!-- jwt is the default; never 'none' outside a laptop -->
</property>

The second development runs in the opposite direction: Hive as a client of someone else's catalog. The Hive 4.2.0 quickstart configures HiveServer2 to use an external Iceberg REST catalog such as Apache Polaris or Apache Gravitino by setting metastore.client.impl to org.apache.iceberg.hive.client.HiveRESTCatalogClient, naming a default catalog, and setting iceberg.catalog.<name>.type to rest with a URI and OAuth2 credentials. In that design the Hive engine keeps working while the source of truth for table metadata moves to a catalog shared with non-Hive tools.

Together these mean the catalog decision is now explicit. You can keep HMS as the catalog and let REST clients in, or move to a dedicated REST catalog and keep Hive as one of its clients. Both are supported upstream; neither requires rewriting data files for Iceberg tables. The standalone metastore packaging in 4.1 also matters operationally, because it lets you upgrade and scale HMS on its own schedule, independent of HiveServer2. High availability for that tier is covered in the Hive Metastore guide.

The engine: Tez only, and a smaller niche

With MapReduce deprecated and Hive on Spark removed, the HiveQL engine has converged on Tez, with LLAP as the low-latency option. That simplifies operations: one DAG engine to tune and one UI.

Hive's niche has narrowed but is real. It remains a strong choice for large SQL-defined batch transformations, especially ones that rely on Hive ACID merge semantics, materialized views with automatic query rewriting, or existing HiveQL that would be costly to port. Where LLAP is in production, the LLAP guide covers its sizing and cache behaviour.

The JDK 21 floor is the hidden cost. Custom UDFs, SerDes, authorization hooks and metastore listeners are often old, unowned jars. Each must be rebuilt, and some depend on libraries that never moved past Java 8. Inventory them before choosing a target version, not after.

The table format: Hive ACID versus Iceberg

Hive ACID was the project's answer to updates on a data lake: transactional ORC tables, a transaction manager in HMS, and background compaction that merges delta directories into new bases. Only engines that implement Hive ACID semantics can write it safely.

Iceberg solves the same problem with a format any engine can implement: immutable data files, manifest files and snapshot metadata, with an atomic pointer swap in the catalog at commit. A table written by Hive can be read and written by Spark, Trino, Impala and Flink with the same semantics. In Hive 4 you create one with STORED BY ICEBERG, and existing external tables can be migrated in place by switching the storage handler, which rewrites metadata but not data files.

-- New tables: Iceberg from day one (Hive 4)
CREATE TABLE sales.orders (
  order_id BIGINT, customer_id BIGINT, amount DECIMAL(12,2), order_ts TIMESTAMP)
PARTITIONED BY SPEC (day(order_ts))
STORED BY ICEBERG
TBLPROPERTIES ('format-version'='2');

-- Existing external Parquet/ORC table: migrate metadata in place, data files untouched
ALTER TABLE sales.orders_legacy
SET TBLPROPERTIES ('storage_handler'='org.apache.iceberg.mr.hive.HiveIcebergStorageHandler');

Hive ACID managed tables are the hard case: their delta layout is not Iceberg's, so moving them generally means rewriting the data, for example with CREATE TABLE ... STORED BY ICEBERG AS SELECT followed by a swap; the CTAS target exists, empty, until the query commits, so swap only after it finishes. Budget compute and a cut-over window for each large ACID table. Iceberg table design on Hive is covered in Hive Iceberg tables.

Four realistic futures for a Hive estate

FutureWhat you doChoose it whenMain cost
A. Modernise in placeUpgrade to 4.x on Tez, keep HMS, keep Hive ACID where it worksHiveQL ETL is large and stable, one engine dominatesJDK 21 rebuilds, schema upgrade, still Hive-only ACID
B. Keep HMS, move computeHMS stays the catalog; Spark, Trino or Impala run most queriesMany engines already read the tables; Hive jobs are shrinkingPorting HiveQL, reconciling SQL dialects
C. Iceberg on a shared catalogNew tables in Iceberg; serve HMS over REST or adopt a REST catalogSeveral engines must write the same tablesRewriting ACID tables, catalog security and ownership
D. Leave the Hadoop stackMove tables and jobs to a cloud warehouse or managed lakehouseHardware refresh or data centre exit is already plannedFull migration, egress, retraining, lock-in

These are not exclusive. A common sequence is A, then C for new tables, then B as workloads move, with D only if infrastructure strategy demands it. Avoid doing D implicitly, one team at a time, which leaves two catalogs with no owner. The engine trade-offs between Hive and Impala in particular are compared in Impala versus Hive.

Worked example: inventory the metastore before you decide

Decisions about Hive's future are usually made from opinions. Make them from an inventory instead. The metastore's backing database already knows every table's type, format, size hints and transactional status. The query below runs against the standard HMS schema (table and column names as in the upstream schema scripts; written for a PostgreSQL backend, so drop the double quotes on MySQL) and classifies every table into a migration bucket. Run it against a read replica.

SELECT d."NAME" AS db, t."TBL_NAME" AS tbl, t."TBL_TYPE" AS tbl_type,
       s."INPUT_FORMAT" AS input_format,
       MAX(CASE WHEN p."PARAM_KEY" = 'transactional'    THEN p."PARAM_VALUE" END) AS txn,
       MAX(CASE WHEN p."PARAM_KEY" = 'table_type'       THEN p."PARAM_VALUE" END) AS table_type,
       MAX(CASE WHEN p."PARAM_KEY" = 'totalSize'        THEN p."PARAM_VALUE" END) AS total_size,
       MAX(CASE WHEN p."PARAM_KEY" = 'transient_lastDdlTime' THEN p."PARAM_VALUE" END) AS last_ddl
FROM "TBLS" t
JOIN "DBS" d ON d."DB_ID" = t."DB_ID"
LEFT JOIN "SDS" s ON s."SD_ID" = t."SD_ID"
LEFT JOIN "TABLE_PARAMS" p ON p."TBL_ID" = t."TBL_ID"
GROUP BY d."NAME", t."TBL_NAME", t."TBL_TYPE", s."INPUT_FORMAT";
import csv, collections

def bucket(r):
    if (r["table_type"] or "").upper() == "ICEBERG":
        return "iceberg: already portable"
    if r["tbl_type"] == "VIRTUAL_VIEW":
        return "view: port SQL, check dialect"
    if (r["txn"] or "").lower() == "true":
        return "hive acid: rewrite to migrate"
    fmt = (r["input_format"] or "").lower()
    if "parquet" in fmt or "orc" in fmt:
        return "external columnar: in-place Iceberg candidate"
    return "text/other: convert format first"

rows = list(csv.DictReader(open("hms_tables.csv")))
size = collections.Counter()
count = collections.Counter()
for r in rows:
    b = bucket(r)
    count[b] += 1
    size[b] += int(r["total_size"] or 0)
for b in sorted(count, key=lambda k: -size[k]):
    print(f"{b:48s} {count[b]:6d} tables {size[b] / 1e12:8.2f} TB")

Suppose the output shows 4,000 tables. 2,600 are external Parquet holding 900 TB, 300 are Hive ACID holding 120 TB, 900 are text staging tables of a few TB, and 200 are views. That estate points clearly at future C. Most of the data can become Iceberg with metadata-only migrations. The 300 ACID tables are the real project, and they are small enough to rewrite over a few weekends. The views are where dialect work concentrates if compute later moves to Trino or Spark. Statistics such as totalSize are only as fresh as the last ANALYZE or write, so treat them as estimates and cross-check large tables against storage listings.

Failure modes to plan for

  • Two catalogs drift apart. A team registers tables in a new REST catalog while jobs still write through HMS. Readers see different snapshots depending on the engine. Fix: one catalog owns each table, and the cut-over is a recorded event, not a gradual drift.
  • Metastore database exhaustion. Partition-heavy tables with millions of partitions make HMS calls slow and the backing database large. Iceberg moves partition tracking into table metadata, which is one of the strongest operational arguments for migrating.
  • Silent dialect differences. HiveQL, Spark SQL and Trino differ in implicit casts, null ordering, integer division and timestamp semantics. A ported job can succeed and produce different numbers. Run old and new side by side and compare row counts and aggregates before switching.
  • Compaction debt on ACID tables. If compaction falls behind, readers open thousands of delta files and queries slow down. Check compaction status before migrating, because rewriting a table with a large delta backlog costs far more.
  • Security gaps on new endpoints. Turning on the REST servlet with simple or none authentication exposes table metadata, and often storage locations, to anyone who can reach the port.

Trade-offs worth arguing about

Keeping HMS is cheap and widely compatible, but it ties you to a Thrift service and a relational schema designed for Hive's needs. A dedicated REST catalog offers finer-grained access control and credential vending in some implementations, but adds a new critical service and a new team to own it. Moving to Iceberg frees writers but adds snapshot expiry and compaction routines, and moving compute off Hive only reduces engines if the Hive jobs actually disappear.

What to do next

  1. Find out which Hive version and vendor support window you are on; if it is upstream 3.x, treat the upgrade as a security item.
  2. Run the metastore inventory against a replica and bucket every table by format, transactional status and size.
  3. Join the inventory with a month of audit logs to find which engines read which tables, and archive what nobody reads.
  4. List every custom UDF, SerDe, hook and listener jar, with an owner and a plan to rebuild it for JDK 21.
  5. Move any job still on the MapReduce or Spark engine to Tez and compare outputs.
  6. Make Iceberg the default for new tables, and pilot in-place migration on one large external Parquet table.
  7. Decide explicitly whether HMS remains the catalog; if Iceberg-native clients need access, enable the REST servlet with jwt or oauth2 on a test metastore first.
  8. Pick one of the four futures, write it down with dates, and review it against the inventory each quarter.
Key takeaway: Hive is three components with different futures. The engine has narrowed to Tez and a batch SQL niche, the metastore is becoming the shared lakehouse catalog with an Iceberg REST API alongside Thrift, and Iceberg is replacing Hive ACID as the table format for anything several engines touch. Upstream 1.x to 3.x are end of life and 4.2 requires JDK 21. Inventory the metastore, plan each component separately, and choose a written direction instead of letting the estate drift.