Azure HDInsight is Microsoft's managed service for running open-source Hadoop-stack clusters: Spark, Hive on Tez and LLAP, Kafka, HBase and the classic Hadoop MapReduce stack, deployed onto Azure virtual machines with Ambari for management. You choose a cluster type and size, Azure provisions and patches the nodes, and you keep the knobs you know from on-premises Hadoop. That combination is both its appeal and its trap: it feels like your old cluster, so teams run it like their old cluster, and pay for idle machines around the clock.
This article explains the service as it stands in October 2026, from first principles: what is supported, how a cluster is put together, why storage and metadata must live outside it, a worked deployment you can adapt, how autoscale really behaves, the security model, the common failures, and when another service is the better choice. Version and support facts were checked against Microsoft Learn on 2026-10-03; re-check them before you plan, because Microsoft changes defaults without notice. If you are moving an on-premises estate, read this alongside the Hadoop Cloud Migration Playbook.
Where HDInsight stands in 2026
There is exactly one supported version: HDInsight 5.1, released in November 2023 and on standard support with no retirement date announced. HDInsight 4.0 and 5.0 both retired on 31 March 2025, which means no new clusters, no fixes and the right for Microsoft to remove existing ones. HDInsight on AKS, the container-based preview that offered Flink and Trino, was retired on 31 January 2025 and is not an option for new work. If a design document or a stub page tells you otherwise, it is out of date.
Version 5.1 runs on Ubuntu 18.04.5 with a 5.4 kernel. Canonical's standard support for 18.04 ended long ago, so HDInsight relies on Extended Security Maintenance for OS patches. Microsoft ships maintenance images every few months; the most recent release note at the time of writing is dated 29 June 2026 and is a security-fix release. The component versions are fixed per HDInsight version:
| Component | HDInsight 5.1 |
|---|---|
| Apache Spark | 3.3.1 |
| Apache Hadoop | 3.3.4 |
| Apache Hive / Tez | 3.1.2 / 0.9.1 |
| Apache Kafka | 3.2.0 |
| Apache HBase / Phoenix | 2.4.11 / 5.1.2 |
| Apache Ranger | 2.3.0 |
| Apache Oozie | 5.2.1 |
| Apache ZooKeeper | 3.6.3 |
| Apache Ambari | 2.7.3 |
Two consequences follow. First, you cannot pick a newer Spark or Kafka than the table says; if you need Spark 3.5 or 4.x features, HDInsight is the wrong service. Second, Sqoop and Pig were dropped in 5.1, so legacy ingestion jobs that call them need rewriting before they can move.
Cluster anatomy: who does what
Every cluster is a set of Azure VMs with fixed roles. Two head nodes run the master services: the YARN ResourceManager, the HDFS NameNode for the small local file system, HiveServer2, Ambari server and, depending on cluster type, Livy, the Spark history server or the HBase master. They run active and standby, and a ZooKeeper ensemble of three nodes arbitrates failover. Gateway nodes sit in front and are the only path from the public endpoint, https://CLUSTER.azurehdinsight.net, to the cluster; they terminate HTTPS and enforce the cluster login. Worker nodes run the YARN NodeManager and do the actual work. On Kafka clusters the workers are brokers and on HBase clusters they are region servers, which is why those two types behave differently when you resize them.
The defining design decision is that the default file system is not HDFS on the workers. It is Azure storage, usually Azure Data Lake Storage Gen2 addressed with the abfs:// scheme. Your tables, jar files, Spark event logs and YARN application logs all land there. The workers still have local disks and a local HDFS for shuffle spill and scratch data, but nothing you need tomorrow should live on them. Metadata follows the same rule: Hive, Oozie and Ambari each keep state in a SQL database. By default HDInsight creates a small internal database that is deleted with the cluster. A custom metastore on Azure SQL Database that you own survives deletion, can be shared by several clusters, and must be chosen when the cluster is created.
That is what makes HDInsight affordable: a nightly cluster can be created at 01:00, run against existing data and tables, and be deleted at 04:00. A cluster kept alive only because it holds the metastore is a design error.
Choosing a cluster type
A cluster has exactly one type, set at creation. Pick it by the workload, not by habit:
| Type | Use it for | Watch out for |
|---|---|---|
| Spark | Batch ETL, PySpark and Scala jobs, notebooks via Jupyter or Zeppelin | Spark is pinned at 3.3; executor sizing is yours |
| Hadoop | Hive on Tez batch queries, MapReduce, Oozie workflows | Mostly legacy; prefer Spark for new pipelines |
| Interactive Query | Hive LLAP for low-latency SQL over shared tables | Autoscale is schedule-based only |
| Kafka | Event ingestion with HDInsight-managed brokers | No autoscale; broker disks and partition rebalancing are yours |
| HBase | Wide-column random reads and writes, Phoenix SQL | No autoscale; region server scaling is manual |
To mix types, run several clusters against one storage account and metastore, for example a long-lived Interactive Query cluster for analysts plus short-lived Spark clusters for ETL.
Worked example: an ephemeral Spark cluster with Livy
Suppose a team runs a nightly job that reads raw clickstream JSON, sessionises it and writes a partitioned Parquet table. The data is in an ADLS Gen2 account, lakeprod, container raw. The plan: create a 5.1 Spark cluster that authenticates to storage with a user-assigned managed identity, submit the job over Livy, and delete the cluster when it finishes. The identity needs the Storage Blob Data Owner role on the account before creation. The Azure CLI call below uses flags from the current az hdinsight create reference; check them against your CLI version.
az hdinsight create \
--name etl-nightly-0103 --resource-group rg-analytics --location westeurope \
--type spark --version 5.1 --component-version Spark=3.3 \
--http-user admin --http-password "$CLUSTER_PW" \
--ssh-user sshuser --ssh-password "$SSH_PW" \
--storage-account lakeprod --storage-filesystem hdi-etl-nightly \
--storage-account-managed-identity /subscriptions/$SUB/resourceGroups/rg-analytics/providers/Microsoft.ManagedIdentity/userAssignedIdentities/hdi-msi \
--workernode-count 4 --workernode-size Standard_E8_v3 \
--vnet-name vnet-analytics --subnet hdiThe job itself is ordinary PySpark. Note that every path is an abfs:// URI, so the code would run unchanged on the next cluster:
from pyspark.sql import SparkSession, functions as F, Window
spark = SparkSession.builder.appName("sessionise").getOrCreate()
src = "abfs://raw@lakeprod.dfs.core.windows.net/clicks/dt=2026-10-02/"
dst = "abfs://curated@lakeprod.dfs.core.windows.net/sessions/"
clicks = spark.read.json(src)
w = Window.partitionBy("user_id").orderBy("ts")
gap = F.col("ts").cast("long") - F.lag("ts").over(w).cast("long")
sessions = (clicks
.withColumn("new_session", (gap.isNull() | (gap > 1800)).cast("int"))
.withColumn("session_no", F.sum("new_session").over(w))
.groupBy("user_id", "session_no")
.agg(F.min("ts").alias("start"), F.max("ts").alias("end"), F.count("*").alias("events")))
(sessions.withColumn("dt", F.lit("2026-10-02"))
.write.mode("overwrite").partitionBy("dt").parquet(dst))Submission goes through Livy's REST API behind the gateway. Livy requires the X-Requested-By header on POST requests, and the gateway uses the cluster login for basic authentication:
curl -s -u admin:"$CLUSTER_PW" -H "Content-Type: application/json" -H "X-Requested-By: admin" \
-X POST https://etl-nightly-0103.azurehdinsight.net/livy/batches \
-d '{"file": "abfs://jobs@lakeprod.dfs.core.windows.net/sessionise.py",
"conf": {"spark.sql.shuffle.partitions": "64"}}'
# poll: GET /livy/batches/<id>/state until "success" or "dead", then:
az hdinsight delete --name etl-nightly-0103 --resource-group rg-analytics --yesIn production, drive the same three steps from an orchestrator such as Azure Data Factory or Airflow, and make deletion unconditional: a failed job that leaves its cluster running is the most common source of surprise bills.
Autoscale: what it does and does not do
Autoscale changes the number of worker nodes only; head and ZooKeeper nodes never scale. It comes in two forms. Load-based autoscale samples YARN metrics every 60 seconds: total pending CPU and memory needed to start waiting containers against total free CPU and memory on active workers. If pending exceeds free for roughly three to five minutes it adds enough nodes to cover the shortfall; if free exceeds pending for the same time it removes nodes, preferring workers with no Application Master containers, and decommissions them gracefully so running Spark and Hadoop tasks finish first. Schedule-based autoscale sets a node count at fixed times of day and, importantly, does not decommission gracefully, so a scale-down that lands mid-job can kill tasks.
Support differs by type. Spark and Hadoop clusters can use either form. Interactive Query supports only schedule-based scaling, and scaling changes the LLAP daemon count without persisting it in Ambari, so a manual Hive restart resets it. Kafka and HBase do not support autoscale at all. Never set the minimum below three workers: the local HDFS replicates three ways, and with fewer nodes the NameNode can get stuck in safe mode, the same failure described in HDFS Safe Mode. Finally, a scale operation takes ten to twenty minutes end to end, so schedule a 09:00 capacity bump for 08:30, and remember that persisted script actions run on every new node and add to that time.
{ "name": "workernode", "targetInstanceCount": 4,
"autoscale": { "capacity": { "minInstanceCount": 3, "maxInstanceCount": 12 } } }
Security: network, identity and authorisation
Start with the network. Deploy clusters into your own virtual network so workers can reach private endpoints for storage and SQL, and restrict inbound traffic with network security groups using the HDInsight service tag for the management plane. Without that, the gateway endpoint is reachable from the internet and protected only by the cluster login password. For storage, prefer a managed identity over account keys: keys grant full access to everything in the account and tend to be copied into notebooks.
For multi-user clusters, the Enterprise Security Package (ESP) joins the cluster to a Microsoft Entra Domain Services managed domain so users authenticate with their corporate identity through Kerberos, and installs Apache Ranger for fine-grained authorisation on Hive tables, columns and HDFS paths. ESP is the only way to get per-user policies, auditing and row filtering on HDInsight; without it everyone who has the cluster login is effectively the same user. It also brings operational weight: a managed domain to run and Kerberos tickets to debug. The policy model is the same Ranger you may already know from Apache Ranger.
Customising nodes with script actions
Script actions are Bash scripts, stored at a URI the cluster can read, that HDInsight runs as root on chosen node types. They are how you install a Python wheel, a JDBC driver or a monitoring agent. A script marked persisted is re-run on every node added later, including by autoscale. Treat them like configuration management: make them idempotent, pin versions, log to a known path, and keep them short, because they sit on the scale-up critical path. Never use a script action to change a service configuration that Ambari also manages; Ambari will overwrite it on the next restart. Change those settings through Ambari's REST API instead, so they survive.
#!/usr/bin/env bash
# persisted script action: a separate env with a pinned package, on head and worker nodes
set -euo pipefail
LOG=/var/log/hdi-scriptaction-geo.log
BASE=${BASE:-/usr/bin/miniforge/envs/py38/bin/python} # documented Spark 3.x PYSPARK_PYTHON
ENV=/opt/venvs/geo
if [ ! -x "$ENV/bin/python" ]; then
"$BASE" -m venv --system-site-packages "$ENV" >>"$LOG" 2>&1
fi
"$ENV/bin/pip" install --no-cache-dir "h3==3.7.7" >>"$LOG" 2>&1 # no-op if already installedMicrosoft's guidance is not to install into the built-in Python environments, because the cluster's own tooling depends on them. Build a separate environment as above, confirm the base path with echo $PYSPARK_PYTHON on a node first, and then point PYSPARK_PYTHON at the new environment in the Spark and Livy env settings through Ambari.
Failure modes and how to recognise them
| Symptom | Usual cause | Fix |
|---|---|---|
| Bill keeps growing with no jobs | Clusters are billed per minute from creation to deletion, idle or not | Ephemeral clusters; unconditional delete in the orchestrator |
| Tables vanish after a rebuild | Default internal metastore deleted with the cluster | External Azure SQL metastore set at creation |
| Scale-down kills jobs | Schedule-based scaling is not graceful | Load-based for Spark, or schedule around job windows |
| Cluster stuck after scale-down | Fewer than 3 workers, HDFS safe mode | Minimum of 3 workers |
| Config change lost after restart | Edited files directly instead of via Ambari | Apply through Ambari or its REST API |
| 403 from storage | Managed identity lacks Storage Blob Data role | Grant the role before creation |
| Slow scale-up | Heavy persisted script actions | Bake fewer, smaller scripts; cache artifacts in storage |
There is no in-place upgrade between versions or images: to pick up a new image you create a new cluster, one more reason to keep clusters stateless. Watch YARN queue pressure in Ambari or Azure Monitor; the YARN overview explains the scheduler metrics autoscale relies on.
When to choose something else
HDInsight is the right answer when you need open-source Hadoop-stack compatibility on Azure with VM-level control: existing Hive and Oozie estates, HBase with Phoenix, self-managed Kafka semantics, or custom native libraries on every node. It is a weak answer when you want current Spark versions, serverless billing or a managed lakehouse. For new Spark and SQL analytics, evaluate Azure Databricks and Microsoft Fabric; for Kafka-compatible ingestion without brokers, evaluate Event Hubs. Given the aging base OS, keep code portable with abfs:// paths, standard Spark APIs and an external metastore. Cost modelling for that decision is covered in Hadoop Cost Management.
What to do next
- Inventory every HDInsight cluster and confirm it runs 5.1; anything on 4.0 or 5.0 is retired and needs a rebuild now.
- Move Hive, Oozie and Ambari metadata to an Azure SQL database you own, and recreate clusters against it.
- Switch storage authentication from account keys to a managed identity with the narrowest Storage Blob Data role that works.
- Make batch clusters ephemeral: create, run, delete from your orchestrator, with deletion in a finally step.
- Enable load-based autoscale on long-lived Spark clusters with a minimum of three workers and a ceiling that fits your core quota.
- Rewrite script actions to be idempotent and short, and move service configuration changes into Ambari.
- Decide per workload whether HDInsight is still the right home, and record the reason.