Apache Ambari is the open-source cluster manager for the Hadoop ecosystem. It installs HDFS, YARN, Hive, HBase, Kafka, ZooKeeper and friends onto a fleet of hosts, renders their configuration, starts and stops them in dependency order, runs health checks, wires in Kerberos and drives upgrades, all behind a web UI and a REST API. For a team running Hadoop on its own hardware, it is the difference between operating forty daemons by hand and operating one declared cluster.
This page explains how Ambari actually works (the server, the agents, the database and the desired-state model), then covers stack definitions, the REST API, configuration versioning, blueprints and alerts, and walks through a safe NameNode heap change on a highly available cluster. It ends with the failure modes that bite in production and a checklist. It also covers where the project stands in 2026, because that has changed more than most tooling.
Where Ambari stands in 2026
Ambari grew up at Hortonworks as the manager for the Hortonworks Data Platform (HDP). After the Cloudera merger, public HDP releases stopped, and in January 2022 the Ambari PMC voted to move the project to the Apache Attic for lack of active contributors. The community later brought it back out, and in April 2025 released Ambari 3.0.0. Its biggest change is that Apache Bigtop is now the default packaging, so the stack Ambari installs is built from community Bigtop packages rather than a vendor distribution. The release notes also list support for Rocky Linux 8 and 9 and openEuler 22.03, Java 17, a move to Python 3, new Grafana dashboards and fewer websocket connections.
The practical consequence: a new Ambari deployment in 2026 means Ambari 3.x plus a Bigtop stack, with packages mirrored into your own repository. Existing HDP 3.1 clusters under older Ambari versions still run, but they receive no public fixes, so treat them as a migration project. The alternatives are Cloudera Manager, a commercial product tied to Cloudera's distribution, Kubernetes operators for the components that have them, or plain configuration management (Ansible and the like), which gives you templates but none of Ambari's service-aware orchestration.
Architecture: server, agents, database
Ambari has three moving parts. The server is a Java process that hosts the web UI and the REST API, holds the cluster model, schedules operations and evaluates alert state. The database, PostgreSQL or MySQL in most deployments, is where that model lives: hosts, services, components, every version of every configuration, request history and alert definitions. The agent is a Python daemon on every host. It registers with the server, sends heartbeats with host and component status, receives commands (install, configure, start, stop, status checks, custom actions), executes them with service scripts, and runs the alert checks assigned to its host.
One property matters more than any other when you run this: the Hadoop daemons do not depend on Ambari at run time. If the server or its database goes down, HDFS keeps serving and YARN keeps scheduling; you lose management, alerts and the UI, not the data plane. That makes Ambari server downtime an operational blind spot rather than an outage.
The desired-state model
Everything in Ambari is a comparison between desired state and actual state. Each host component (a DataNode on host w7, say) has a desired state stored by the server and an actual state reported by the agent. The states you will meet are INSTALLED, which confusingly means installed and stopped, STARTED, and transitional or error states such as INSTALL_FAILED and UNKNOWN. When you ask Ambari to start HDFS, you are not running a command: you set desired state to STARTED and the server generates an ordered set of commands, respecting role dependencies such as JournalNodes before NameNodes, as an asynchronous request made of stages and tasks.
Configuration follows the same model. Config is grouped into types such as hdfs-site, core-site and hadoop-env. Each change creates a new immutable version identified by a tag, and the cluster points at one desired tag per type. When an agent runs configure or start, it renders files on disk from the desired version. So the files under /etc/hadoop/conf are outputs, not inputs. Edit one by hand and the next restart through Ambari silently overwrites it. Config groups let a subset of hosts (DataNodes with different disk layouts, say) override particular properties, and the UI flags components whose running config is older than the desired one as needing a restart.
Stacks and service definitions
Ambari knows nothing about Hadoop in its core. A stack is a versioned directory of service definitions, and each service has a metainfo.xml listing its components (with a category of MASTER, SLAVE or CLIENT and a cardinality such as 1+), its packages, its configuration types with defaults, and a command script per component. Command scripts are Python classes built on Ambari's resource_management library, which offers idempotent resources for packages, files, directories and commands. A custom service, such as an in-house indexing daemon you want managed alongside HDFS, is a new service directory with a script like this:
# ORDER_INDEXER/package/scripts/order_indexer.py - lifecycle script (sketch)
from resource_management.libraries.script.script import Script
from resource_management.core.resources.system import Execute, File
from resource_management.core.source import InlineTemplate
from resource_management.libraries.functions.check_process_status import check_process_status
class OrderIndexer(Script):
def install(self, env):
self.install_packages(env) # packages declared in metainfo.xml
def configure(self, env):
import params # derived from the command JSON the server sent
env.set_params(params)
File('/etc/order-indexer/indexer.properties',
content=InlineTemplate(params.indexer_properties_template),
owner=params.indexer_user, mode=0o640)
def start(self, env):
self.configure(env) # always re-render from desired state
import params
Execute(f'/usr/lib/order-indexer/bin/indexer --daemon --pid {params.pid_file}',
user=params.indexer_user,
not_if=f'test -f {params.pid_file} && ps -p $(cat {params.pid_file})')
def stop(self, env):
import params
Execute(f'kill $(cat {params.pid_file})', user=params.indexer_user)
File(params.pid_file, action='delete')
def status(self, env):
import params
check_process_status(params.pid_file) # raises if the process is not running
if __name__ == '__main__':
OrderIndexer().execute()The pattern to copy is that start calls configure first, so configuration always comes from the server's desired state, and that status is cheap and side-effect free, because agents run it constantly. Take the import paths from the stack you actually run rather than from this sketch: they have moved between Ambari versions.
The REST API
Everything the UI does goes through /api/v1, so everything can be scripted. Reads are plain GETs with a fields parameter to trim responses. Writes need the X-Requested-By header, a CSRF guard, and usually return a request id you poll.
AMBARI=https://ambari.example.com:8443 # port chosen in setup-security; 8080 is plain HTTP
C=prod1
AUTH="-u svc-ops:$AMBARI_PASSWORD" # a service account, not admin
H='X-Requested-By: ambari' # required on every modifying call
# Current desired config tag for every config type
curl -s $AUTH "$AMBARI/api/v1/clusters/$C?fields=Clusters/desired_configs"
# Maintenance mode: suppress alerts and exclude HDFS from bulk operations
curl -s $AUTH -H "$H" -X PUT "$AMBARI/api/v1/clusters/$C/services/HDFS" \
-d '{"RequestInfo":{"context":"CHG-1432 maintenance"},"Body":{"ServiceInfo":{"maintenance_state":"ON"}}}'
# Stop a service: desired state INSTALLED means installed and stopped
curl -s $AUTH -H "$H" -X PUT "$AMBARI/api/v1/clusters/$C/services/YARN" \
-d '{"RequestInfo":{"context":"Stop YARN"},"Body":{"ServiceInfo":{"state":"INSTALLED"}}}'
# -> {"href": ".../requests/57", "Requests": {"id": 57, "status": "Accepted"}}
# Poll the asynchronous request, or abort it if it hangs
curl -s $AUTH "$AMBARI/api/v1/clusters/$C/requests/57?fields=Requests/request_status,Requests/progress_percent"
curl -s $AUTH -H "$H" -X PUT "$AMBARI/api/v1/clusters/$C/requests/57" \
-d '{"Requests":{"request_status":"ABORTED","abort_reason":"hung on w17"}}'The most dangerous call is the configuration update. A PUT of desired_config replaces the entire config type, so sending only the property you want to change deletes every other property in that type. Always read, merge and write back, and keep the old tag for rollback:
import json, time, requests
S = requests.Session()
S.auth = ("svc-ops", PASSWORD)
S.headers["X-Requested-By"] = "ambari"
BASE = f"{AMBARI}/api/v1/clusters/{CLUSTER}"
def current(cfg_type):
desired = S.get(BASE, params={"fields": "Clusters/desired_configs"}).json()
tag = desired["Clusters"]["desired_configs"][cfg_type]["tag"]
item = S.get(f"{BASE}/configurations", params={"type": cfg_type, "tag": tag}).json()["items"][0]
return tag, item["properties"], item.get("properties_attributes", {})
def set_props(cfg_type, changes, note):
old_tag, props, attrs = current(cfg_type)
body = {"Clusters": {"desired_config": [{
"type": cfg_type,
"tag": f"version{int(time.time() * 1000)}",
"properties": {**props, **changes}, # the PUT replaces the WHOLE type
"properties_attributes": attrs,
"service_config_version_note": note}]}}
S.put(BASE, data=json.dumps(body)).raise_for_status()
return old_tag # keep it for rollback
old = set_props("hadoop-env", {"namenode_heapsize": "16384m"}, "CHG-1432 NN heap 8g to 16g")
Worked example: a NameNode heap change on an HA cluster
A 40-node cluster runs HDFS with active and standby NameNodes. File count has grown and the active NameNode is spending too long in garbage collection, so the change is to raise namenode_heapsize in hadoop-env from 8 GB to 16 GB without downtime. First, confirm both NameNode hosts have the memory and that the standby is healthy and caught up. Then turn on maintenance mode for HDFS so the restarts do not page anyone, and apply the change with the read-merge-write function above, noting the change ticket in the version note. Ambari now marks both NameNodes as needing a restart.
Restart the standby first, from the UI or by setting its host component state to INSTALLED and then STARTED. Wait until it has left safe mode and its edit-log tailing is current. Fail over so the restarted node becomes active, confirm clients are healthy, and restart the other NameNode. Finally turn maintenance mode off and confirm that no alerts are firing. If anything goes wrong, rollback is another desired-config PUT pointing at the old properties, or the 'make current' action on the previous service config version in the UI, followed by the same rolling restart.
Blueprints: clusters as data
A blueprint is a JSON document describing host groups (which components run together), stack version and configuration, without hostnames. A cluster creation template then maps real hosts, or host counts with predicates, onto those groups. Exporting a working cluster as a blueprint is the fastest way to build an identical staging environment or to rebuild after a disaster.
# Export a running cluster as a blueprint: topology and configs, no hostnames
curl -s $AUTH "$AMBARI/api/v1/clusters/prod1?format=blueprint" > prod1-blueprint.json
# Register it under a name (stack name and version come from GET /api/v1/stacks)
curl -s $AUTH -H "$H" -X POST "$AMBARI/api/v1/blueprints/hdfs-yarn-ha" -d @prod1-blueprint.json
# cluster-template.json maps host groups onto real hosts
{
"blueprint": "hdfs-yarn-ha",
"default_password": "set-then-rotate",
"host_groups": [
{"name": "master_1", "hosts": [{"fqdn": "m1.dc2.example.com"}]},
{"name": "master_2", "hosts": [{"fqdn": "m2.dc2.example.com"}]},
{"name": "worker", "host_count": "8", "host_predicate": "Hosts/cpu_count>=16"}
]
}
# Create the cluster; the response is a request id to poll like any other
curl -s $AUTH -H "$H" -X POST "$AMBARI/api/v1/clusters/staging2" -d @cluster-template.jsonPassword properties are not exported, so supply them in the template or through a credential store, and review the exported configuration before reusing it, because it contains site-specific values such as directory paths and hostnames baked into properties.
Security, alerts and metrics
Ambari's Kerberos wizard creates service principals in an MIT KDC or Active Directory, distributes keytabs to hosts and rewrites the relevant configuration; Hadoop Kerberos architecture explains what it is automating. It also installs and configures Apache Ranger for authorisation. For Ambari itself, enable TLS on the server with ambari-server setup-security, connect users to LDAP with ambari-server setup-ldap, and give automation its own account with only the permissions it needs.
Alert definitions ship with each service and come in several types: PORT (is a socket open), WEB (does an HTTP endpoint answer), METRIC (is a JMX value within thresholds), SCRIPT (arbitrary Python check) and AGGREGATE (what percentage of DataNodes are alerting). Agents evaluate them on schedules and the server sends notifications to email or SNMP targets. Ambari Metrics collects host and service metrics for its dashboards and Grafana. Because every alert flows through the server, a dead Ambari server means silence. Monitor the server from outside Ambari.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| A hand-made config change disappears | Agents re-render files from desired state on restart. | Make every change through Ambari; use config groups for host-specific values. |
| Properties vanish after a scripted change | Partial desired_config PUT replaced the whole type. | Read, merge, write; roll back to the previous tag. |
| Host shows heartbeat lost | Agent down, FQDN mismatch between hostname -f and registration, or TLS trust failure. | Check /var/log/ambari-agent/ambari-agent.log and the server hostname in ambari-agent.ini. |
| Operations queue behind one task | A hung command on one host blocks its request. | Abort the request by API, fix the host, retry. |
| Install fails across hosts | Package repository unreachable or unsigned. | Mirror the Bigtop repository internally and pin it in the stack's repository settings. |
| Cluster unmanageable after a server loss | Ambari database lost with no backup. | Nightly database dumps plus an exported blueprint; test restores. |
Trade-offs
- Service awareness: Ambari understands role ordering, HA topologies and rolling restarts, which general configuration management does not, but it owns your config files and fights anyone editing them.
- Open source with a small community: no licence cost and no vendor lock-in, but fixes arrive at community pace, and you own packaging mirrors and testing.
- Single server: simple to run and harmless to the data plane when down, but a gap in alerting and control.
- UI versus code: the UI is excellent for exploration, but production changes should go through scripted API calls or blueprints so they are reviewable and repeatable.
What to do next
- If you are on an old Ambari with an HDP stack, inventory versions and plan a move to Ambari 3.x with a Bigtop stack in a staging cluster first.
- Set up nightly backups of the Ambari database and an exported blueprint, and rehearse a restore.
- Replace hand-edited config files with Ambari config groups and forbid direct edits.
- Wrap config changes in a read-merge-write script that records the previous tag and a change-ticket note.
- Enable TLS and LDAP on the server and create a least-privilege automation account.
- Monitor the Ambari server and its database from an external system, and review which alerts actually page.
- Write a runbook for rolling restarts and DataNode decommissioning using the API calls above.