Most Ambari material stops at installing a cluster and clicking through the UI. The operators who keep Ambari-managed clusters healthy for years spend their time on the next layer down: packaging in-house services so Ambari can install, configure and restart them like HDFS, keeping alert definitions in version control, keeping the metrics system alive, upgrading Ambari without disturbing the cluster and rebuilding the server when its host dies.
This page covers that layer. It assumes you know Ambari's server, agent and database architecture, its desired-state model, stacks, the REST API and blueprints, all of which are explained in Apache Ambari: server, agents, stacks, blueprints and the REST API. On status: the project went to the Apache Attic in January 2022, came back, and released Ambari 3.0.0 in April 2025 with Apache Bigtop as its default stack packaging; the 2.7.x line is end-of-life. Preview documentation for 3.1.0 describes a new Prometheus-compatible monitoring system replacing Ambari Metrics, but at the time of writing that release is unpublished, so this page describes Ambari Metrics as shipped in 3.0 and flags the change where it matters.
How a command reaches a host
When you restart a service, the server turns the request into stages, each holding tasks for specific hosts, and records them in the database. On the next heartbeat each agent receives its tasks. For each task it writes a command file under /var/lib/ambari-agent/data/, named command-<id>.json, containing the role, the command (INSTALL, START, STOP and so on), every configuration type the service depends on and host-level parameters. It then runs the component's Python command script, writing stdout and stderr to matching output-<id>.txt and errors-<id>.txt files.
Those three files are the first place to look when an operation fails. The command file shows exactly what Ambari believed the configuration to be, and the output shows each resource the script applied. The agent caches service scripts copied from the server under /var/lib/ambari-agent/cache/, so an edited script on the server only takes effect after agents resync. Restarting the agent forces that.
Writing a custom service
Any process you want Ambari to manage, such as an in-house ingest daemon, a monitoring agent or a component the stack does not ship, can be described as a custom service. The example here, QueueMon, is a small daemon that watches an ingest queue. A service is a directory under the stack definition:
/var/lib/ambari-server/resources/stacks/<STACK>/<VERSION>/services/QUEUEMON/
metainfo.xml # service, components, cardinality, dependencies
configuration/
queuemon-env.xml # properties shown and versioned in Ambari
package/
scripts/
params.py # read configs from the command JSON
queuemon_server.py # install / configure / start / stop / status
templates/
queuemon.conf.j2 # rendered on every configure
alerts/
alert_queue_depth.py # SCRIPT alert
alerts.json # alert definitions shipped with the serviceThe metainfo.xml file declares the service, its components, how many instances each may have and which other components must sit on the same host. The category is MASTER, SLAVE or CLIENT; cardinality is a number, a range such as 1+, or ALL.
<metainfo>
<schemaVersion>2.0</schemaVersion>
<services>
<service>
<name>QUEUEMON</name>
<displayName>QueueMon</displayName>
<comment>Watches ingest queue depth</comment>
<version>1.2.0</version>
<components>
<component>
<name>QUEUEMON_SERVER</name>
<displayName>QueueMon Server</displayName>
<category>MASTER</category>
<cardinality>1</cardinality>
<commandScript>
<script>scripts/queuemon_server.py</script>
<scriptType>PYTHON</scriptType>
<timeout>600</timeout>
</commandScript>
<dependencies>
<dependency>
<name>HDFS/HDFS_CLIENT</name>
<scope>host</scope>
<auto-deploy><enabled>true</enabled></auto-deploy>
</dependency>
</dependencies>
</component>
</components>
<requiredServices><service>HDFS</service></requiredServices>
<configuration-dependencies><config-type>queuemon-env</config-type></configuration-dependencies>
</service>
</services>
</metainfo>The scripts use Ambari's resource_management library, which applies resources (directories, files, templates, commands) idempotently in the style of a configuration-management tool. Since Ambari 3.0 this code runs on Python 3. The configure method re-renders configuration from desired state every time, which is what makes Ambari authoritative; status is polled frequently and must be cheap.
# params.py
from resource_management.libraries.script.script import Script
config = Script.get_config()
env = config['configurations']['queuemon-env']
user, port = env['queuemon_user'], env['queuemon_port']
pid_file = '/var/run/queuemon/queuemon.pid'
# queuemon_server.py
from resource_management.libraries.script.script import Script
from resource_management.core.resources.system import Directory, Execute, File
from resource_management.core.source import Template
from resource_management.libraries.functions.check_process_status import check_process_status
class QueueMonServer(Script):
def install(self, env):
self.install_packages(env) # packages listed in metainfo osSpecifics
self.configure(env)
def configure(self, env):
import params
env.set_params(params)
Directory('/var/run/queuemon', owner=params.user, create_parents=True)
File('/etc/queuemon/queuemon.conf', owner=params.user,
content=Template('queuemon.conf.j2')) # re-rendered every time: Ambari owns this file
def start(self, env):
import params
self.configure(env)
Execute(f'/usr/bin/queuemon --port {params.port} --pid {params.pid_file} --daemon',
user=params.user, not_if=f'test -f {params.pid_file} && kill -0 $(cat {params.pid_file})')
def stop(self, env):
import params
Execute(f'kill $(cat {params.pid_file})', user=params.user, only_if=f'test -f {params.pid_file}')
File(params.pid_file, action='delete')
def status(self, env):
import params
check_process_status(params.pid_file) # raises ComponentIsNotRunning if dead
if __name__ == '__main__':
QueueMonServer().execute()Restart the Ambari server so it rereads the stack, then add the service through the Add Service wizard or the REST API. For distribution to many clusters, management packs bundle services, stack extensions and their own versions into one archive, installed with ambari-server install-mpack. Confirm your Ambari release supports the mpack features you need, because mpack support has varied between versions.
Ordering, dependencies and idempotence
Ambari decides start order from role_command_order.json files at stack and service level, which say, for example, that QUEUEMON_SERVER-START waits for NAMENODE-START. Without an entry, Ambari may start your daemon before HDFS is ready during a full cluster restart. Host-scoped dependencies with auto-deploy make Ambari install the HDFS client wherever QueueMon lands. Cluster-scoped dependencies only require the component to exist somewhere.
Write every method so it can run twice. Agents retry, operators click Restart twice, and a restart after a host crash finds stale PID files. Guard commands with not_if and only_if, and never let status start anything. Test the script on one host by rerunning a saved command file before rolling it out, because a bug in configure fails every restart across the cluster.
Alerts as code
Each service ships its alert definitions in alerts.json. The simple types, PORT, WEB and METRIC, cover sockets, HTTP endpoints and JMX thresholds. A SCRIPT alert runs Python inside the agent on the alert's interval, which makes it the right tool for business-level checks such as queue depth:
{
"QUEUEMON": {
"QUEUEMON_SERVER": [
{
"name": "queuemon_queue_depth",
"label": "QueueMon ingest queue depth",
"interval": 2,
"scope": "ANY",
"enabled": true,
"source": {
"type": "SCRIPT",
"path": "QUEUEMON/package/alerts/alert_queue_depth.py",
"parameters": [
{"name": "warn_depth", "display_name": "Warning depth", "value": 50000, "type": "NUMERIC"},
{"name": "crit_depth", "display_name": "Critical depth", "value": 200000, "type": "NUMERIC"}
]
}
}
]
}
}# alert_queue_depth.py -- runs inside the agent on the alert's interval
import json, urllib.request
def get_tokens():
return ('{{queuemon-env/queuemon_port}}',) # config values the agent substitutes
def execute(configurations={}, parameters={}, host_name=None):
port = configurations['{{queuemon-env/queuemon_port}}']
try:
with urllib.request.urlopen(f'http://{host_name}:{port}/stats', timeout=5) as r:
depth = json.load(r)['depth']
except Exception as e:
return ('UNKNOWN', [f'stats endpoint unreachable: {e}'])
if depth >= float(parameters.get('crit_depth', 200000)):
return ('CRITICAL', [f'queue depth {depth}'])
if depth >= float(parameters.get('warn_depth', 50000)):
return ('WARNING', [f'queue depth {depth}'])
return ('OK', [f'queue depth {depth}'])get_tokens declares which configuration values the agent must substitute, and execute returns a state and the text for the alert message. Parameters become editable thresholds in the UI. Once a cluster is running, definitions live in the database, so edits made in the UI drift from the JSON in your repository. Export them regularly through the API, keep the export in git and review differences. Notification targets are resources too:
A=https://ambari.example.com:8443/api/v1; C=prod; AUTH="-u automation:$PW -H X-Requested-By:ambari"
# Alert definitions are resources: export them, keep them in git, diff before applying
curl -s $AUTH "$A/clusters/$C/alert_definitions?fields=*" > alert_definitions.json
# Page only on CRITICAL, by e-mail, for every alert group
curl -s $AUTH -X POST "$A/alert_targets" -d '{"AlertTarget": {
"name": "oncall-email", "notification_type": "EMAIL", "global": true,
"alert_states": ["CRITICAL"],
"properties": {"ambari.dispatch.recipients": ["oncall@example.com"]}}}'
# Ask the metrics collector for the last hour of a NameNode metric
curl -s "http://ams-collector.example.com:6188/ws/v1/timeline/metrics?metricNames=jvm.JvmMetrics.MemHeapUsedM&appId=namenode&hostname=nn1.example.com&startTime=$(( $(date +%s) - 3600 ))000&endTime=$(date +%s)000"
The metrics system
Ambari Metrics (AMS) has three parts. A Metrics Monitor on every host collects operating-system metrics. Hadoop sinks inside NameNode, ResourceManager, HBase and other daemons push service metrics. Both send to the Metrics Collector, which stores time series in an HBase instance queried through Apache Phoenix and serves them on port 6188 to Ambari's UI and the bundled Grafana. In embedded mode the collector's HBase writes to local disk on the collector host; in distributed mode it writes to HDFS, which scales further but means losing HDFS also loses metrics. Aggregators roll raw data into minute, hourly and daily tables per host and per cluster, each with its own TTL.
Most AMS trouble is sizing. The collector's HBase heap and region count must grow with the number of hosts and metrics. An undersized collector shows gaps in graphs, slow dashboards and sinks logging write failures, and METRIC alerts that read from it go UNKNOWN. Cut what is collected before adding hardware: whitelisting metric names, shorter raw-data TTLs and fewer per-host aggregates usually fix a struggling collector. Treat AMS as operational telemetry with short retention; export to a long-term system if you need capacity-planning history. If you plan to adopt Ambari 3.1 when it ships, expect the collector and its tuning to be replaced by the new Prometheus-compatible stack, and avoid building heavy custom tooling on the AMS API in the meantime.
Upgrading Ambari itself
Upgrading Ambari is separate from upgrading the stack it manages, and much less risky, because cluster services keep running while the control plane is down. The sequence is: back up, stop server and agents, upgrade packages, run the schema upgrade, start, verify heartbeats. Stack upgrades of HDFS, YARN and the rest are a different procedure; rolling upgrades explains the HDFS side.
# 0. Backups first: database dump, server config, keys, custom stack content
pg_dump -U ambari ambari > ambari-$(date +%F).sql
tar czf ambari-etc-$(date +%F).tgz /etc/ambari-server/conf /var/lib/ambari-server/keys \
/var/lib/ambari-server/resources/stacks /var/lib/ambari-server/resources/mpacks
# 1. Stop the control plane (cluster services keep running)
ambari-server stop
pdsh -w ^all_hosts 'ambari-agent stop'
# 2. Point the package manager at the new Ambari repository, then upgrade
dnf -y upgrade ambari-server
ambari-server upgrade # migrates the database schema
pdsh -w ^all_hosts 'dnf -y upgrade ambari-agent'
# 3. Start and verify
ambari-server start
pdsh -w ^all_hosts 'ambari-agent start'
curl -s $AUTH "$A/hosts?fields=Hosts/host_status,Hosts/last_heartbeat_time"Rehearse on a staging copy restored from the production database dump, because the schema upgrade is where custom services, unusual configuration groups and old leftover data cause trouble. Keep the old packages and the dump until the new version has run for a week. A failed schema upgrade is recovered by reinstalling the old version and restoring the dump, not by fixing tables by hand.
Rebuilding a lost server
Ambari has no built-in server high availability. If its host dies the cluster keeps running, since NameNodes and ResourceManagers do not depend on it, but you lose control, configuration history and alerting until it returns. Recovery needs four things captured in advance: a recent database dump, the server configuration directory, the keys directory (TLS and encryption keys for stored credentials) and the stack resources tree including custom services. Install the same Ambari version on a new host, restore the database and files, start the server, then point every agent at the new hostname in the server section of ambari-agent.ini if the name changed. Agents re-register on heartbeat and the desired state is intact.
If the database is gone and only a blueprint export survives, you can recreate the cluster definition but not configuration history or alert customisations. That is why the nightly dump matters more than the blueprint. Keep the Kerberos admin credentials and keytab procedures in the runbook, since a kerberized cluster's recovery also depends on the KDC; see Hadoop Kerberos. Monitor the Ambari server and the HA pairs it manages, such as NameNode HA, from a system outside Ambari, because a dead server raises no alert about itself.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Custom service edit has no effect | Agents still run cached scripts | Restart agents to resync the cache |
| Every restart of a service fails at configure | Script bug or missing config key in params.py | Rerun a saved command file on one host; read errors-N.txt |
| Service shows stopped while running | status reads the wrong PID file or is too slow | Fix the pid path; keep status cheap |
| Daemon starts before HDFS during full restart | No role_command_order entry | Add the dependency to role_command_order.json |
| Graphs have gaps, METRIC alerts UNKNOWN | Undersized metrics collector HBase | Whitelist metrics, cut TTLs, raise heap |
| Server fails to start after upgrade | Schema upgrade broke on unusual data | Reinstall old version, restore dump, fix in staging |
Trade-offs
- Custom services versus external tooling: managing your daemons in Ambari gives one restart, config and alert model, at the cost of writing and testing Python against an API with sparse documentation.
- Alerts in Ambari versus a central monitoring system: built-in definitions understand topology and HA, but they die with the server and suit operational checks better than long-term SLOs.
- Embedded versus distributed metrics: embedded is simple and independent of HDFS; distributed scales further but couples observability to the system being observed.
- Staying on Ambari versus migrating: 3.x is alive again but small. Every custom service you write is something to port if you move to another manager.
What to do next
- Read the command, output and error files of one recent operation to learn what your agents actually run.
- Package one in-house daemon as a custom service on a staging cluster, with idempotent scripts and a role order entry.
- Export alert definitions and targets through the API into git and review the diff weekly.
- Check metrics collector health and trim collected metrics before it falls over.
- Schedule nightly database dumps plus the config, keys and stack resources directories, and rehearse a rebuild on a fresh host.
- Rehearse the next Ambari upgrade on a staging copy of the production database.
- Track the Ambari 3.1 release and plan how a monitoring replacement would affect your METRIC alerts and dashboards.