Most Ambari material stops at installing a cluster and clicking through the UI. The operators who keep Ambari-managed clusters healthy for years spend their time on the next layer down: packaging in-house services so Ambari can install, configure and restart them like HDFS, keeping alert definitions in version control, keeping the metrics system alive, upgrading Ambari without disturbing the cluster and rebuilding the server when its host dies.

This page covers that layer. It assumes you know Ambari's server, agent and database architecture, its desired-state model, stacks, the REST API and blueprints, all of which are explained in Apache Ambari: server, agents, stacks, blueprints and the REST API. On status: the project went to the Apache Attic in January 2022, came back, and released Ambari 3.0.0 in April 2025 with Apache Bigtop as its default stack packaging; the 2.7.x line is end-of-life. Preview documentation for 3.1.0 describes a new Prometheus-compatible monitoring system replacing Ambari Metrics, but at the time of writing that release is unpublished, so this page describes Ambari Metrics as shipped in 3.0 and flags the change where it matters.

Advertisement

How a command reaches a host

Ambari serverrequest -> stages -> tasksAmbari databasedesired state, requests, alertsresources/stacks/.../servicesmetainfo.xml, configuration/, scripts/Ambari agent (each host)heartbeat, command queueagent cacheservice scripts copied from servercommand-N.jsonconfigs + hostLevelParams + rolePython command scriptinstall / configure / start / stop / statusresource_managementExecute, File, Directory, TemplateService processconfigs rendered from desired statetasksyncinvokeA custom service plugs in at the bottom left: Ambari ships your scripts to every agent and calls them exactly like a built-in service.
From request to process: the server splits a request into tasks, agents fetch the service's scripts and run them with a JSON command file containing the desired configuration.

When you restart a service, the server turns the request into stages, each holding tasks for specific hosts, and records them in the database. On the next heartbeat each agent receives its tasks. For each task it writes a command file under /var/lib/ambari-agent/data/, named command-<id>.json, containing the role, the command (INSTALL, START, STOP and so on), every configuration type the service depends on and host-level parameters. It then runs the component's Python command script, writing stdout and stderr to matching output-<id>.txt and errors-<id>.txt files.

Those three files are the first place to look when an operation fails. The command file shows exactly what Ambari believed the configuration to be, and the output shows each resource the script applied. The agent caches service scripts copied from the server under /var/lib/ambari-agent/cache/, so an edited script on the server only takes effect after agents resync. Restarting the agent forces that.

Writing a custom service

Any process you want Ambari to manage, such as an in-house ingest daemon, a monitoring agent or a component the stack does not ship, can be described as a custom service. The example here, QueueMon, is a small daemon that watches an ingest queue. A service is a directory under the stack definition:

/var/lib/ambari-server/resources/stacks/<STACK>/<VERSION>/services/QUEUEMON/
  metainfo.xml                 # service, components, cardinality, dependencies
  configuration/
    queuemon-env.xml           # properties shown and versioned in Ambari
  package/
    scripts/
      params.py                # read configs from the command JSON
      queuemon_server.py       # install / configure / start / stop / status
    templates/
      queuemon.conf.j2         # rendered on every configure
    alerts/
      alert_queue_depth.py     # SCRIPT alert
  alerts.json                  # alert definitions shipped with the service

The metainfo.xml file declares the service, its components, how many instances each may have and which other components must sit on the same host. The category is MASTER, SLAVE or CLIENT; cardinality is a number, a range such as 1+, or ALL.

<metainfo>
  <schemaVersion>2.0</schemaVersion>
  <services>
    <service>
      <name>QUEUEMON</name>
      <displayName>QueueMon</displayName>
      <comment>Watches ingest queue depth</comment>
      <version>1.2.0</version>
      <components>
        <component>
          <name>QUEUEMON_SERVER</name>
          <displayName>QueueMon Server</displayName>
          <category>MASTER</category>
          <cardinality>1</cardinality>
          <commandScript>
            <script>scripts/queuemon_server.py</script>
            <scriptType>PYTHON</scriptType>
            <timeout>600</timeout>
          </commandScript>
          <dependencies>
            <dependency>
              <name>HDFS/HDFS_CLIENT</name>
              <scope>host</scope>
              <auto-deploy><enabled>true</enabled></auto-deploy>
            </dependency>
          </dependencies>
        </component>
      </components>
      <requiredServices><service>HDFS</service></requiredServices>
      <configuration-dependencies><config-type>queuemon-env</config-type></configuration-dependencies>
    </service>
  </services>
</metainfo>

The scripts use Ambari's resource_management library, which applies resources (directories, files, templates, commands) idempotently in the style of a configuration-management tool. Since Ambari 3.0 this code runs on Python 3. The configure method re-renders configuration from desired state every time, which is what makes Ambari authoritative; status is polled frequently and must be cheap.

# params.py
from resource_management.libraries.script.script import Script
config = Script.get_config()
env = config['configurations']['queuemon-env']
user, port = env['queuemon_user'], env['queuemon_port']
pid_file = '/var/run/queuemon/queuemon.pid'

# queuemon_server.py
from resource_management.libraries.script.script import Script
from resource_management.core.resources.system import Directory, Execute, File
from resource_management.core.source import Template
from resource_management.libraries.functions.check_process_status import check_process_status

class QueueMonServer(Script):
    def install(self, env):
        self.install_packages(env)          # packages listed in metainfo osSpecifics
        self.configure(env)

    def configure(self, env):
        import params
        env.set_params(params)
        Directory('/var/run/queuemon', owner=params.user, create_parents=True)
        File('/etc/queuemon/queuemon.conf', owner=params.user,
             content=Template('queuemon.conf.j2'))   # re-rendered every time: Ambari owns this file

    def start(self, env):
        import params
        self.configure(env)
        Execute(f'/usr/bin/queuemon --port {params.port} --pid {params.pid_file} --daemon',
                user=params.user, not_if=f'test -f {params.pid_file} && kill -0 $(cat {params.pid_file})')

    def stop(self, env):
        import params
        Execute(f'kill $(cat {params.pid_file})', user=params.user, only_if=f'test -f {params.pid_file}')
        File(params.pid_file, action='delete')

    def status(self, env):
        import params
        check_process_status(params.pid_file)   # raises ComponentIsNotRunning if dead

if __name__ == '__main__':
    QueueMonServer().execute()

Restart the Ambari server so it rereads the stack, then add the service through the Add Service wizard or the REST API. For distribution to many clusters, management packs bundle services, stack extensions and their own versions into one archive, installed with ambari-server install-mpack. Confirm your Ambari release supports the mpack features you need, because mpack support has varied between versions.

Advertisement

Ordering, dependencies and idempotence

Ambari decides start order from role_command_order.json files at stack and service level, which say, for example, that QUEUEMON_SERVER-START waits for NAMENODE-START. Without an entry, Ambari may start your daemon before HDFS is ready during a full cluster restart. Host-scoped dependencies with auto-deploy make Ambari install the HDFS client wherever QueueMon lands. Cluster-scoped dependencies only require the component to exist somewhere.

Write every method so it can run twice. Agents retry, operators click Restart twice, and a restart after a host crash finds stale PID files. Guard commands with not_if and only_if, and never let status start anything. Test the script on one host by rerunning a saved command file before rolling it out, because a bug in configure fails every restart across the cluster.

Alerts as code

Each service ships its alert definitions in alerts.json. The simple types, PORT, WEB and METRIC, cover sockets, HTTP endpoints and JMX thresholds. A SCRIPT alert runs Python inside the agent on the alert's interval, which makes it the right tool for business-level checks such as queue depth:

{
  "QUEUEMON": {
    "QUEUEMON_SERVER": [
      {
        "name": "queuemon_queue_depth",
        "label": "QueueMon ingest queue depth",
        "interval": 2,
        "scope": "ANY",
        "enabled": true,
        "source": {
          "type": "SCRIPT",
          "path": "QUEUEMON/package/alerts/alert_queue_depth.py",
          "parameters": [
            {"name": "warn_depth", "display_name": "Warning depth", "value": 50000, "type": "NUMERIC"},
            {"name": "crit_depth", "display_name": "Critical depth", "value": 200000, "type": "NUMERIC"}
          ]
        }
      }
    ]
  }
}
# alert_queue_depth.py -- runs inside the agent on the alert's interval
import json, urllib.request

def get_tokens():
    return ('{{queuemon-env/queuemon_port}}',)      # config values the agent substitutes

def execute(configurations={}, parameters={}, host_name=None):
    port = configurations['{{queuemon-env/queuemon_port}}']
    try:
        with urllib.request.urlopen(f'http://{host_name}:{port}/stats', timeout=5) as r:
            depth = json.load(r)['depth']
    except Exception as e:
        return ('UNKNOWN', [f'stats endpoint unreachable: {e}'])
    if depth >= float(parameters.get('crit_depth', 200000)):
        return ('CRITICAL', [f'queue depth {depth}'])
    if depth >= float(parameters.get('warn_depth', 50000)):
        return ('WARNING', [f'queue depth {depth}'])
    return ('OK', [f'queue depth {depth}'])

get_tokens declares which configuration values the agent must substitute, and execute returns a state and the text for the alert message. Parameters become editable thresholds in the UI. Once a cluster is running, definitions live in the database, so edits made in the UI drift from the JSON in your repository. Export them regularly through the API, keep the export in git and review differences. Notification targets are resources too:

A=https://ambari.example.com:8443/api/v1; C=prod; AUTH="-u automation:$PW -H X-Requested-By:ambari"

# Alert definitions are resources: export them, keep them in git, diff before applying
curl -s $AUTH "$A/clusters/$C/alert_definitions?fields=*" > alert_definitions.json

# Page only on CRITICAL, by e-mail, for every alert group
curl -s $AUTH -X POST "$A/alert_targets" -d '{"AlertTarget": {
  "name": "oncall-email", "notification_type": "EMAIL", "global": true,
  "alert_states": ["CRITICAL"],
  "properties": {"ambari.dispatch.recipients": ["oncall@example.com"]}}}'

# Ask the metrics collector for the last hour of a NameNode metric
curl -s "http://ams-collector.example.com:6188/ws/v1/timeline/metrics?metricNames=jvm.JvmMetrics.MemHeapUsedM&appId=namenode&hostname=nn1.example.com&startTime=$(( $(date +%s) - 3600 ))000&endTime=$(date +%s)000"

The metrics system

Ambari Metrics (AMS) has three parts. A Metrics Monitor on every host collects operating-system metrics. Hadoop sinks inside NameNode, ResourceManager, HBase and other daemons push service metrics. Both send to the Metrics Collector, which stores time series in an HBase instance queried through Apache Phoenix and serves them on port 6188 to Ambari's UI and the bundled Grafana. In embedded mode the collector's HBase writes to local disk on the collector host; in distributed mode it writes to HDFS, which scales further but means losing HDFS also loses metrics. Aggregators roll raw data into minute, hourly and daily tables per host and per cluster, each with its own TTL.

Most AMS trouble is sizing. The collector's HBase heap and region count must grow with the number of hosts and metrics. An undersized collector shows gaps in graphs, slow dashboards and sinks logging write failures, and METRIC alerts that read from it go UNKNOWN. Cut what is collected before adding hardware: whitelisting metric names, shorter raw-data TTLs and fewer per-host aggregates usually fix a struggling collector. Treat AMS as operational telemetry with short retention; export to a long-term system if you need capacity-planning history. If you plan to adopt Ambari 3.1 when it ships, expect the collector and its tuning to be replaced by the new Prometheus-compatible stack, and avoid building heavy custom tooling on the AMS API in the meantime.

Upgrading Ambari itself

Upgrading Ambari is separate from upgrading the stack it manages, and much less risky, because cluster services keep running while the control plane is down. The sequence is: back up, stop server and agents, upgrade packages, run the schema upgrade, start, verify heartbeats. Stack upgrades of HDFS, YARN and the rest are a different procedure; rolling upgrades explains the HDFS side.

# 0. Backups first: database dump, server config, keys, custom stack content
pg_dump -U ambari ambari > ambari-$(date +%F).sql
tar czf ambari-etc-$(date +%F).tgz /etc/ambari-server/conf /var/lib/ambari-server/keys \
    /var/lib/ambari-server/resources/stacks /var/lib/ambari-server/resources/mpacks

# 1. Stop the control plane (cluster services keep running)
ambari-server stop
pdsh -w ^all_hosts 'ambari-agent stop'

# 2. Point the package manager at the new Ambari repository, then upgrade
dnf -y upgrade ambari-server
ambari-server upgrade            # migrates the database schema
pdsh -w ^all_hosts 'dnf -y upgrade ambari-agent'

# 3. Start and verify
ambari-server start
pdsh -w ^all_hosts 'ambari-agent start'
curl -s $AUTH "$A/hosts?fields=Hosts/host_status,Hosts/last_heartbeat_time"

Rehearse on a staging copy restored from the production database dump, because the schema upgrade is where custom services, unusual configuration groups and old leftover data cause trouble. Keep the old packages and the dump until the new version has run for a week. A failed schema upgrade is recovered by reinstalling the old version and restoring the dump, not by fixing tables by hand.

Rebuilding a lost server

Ambari has no built-in server high availability. If its host dies the cluster keeps running, since NameNodes and ResourceManagers do not depend on it, but you lose control, configuration history and alerting until it returns. Recovery needs four things captured in advance: a recent database dump, the server configuration directory, the keys directory (TLS and encryption keys for stored credentials) and the stack resources tree including custom services. Install the same Ambari version on a new host, restore the database and files, start the server, then point every agent at the new hostname in the server section of ambari-agent.ini if the name changed. Agents re-register on heartbeat and the desired state is intact.

If the database is gone and only a blueprint export survives, you can recreate the cluster definition but not configuration history or alert customisations. That is why the nightly dump matters more than the blueprint. Keep the Kerberos admin credentials and keytab procedures in the runbook, since a kerberized cluster's recovery also depends on the KDC; see Hadoop Kerberos. Monitor the Ambari server and the HA pairs it manages, such as NameNode HA, from a system outside Ambari, because a dead server raises no alert about itself.

Failure modes

SymptomCauseFix
Custom service edit has no effectAgents still run cached scriptsRestart agents to resync the cache
Every restart of a service fails at configureScript bug or missing config key in params.pyRerun a saved command file on one host; read errors-N.txt
Service shows stopped while runningstatus reads the wrong PID file or is too slowFix the pid path; keep status cheap
Daemon starts before HDFS during full restartNo role_command_order entryAdd the dependency to role_command_order.json
Graphs have gaps, METRIC alerts UNKNOWNUndersized metrics collector HBaseWhitelist metrics, cut TTLs, raise heap
Server fails to start after upgradeSchema upgrade broke on unusual dataReinstall old version, restore dump, fix in staging

Trade-offs

  • Custom services versus external tooling: managing your daemons in Ambari gives one restart, config and alert model, at the cost of writing and testing Python against an API with sparse documentation.
  • Alerts in Ambari versus a central monitoring system: built-in definitions understand topology and HA, but they die with the server and suit operational checks better than long-term SLOs.
  • Embedded versus distributed metrics: embedded is simple and independent of HDFS; distributed scales further but couples observability to the system being observed.
  • Staying on Ambari versus migrating: 3.x is alive again but small. Every custom service you write is something to port if you move to another manager.

What to do next

  1. Read the command, output and error files of one recent operation to learn what your agents actually run.
  2. Package one in-house daemon as a custom service on a staging cluster, with idempotent scripts and a role order entry.
  3. Export alert definitions and targets through the API into git and review the diff weekly.
  4. Check metrics collector health and trim collected metrics before it falls over.
  5. Schedule nightly database dumps plus the config, keys and stack resources directories, and rehearse a rebuild on a fresh host.
  6. Rehearse the next Ambari upgrade on a staging copy of the production database.
  7. Track the Ambari 3.1 release and plan how a monitoring replacement would affect your METRIC alerts and dashboards.
Key takeaway: Ambari runs every operation by sending agents a JSON command file and invoking a Python command script, so any daemon can be managed like a built-in service by writing metainfo.xml, idempotent resource_management scripts, role ordering and alerts.json. Keep alert definitions exported in version control, size the metrics collector for the cluster, upgrade Ambari itself only after a database backup and staging rehearsal, and capture the database, configuration, keys and stack resources so a lost server can be rebuilt without touching the running cluster.