A Hadoop cluster of any real size runs hundreds of processes: NameNodes, DataNodes, ResourceManagers, NodeManagers, Hive and Impala daemons, ZooKeeper, Kafka brokers, and more. Each one has its own configuration files, its own logs, its own JVM settings and its own place in a startup order. Cloudera Manager is the control plane that does this for Cloudera's distributions. It holds the desired configuration in a database, renders the configuration files, starts and stops every process through an agent on each host, monitors health, and exposes all of it through a web UI and a REST API.

This article explains how the pieces fit: the server and its database, the agent and its process directories, parcels, the configuration model and stale configuration, the Management Service roles, and the API. A worked example then changes a DataNode setting across a cluster and rolls it out with a rolling restart. Facts come from the Cloudera Manager 7.13 documentation as read on 2026-10-03. Ambari, the open-source equivalent with a similar desired-state model, is covered in Apache Ambari.

Architecture at a glance

Cloudera Manager Serverdesired state, commands, UI, APIDatabaseconfigs, history, usersAdmin / automationbrowser, REST APIHTTPSManagement ServiceHost/Service Monitor, AlertsHost Acloudera-scm-agentcloudera-scm-supervisordNameNode, JournalNode/var/run/cloudera-scm-agent/processHost Bcloudera-scm-agentcloudera-scm-supervisordDataNode, NodeManager/opt/cloudera/parcelsHost C ... Ncloudera-scm-agentDataNode, NodeManagerheartbeat 15 sAgents pull desired state on every heartbeat; supervisord starts and restarts the actual processes.
The server stores desired state in its database; agents on every host heartbeat in, receive what should run, and hand processes to supervisord. The Management Service collects metrics and evaluates health.

What Cloudera Manager does and does not do

Cloudera Manager manages the lifecycle of clusters built from Cloudera's runtime: installing software, configuring services, starting and stopping roles, rolling restarts and upgrades, enabling Kerberos and TLS, monitoring health and metrics, raising alerts, and keeping an audit trail of who changed what. It is a desired-state system. You change the configuration the server stores, and the server works out which processes must be reconfigured and restarted to match.

It is not a scheduler, a data catalogue or a security policy engine. YARN schedules work, Atlas catalogues data, and Ranger enforces access policies, and Cloudera Manager installs, configures and monitors all three. It will also happily apply a bad setting to every host at once.

Server, database, agents and the Management Service

There are four parts. The server is a Java application that hosts the web UI and REST API, stores the desired state, plans commands such as restart or deploy, and receives heartbeats from agents. The database, usually PostgreSQL, MySQL, MariaDB or Oracle in production, holds every configuration value, role assignment, command record and user. It is the most important thing to back up: lose it and you lose the cluster's definition, even though the cluster itself keeps running. The agent runs on every managed host, as root, and makes the host match the server's plan. The Management Service is a set of roles that collect metrics and events, evaluate health and send alerts.

The data flow is pull-based. The agent sends a heartbeat to the server every 15 seconds by default, reporting what is running on its host. The reply carries what should be running and with which configuration. If the two differ, the agent acts. When heartbeats stop, the server marks the host Concerning after 5 missed heartbeats and Bad after 10, which is about 75 seconds and 150 seconds at the default interval. A server outage does not stop running services; you lose management and monitoring until it returns.

The agent, supervisord and the process directory

The agent does not run Hadoop daemons itself. It delegates to a modified copy of the open-source process manager supervisord, which runs as cloudera-scm-supervisord and listens locally on port 19001 by default. Supervisord starts processes, sets their effective user, redirects their logs, notices when they die and, if configured, restarts them.

When a heartbeat reply contains a process the host should run, the agent creates a new directory for it under /var/run/cloudera-scm-agent/process, named with an increasing number and the role, for example 1543-hdfs-DATANODE. It unpacks the generated configuration files, keytabs and scripts into that directory, then asks supervisord to start the process from it. A new restart creates a new numbered directory, so the most recent directory for a role holds the configuration the process is actually running with.

The process directories sit on a tmpfs called cm_processes, so secrets such as keytabs never touch a persistent disk. The tmpfs is created on first agent start, reused on a normal start, remounted on a clean restart and gone after a reboot; it can grow to half of physical RAM but only uses what it needs. The agent's own settings live in /etc/cloudera-scm-agent/config.ini, including server_host, server_port, parcel_dir (default /opt/cloudera/parcels) and its log file under /var/log/cloudera-scm-agent.

# Which configuration is the running DataNode really using?
ls -t /var/run/cloudera-scm-agent/process | grep DATANODE | head -1
grep -A1 'dfs.datanode.handler.count' \
  /var/run/cloudera-scm-agent/process/1543-hdfs-DATANODE/hdfs-site.xml

# Agent health and its view of the server
systemctl status cloudera-scm-agent
tail -n 50 /var/log/cloudera-scm-agent/cloudera-scm-agent.log

Parcels: download, distribute, activate

Cloudera Manager can install software from operating-system packages, but the normal path is parcels. A parcel is a single compressed archive containing a whole product, such as the Cloudera Runtime, plus metadata describing its contents. The server manages parcels through three steps, each visible in the UI and the API. Download fetches the parcel from a repository into the server's local parcel repository. Distribute copies it to every host and unpacks it under the parcel directory. Activate switches a stable symbolic link, such as /opt/cloudera/parcels/CDH, to point at the new version, after which roles pick up the new binaries when they restart.

This split is why parcels make upgrades safer than packages. Distribution, the slow part, happens while services keep running. Activation is a symlink change, and roles move to the new version only when restarted, which can be a rolling restart. Several versions can sit side by side on disk, so rolling back means activating the old one again, subject to the limits of each component's on-disk format; an HDFS upgrade that has been finalised cannot be undone by a symlink.

The configuration model and stale configuration

The configuration model has three levels. A service, such as HDFS, holds service-wide settings. Each role type within it, such as DATANODE, has one or more role config groups; the default one is named like hdfs-DATANODE-BASE, and you add groups for hosts with different hardware, for example DataNodes with twelve disks rather than six. Each role instance belongs to exactly one group and can override individual values. Most settings have their own Cloudera Manager name, which maps to a property in a generated file. The full API view reports that file property as the setting's related name, so you can look a setting up by its Hadoop property name rather than guess its Cloudera Manager key. For properties Cloudera Manager does not model, each file has an advanced configuration snippet, often called a safety valve, where you paste raw XML or key-value lines.

Changing a value does not touch running processes. Instead, affected roles are marked as having stale configuration, and the UI shows a restart or refresh indicator. The API exposes this as each role's configStalenessStatus: FRESH, STALE_REFRESHABLE where a refresh command can apply the change without a restart, or STALE where a restart is needed. Client configuration is separate: the files that command-line tools and applications read under /etc/hadoop/conf and similar paths are updated only when you run Deploy Client Configuration. Forgetting this step is a classic cause of servers running new settings while clients still use old ones.

Monitoring roles and health tests

The Management Service is itself a set of roles managed like any other service. The Host Monitor collects host metrics such as CPU, memory, disk and network, and the Service Monitor collects service and role metrics and runs health tests against them. Both keep their time-series data on local disk in LevelDB, so size their storage deliberately and place them on a host that is not already busy. The Event Server stores and indexes events such as health changes and log alerts. The Alert Publisher sends alerts by email or SNMP. The Reports Manager produces reports such as HDFS disk usage by user and directory.

Health tests turn metrics into a status of Good, Concerning or Bad, with thresholds you can tune per role. Out of the box, many thresholds suit small clusters. On a large cluster, tune the noisy ones, such as DataNode free-space and GC-duration tests, before people learn to ignore the alert stream.

The REST API

Everything in the UI is available through a versioned REST API under /api/v<N>. The version depends on the server release: in the 7.13 line, API v57 needs Cloudera Manager 7.13.1 and v58 needs 7.13.2. Do not hard-code the number; ask the server, then build paths from the answer. Long-running actions return a command object with an ID that you poll until it finishes.

CM="https://cm.example.com:$CM_TLS_PORT"   # your server's HTTPS address
curl -s -u admin:"$PW" "$CM/api/version"                       # e.g. v58
V=$(curl -s -u admin:"$PW" "$CM/api/version")
curl -s -u admin:"$PW" "$CM/api/$V/clusters" | jq '.items[].name'
curl -s -u admin:"$PW" "$CM/api/$V/clusters/prod/services/hdfs/roles" \
  | jq -r '.items[] | [.name, .type, .configStalenessStatus, .healthSummary] | @tsv'

Two habits keep API automation safe. First, export the full deployment with GET /api/<version>/cm/deployment before any bulk change, so you have a record of the previous values. Second, make scripts read the current value and skip the change when it is already applied, so a rerun after a failure does nothing harmful.

Worked example: a DataNode setting with a rolling restart

Here is a worked example. A 40-node cluster shows DataNode RPC queues backing up during a nightly ingest, and the team decides to raise the DataNode handler count from its current value to 30. They want to do it with no HDFS downtime, a few DataNodes at a time, and stop if anything goes wrong. The script finds the setting by its related Hadoop property name rather than by guessing the Cloudera Manager key, sets it on the base DataNode group, then starts a rolling restart of only the stale DataNodes.

import os, time, requests

CM, AUTH = os.environ["CM_URL"], ("admin", os.environ["CM_PASSWORD"])
s = requests.Session(); s.auth = AUTH; s.verify = "/etc/pki/cm-ca.pem"
V = s.get(f"{CM}/api/version").text.strip().strip('"')
BASE = f"{CM}/api/{V}/clusters/prod/services/hdfs"
GROUP = f"{BASE}/roleConfigGroups/hdfs-DATANODE-BASE/config"

def wait(cmd):
    while cmd["active"]:
        time.sleep(10)
        cmd = s.get(f"{CM}/api/{V}/commands/{cmd['id']}").json()
    if not cmd.get("success"):
        raise RuntimeError(cmd.get("resultMessage"))
    return cmd

# 1. Find the CM key for dfs.datanode.handler.count via the full view
full = s.get(GROUP, params={"view": "full"}).json()["items"]
item = next(i for i in full if i.get("relatedName") == "dfs.datanode.handler.count")
current = item.get("value", item.get("default"))
print("key", item["name"], "current", current)

# 2. Idempotent update
if str(current) != "30":
    s.put(GROUP, json={"items": [{"name": item["name"], "value": "30"}]}).raise_for_status()

# 3. Rolling restart: stale DataNodes only, 2 at a time, stop after 1 failure
cmd = s.post(f"{BASE}/commands/rollingRestart", json={
    "restartRoleTypes": ["DATANODE"],
    "staleConfigsOnly": True,
    "slaveBatchSize": 2,
    "slaveFailCountThreshold": 1,
    "sleepSeconds": 60,
})
cmd.raise_for_status()
wait(cmd.json())

During the run, each batch of two DataNodes stops, comes back with a new numbered process directory and re-registers with the NameNode before the next batch begins. With replication factor 3 and rack awareness configured, losing two DataNodes briefly leaves every block readable. The 60-second sleep gives block reports time to land. Once it finishes, the team confirms that no DataNode role is still STALE and checks the new value in the newest process directory on one host. Rolling restart is a licensed feature in some editions, and on a cluster where it is unavailable the same script can restart one role at a time through the role-level restart command.

Failure modes

  • Database lost or corrupted. Services keep running, but you cannot manage them. Back up the database on a schedule and test a restore.
  • Server and process disagree. Someone edited a file under the process directory by hand; the next restart silently reverts it. Make every change through Cloudera Manager or a safety valve.
  • Stale client configuration. Servers run new settings while clients use old ones, for example a wrong NameNode address after enabling HA. Run Deploy Client Configuration after relevant changes.
  • Parcel distribution half-done. A full parcel disk on a few hosts leaves them on the old version. Check free space before distributing and confirm activation on every host.
  • Monitoring storage exhausted. Host or Service Monitor LevelDB stores fill their disk, and charts and health tests go blank. Size and monitor them like any data store.
  • Alert fatigue. Default thresholds on a large cluster produce constant Concerning states, and people stop reading them. Tune the noisy tests early.
  • Everyone is an admin. One shared administrator account means no useful audit trail. Use LDAP or SAML with role-based user roles.

Trade-offs

Cloudera Manager trades flexibility for consistency. Its strength is that one system knows the whole cluster: dependencies between services, safe restart orders, Kerberos principals, health and history. The cost is that you are tied to its model and its release cycle, and configuration it does not model ends up in safety valves that are easy to forget. Teams that run everything else through Terraform or Ansible often drive Cloudera Manager through its API from those tools, or use cluster templates, rather than clicking through the UI, so changes are reviewed and repeatable. For related operations, see rolling upgrades, Kerberos on Hadoop and Hadoop monitoring.

What to do next

  1. Back up the Cloudera Manager database today, then restore it into a scratch instance to prove the backup works.
  2. Export /api/<version>/cm/deployment to version control on a schedule, so every configuration change has a diff.
  3. On one host, find the newest process directory for each role and compare a few values with the UI.
  4. List every safety valve in use and write down why it exists; move values into modelled settings where possible.
  5. Check free space under the parcel directory on every host, enough for two full versions.
  6. Review Host Monitor and Service Monitor storage size and retention, and alert on their disk usage.
  7. Tune the five noisiest health tests, then script one routine change end to end through the API, including a rolling restart with a failure threshold.
Key takeaway: Cloudera Manager is a desired-state control plane: the server stores configuration in a database, agents heartbeat every 15 seconds and hand processes to supervisord, and every restart gets a fresh process directory holding the configuration actually in use. Parcels separate slow distribution from fast activation, configuration changes only mark roles stale until you restart or refresh, and client configuration needs its own deploy. Back up the database, drive changes through the versioned API, and roll them out in small batches with a failure threshold.