The ResourceManager is a scheduler, not a historian. It keeps only a bounded list of recently finished applications in memory, and when it restarts or the list rolls over, the record of what ran, for how long, on which containers and with which framework-level detail is gone. The YARN Timeline Server exists to hold that record. It is the store behind the YARN UI's view of completed applications, behind the Tez UI, and behind any framework that wants to publish its own events and query them later.

There are three generations with very different architectures: version 1 is a single server writing to a local LevelDB, version 1.5 moves the write path onto HDFS so that a slow server cannot stall applications, and version 2 distributes collectors across NodeManagers and stores everything in HBase. This article explains each from first principles, shows how an ApplicationMaster publishes data and how you query it, and ends with the operational problems that actually page people: an ever-growing LevelDB, clients blocking on retries, and an active directory full of small files.

Advertisement

What the Timeline Server is for

Two kinds of history matter on a YARN cluster. Generic history is what YARN itself knows: application ID, user, queue, start and finish times, final status, attempts and the containers each attempt ran. Framework-specific history is what only the framework knows: for Tez, the DAG, its vertices, task attempts and counters; for a custom AM, whatever its authors care about. Before the Timeline Server, each framework built its own history server (the MapReduce JobHistory Server is the surviving example), and generic history had no durable home at all.

The Timeline Server provides both in one place. With yarn.resourcemanager.system-metrics-publisher.enabled set, the ResourceManager publishes lifecycle events for every application, attempt and container. With yarn.timeline-service.generic-application-history.enabled, the server answers the generic history API that the YARN web UI and yarn application -status fall back to once the ResourceManager has forgotten an application. Frameworks publish their own entities through a client library. The scheduler side of the picture, which this article deliberately does not repeat, is in the ResourceManager deep dive.

The data model: entities, events, filters and domains

The version 1 data model is small and worth learning exactly, because every query you will write depends on it. An entity is identified by a type and an ID, such as TEZ_DAG_ID and dag_1696_0001_1. It carries a start time, a list of time-stamped events (each with a type and an info map), primary filters (indexed key-value pairs you can query on), other info (unindexed key-value pairs returned with the entity), and related entities (links to other type/ID pairs, for example a DAG to its vertices). Every entity belongs to a domain, a namespace with its own reader and writer ACLs.

The distinction between primary filters and other info is the single most important modelling decision. Only primary filters can appear in a primaryFilter or secondaryFilter query, so anything you want to search by, such as user, queue or a business job name, must be a primary filter. Everything else should be other info, because each primary filter costs extra index writes. A posted entity looks like this:

{
  "entities": [{
    "entitytype": "ETL_JOB",
    "entity": "nightly_orders_2026-10-03",
    "starttime": 1759449600000,
    "domain": "etl",
    "primaryfilters": { "user": ["etl_svc"], "pipeline": ["orders"] },
    "otherinfo": { "inputBytes": 81234567890, "sparkConfHash": "a41f9c" },
    "relatedentities": { "YARN_APPLICATION": ["application_1759400000000_0042"] },
    "events": [
      { "timestamp": 1759449600000, "eventtype": "STARTED", "eventinfo": {} },
      { "timestamp": 1759453200000, "eventtype": "FINISHED",
        "eventinfo": { "status": "SUCCEEDED" } }
    ]
  }]
}
Advertisement

Three generations, three write paths

Version 1 is one daemon, the ApplicationHistoryServer, started with yarn --daemon start timelineserver. Clients send entities over HTTP, and the server writes them to a store class, by default org.apache.hadoop.yarn.server.timeline.LeveldbTimelineStore under yarn.timeline-service.leveldb-timeline-store.path. It serves the REST API on yarn.timeline-service.webapp.address (port 8188 by default) and RPC on port 10200. Retention is governed by yarn.timeline-service.ttl-enable and yarn.timeline-service.ttl-ms, which defaults to 604800000 ms, seven days. The design is simple and the limits follow from it: a single writer, a single local disk, and synchronous writes from every application on the cluster.

Version 1.5 keeps the version 1 read API but changes the write path. With yarn.timeline-service.version set to 1.5 and the store class set to org.apache.hadoop.yarn.server.timeline.EntityGroupFSTimelineStore, clients append entities to files in an HDFS active directory (yarn.timeline-service.entity-group-fs-store.active-dir, conventionally /ats/active/, created by an administrator with sticky, world-writable permissions before the server starts). The server scans that directory, loads small summary entities into a LevelDB summary store, and loads detailed entity groups on demand into a cache when someone queries them. Finished applications are moved to a done directory and cleaned after a retention period. The point is decoupling: a slow or down Timeline Server no longer blocks a running ApplicationMaster, because HDFS is the buffer.

Version 2 is a redesign. Each running application gets a collector, hosted by an auxiliary service on the NodeManager running its AM (timeline_collector in yarn.nodemanager.aux-services), and the ResourceManager runs its own collector for YARN system events. Collectors write to HBase through a pluggable writer class. A separate, stateless TimelineReader daemon serves queries from HBase. Writes scale with the cluster, because there is no central write bottleneck, and reads scale by adding readers.

v1: Timeline ServerHTTP put, local LevelDBv1.5: clients write HDFS/ats/active, server scansv2: per-app collectorsNM aux service, asyncLevelDBsingle node, TTLHDFS + summaryLevelDB + entity cacheHBaseflow / run / app tablesReadersREST /ws/v1 or /ws/v2Each generation moves the write path further from a single daemonv1 blocks writers on one server; v1.5 buffers in HDFS; v2 spreads writers over NodeManagers
Three generations of the Timeline Server. The read API is similar in spirit across them, but where writes land, and therefore what fails under load, is completely different.

Version 2: flows, metrics and aggregation

Version 2 adds structure that version 1 lacked. Applications are grouped into flows: a flow is a logical job, such as a nightly pipeline, a flow run is one execution of it, and each run contains one or more YARN applications, which in turn contain entities such as attempts, containers and custom types. An application declares its flow through application tags of the form TIMELINE_FLOW_NAME_TAG:orders_nightly, TIMELINE_FLOW_VERSION_TAG:v7 and TIMELINE_FLOW_RUN_ID_TAG:1759449600000; without them the flow name defaults to something derived from the application name, which is rarely what you want to aggregate on.

Entities in version 2 carry metrics as time series, not just key-value info, and the system aggregates them: application-level metrics roll up from entities using SUM or MAX, and flow-run and flow-level views are maintained in HBase so that a question such as "how many container-seconds did every run of this pipeline use this week" is a single read rather than a scan of every application. That aggregation is why version 2 needs HBase: it relies on HBase coprocessors and a row-key design that keeps a flow's runs adjacent. Create the schema once with the schema creator before the first write:

# On a host with the HBase client configuration and the timeline-service jars
hadoop org.apache.hadoop.yarn.server.timelineservice.storage.TimelineSchemaCreator -create

# yarn-site.xml (excerpt; shorthand for <property> name/value entries)
yarn.timeline-service.enabled              = true
yarn.timeline-service.version              = 2.0f
yarn.system-metrics-publisher.enabled      = true
yarn.nodemanager.aux-services              = mapreduce_shuffle,timeline_collector
yarn.nodemanager.aux-services.timeline_collector.class =
    org.apache.hadoop.yarn.server.timelineservice.collector.PerNodeTimelineCollectorsAuxService
yarn.timeline-service.reader.webapp.address = reader-host:8188

Check the supported HBase versions in the Timeline Service v.2 page for your exact Hadoop release before planning the deployment; they have changed across 3.x releases, and running against an unsupported HBase major version fails in ways that look like schema corruption.

Publishing from an ApplicationMaster

Publishing is the part developers touch. In version 1 and 1.5 an ApplicationMaster creates a TimelineClient, builds entities and puts them. The response lists rejected entities with an error code, and ignoring it is the most common publishing bug, because a rejected entity looks exactly like a successful one until somebody queries for it.

TimelineClient client = TimelineClient.createTimelineClient();
client.init(conf);
client.start();

TimelineEntity job = new TimelineEntity();
job.setEntityType("ETL_JOB");
job.setEntityId("nightly_orders_2026-10-03");
job.setStartTime(System.currentTimeMillis());
job.addPrimaryFilter("pipeline", "orders");
job.addOtherInfo("inputBytes", 81234567890L);

TimelineEvent started = new TimelineEvent();
started.setEventType("STARTED");
started.setTimestamp(System.currentTimeMillis());
job.addEvent(started);

TimelinePutResponse resp = client.putEntities(job);
for (TimelinePutResponse.TimelinePutError e : resp.getErrors()) {
  LOG.warn("timeline rejected {}/{}: code {}", e.getEntityType(), e.getEntityId(), e.getErrorCode());
}

In version 2 the client must know where its collector lives, which only the ResourceManager can tell it. The AM creates a TimelineV2Client for its application ID and registers it with its AMRMClient; each allocate heartbeat then delivers the collector address. Writes can be synchronous or asynchronous, and asynchronous is the default choice for anything on the hot path.

TimelineV2Client tl = TimelineV2Client.createTimelineClient(appId);
tl.init(conf);
tl.start();
amRMClient.registerTimelineV2Client(tl);   // collector address arrives via allocate

org.apache.hadoop.yarn.api.records.timelineservice.TimelineEntity e =
    new org.apache.hadoop.yarn.api.records.timelineservice.TimelineEntity();
e.setType("ETL_STAGE");
e.setId("load_orders");
e.setCreatedTime(System.currentTimeMillis());

TimelineMetric rows = new TimelineMetric();
rows.setId("ROWS_WRITTEN");
rows.addValue(System.currentTimeMillis(), 1250000L);
e.addMetric(rows);

tl.putEntitiesAsync(e);   // buffered; use putEntities(e) where loss is unacceptable

The container lifecycle on the NodeManager side, which is what hosts the collector, is covered in the NodeManager article, and the AM protocol that delivers the collector address is in the ApplicationMaster article.

Querying history: a worked example

Reading is plain HTTP and JSON. A worked example: a nightly pipeline missed its deadline and you want to know which run slowed down. In version 1 or 1.5 you query by type and primary filter, newest first:

# v1 / v1.5: last 7 runs of the orders pipeline, events only
curl -s "http://ats-host:8188/ws/v1/timeline/ETL_JOB?primaryFilter=pipeline:orders&limit=7&fields=events,otherinfo"

# one entity, everything
curl -s "http://ats-host:8188/ws/v1/timeline/ETL_JOB/nightly_orders_2026-10-03"

# generic history for an application the RM has forgotten
curl -s "http://ats-host:8188/ws/v1/applicationhistory/apps/application_1759400000000_0042"

# v2: runs of a flow, then applications inside one run
curl -s "http://reader-host:8188/ws/v2/timeline/users/etl_svc/flows/orders_nightly/runs"
curl -s "http://reader-host:8188/ws/v2/timeline/clusters/prod/apps/application_1759400000000_0042/entities/ETL_STAGE?fields=METRICS"

Compute the duration of each run from its STARTED and FINISHED events, then compare inputBytes across runs. If duration rose with input size, the job is fine and the deadline is wrong; if duration rose while input stayed flat, open the slow run's related YARN application and compare its container count and queue with a normal run. That two-step query, framework entity first and generic history second, is the pattern the data model was designed for, and it only works if you made the right fields primary filters when you wrote the publisher.

Failure modes and operations

FailureWhat you seeWhat to do
LevelDB grows without bound (v1, v1.5 summary)Disk fills on the Timeline host; reads slow as compaction falls behindEnable TTL, size retention to what anyone actually queries, consider the rolling LevelDB store, monitor the store directory size
Synchronous puts against a slow server (v1)AMs stall on publishing; Tez DAGs take longer while the cluster is idleMove to 1.5 so writes land in HDFS; bound client retries with yarn.timeline-service.client.max-retries (default 30)
Timeline Server down at job submissionClients that fetch a timeline delegation token fail to submit, even if they never publishTreat the Timeline Server as a dependency of submission on secure clusters and alert on it like the RM; yarn.timeline-service.client.best-effort lets submission proceed without the token
Small files in the active directory (v1.5)NameNode object count climbs; directory scans slowKeep done-directory retention short, watch file counts under /ats, avoid one file per tiny entity
HBase hotspotting or unavailability (v2)Collectors buffer, then drop asynchronous writes; reader latency spikesSize HBase for write load, monitor region balance, use synchronous puts only for records you cannot lose
Rejected entities ignoredData silently missing in the UILog or count every TimelinePutError; check domain ACLs and entity size

Two operational habits prevent most incidents. First, measure the store: alert on LevelDB directory size, HDFS object counts under the active and done directories, or HBase write latency, depending on version. Second, decide retention by asking who queries history and how far back, then set the TTL to that and not to "forever"; a Timeline Server that keeps everything becomes the slowest component on the cluster.

Choosing a version

NeedVersion 1Version 1.5Version 2
Small cluster, light publishingFine, simplest to runWorks, more moving partsOverkill
Tez UI on a busy clusterWriters stall on the serverDesigned for this caseCheck your Tez version's support first
Flow-level aggregation and metricsNoNoYes, its main purpose
Operational dependencyLocal diskHDFS plus a local summary storeHBase cluster plus readers
Write scalabilityOne daemonHDFS throughputScales with NodeManagers

The honest default for most on-premises clusters running Tez is version 1.5: it removes the write-path bottleneck without adding an HBase dependency. Choose version 2 when you need flow aggregation or metrics history at scale and already run HBase well. Stay on version 1 only where publishing is light and history is mostly the generic kind. If you are mapping out the whole resource-management layer first, start from the YARN overview.

What to do next

  1. Find out which version you run: read yarn.timeline-service.version and yarn.timeline-service.store-class from the live configuration, not from a template.
  2. Measure the store today (LevelDB directory size, file counts under the HDFS active and done directories, or HBase table sizes) and record the growth rate over a week.
  3. Set TTL and done-directory retention from the longest look-back anyone actually uses, and write that decision down.
  4. Review every publisher you own: make query keys primary filters, keep bulky data in other info, and log every rejected entity.
  5. On secure clusters, add the Timeline Server to the submission-path health checks and alert on it.
  6. If writers stall on a v1 server, plan a move to 1.5; if you need flow-level aggregation, prototype v2 against your existing HBase and run the schema creator in a test environment first.
  7. Run the two-step query from the worked example against a real slow job, so the path from framework entity to generic history is rehearsed before an incident.
Key takeaway: The Timeline Server is YARN's durable history: generic application, attempt and container records published by the ResourceManager, plus framework entities published by ApplicationMasters. Version 1 writes synchronously to one LevelDB server, version 1.5 buffers writes in HDFS so a slow server cannot stall jobs, and version 2 spreads collectors across NodeManagers onto HBase with flows and metric aggregation. Model queries as primary filters, check every put response, set retention deliberately, and monitor whichever store your version depends on.