When a Hive query on Tez is slow, the ResourceManager UI tells you almost nothing: one YARN application, one ApplicationMaster container, a start and an end time. The interesting structure, the DAG of vertices, the hundreds of tasks inside each vertex, the failed attempts and the counters that explain where time went, lives inside the Tez ApplicationMaster. And the AM exits shortly after the query finishes, or idles in a Hive session until it times out. If you want to see what happened yesterday at 02:00, something has to have written that structure down.
That something is the YARN Application Timeline Server (ATS), and the Tez UI is a browser application that reads from it. This article explains the whole pipeline from first principles: what the AM logs, how ATS v1 and v1.5 store it, what the entity model looks like, how to configure it, how to query it without the UI, how to read a slow DAG, and what breaks in production. It assumes you know roughly how Hive turns a query into a Tez DAG; Hive on Tez execution covers that part.
Why DAG history needs its own service
YARN was designed around applications, not around what an application does internally. The RM knows that application_1727840000000_0412 ran in queue etl with 40 containers; it does not know that container 17 ran task 3 of vertex Reducer 2, or that this vertex spilled 60 GB to disk. MapReduce solved this with its own JobHistory Server. Tez, which can run many DAGs in one long-lived session AM, needed a generic store instead, and YARN's Timeline Server was built to be that store for any framework.
The split of responsibilities is worth memorising because every troubleshooting session starts with it:
- Running DAGs: the UI asks the RM for the application, then the AM itself for live progress (through the RM proxy). If the AM is gone, live data is gone.
- Finished DAGs: everything comes from ATS. If the AM never wrote its events, or ATS expired them, the DAG does not exist as far as the UI is concerned.
- Hive query text and plan: written separately by HiveServer2's ATS hook, linked to the DAG by query and DAG identifiers.
The data path, step by step
Inside the AM, every state change produces a history event: DAG submitted, initialised, started and finished; vertex initialised, started and finished; task started and finished; task attempt started and finished; container launched and stopped. These events go to whichever history logging service tez.history.logging.service.class names. The two ATS-backed implementations differ in where the bytes go:
| ATS v1 | ATS v1.5 | |
|---|---|---|
| Tez class | ATSHistoryLoggingService | ATSV15HistoryLoggingService |
| Write path | AM sends entities to the Timeline Server over HTTP | AM appends entity logs to files on HDFS |
| Server store | LeveldbTimelineStore (the default store class) | EntityGroupFSTimelineStore |
| Read path | Server queries its LevelDB | Server scans HDFS, loads entity groups into a cache |
| Main weakness | Timeline Server is in the write path of every DAG | Read latency on first access; HDFS directory hygiene |
Version 1.5 exists because v1 put a single, non-replicated server in the write path of every job on the cluster. When it fell behind, AMs had to buffer events, and busy clusters could see slow DAG completion and gaps in history. In v1.5 the AM writes to an active-dir on HDFS, which scales with the filesystem; the server reads those files in the background and moves finished applications to a done-dir. Tez ships a cache plugin, TimelineCachePluginImpl, that tells the server which group of files holds a requested DAG, so a UI request for one DAG loads one group rather than scanning everything.
The entity model
ATS stores entities: a type, an id, a start time, a list of timestamped events, primary filters (indexed key-value pairs), other info (unindexed key-value pairs) and related entities. Tez maps its object hierarchy onto that model with one entity type per level:
| Entity type | One per | Typical use |
|---|---|---|
TEZ_APPLICATION | YARN application | Tez version, AM configuration |
TEZ_DAG_ID | DAG | Status, duration, DAG plan, DAG-level counters |
TEZ_VERTEX_ID | Vertex | Task counts, timings, vertex counters |
TEZ_TASK_ID | Task | Successful attempt, task duration |
TEZ_TASK_ATTEMPT_ID | Attempt | Container, node, diagnostics, attempt counters |
HIVE_QUERY_ID | Hive query (from ATSHook) | Query text, user, explain plan |
Child entities carry their parent ids as primary filters, which is what makes questions like "all vertices of this DAG" cheap. That design also explains a cost: a DAG with 5,000 tasks and some retries produces well over 10,000 entities, each with counters. History volume grows with task count, not with query count, so one badly split table can produce more timeline data than a thousand small queries.
Configuration that actually works
The Tez UI documentation lists the minimum: enable the timeline service, enable CORS so a browser page on another origin can call it, publish YARN's own application metrics, and tell Tez to log to ATS. For a v1.5 setup, on a Hadoop release that includes it, the pieces look like this:
<!-- yarn-site.xml -->
<property><name>yarn.timeline-service.enabled</name><value>true</value></property>
<property><name>yarn.timeline-service.version</name><value>1.5</value></property>
<property><name>yarn.timeline-service.hostname</name><value>ats01.example.com</value></property>
<property><name>yarn.timeline-service.http-cross-origin.enabled</name><value>true</value></property>
<property><name>yarn.resourcemanager.system-metrics-publisher.enabled</name><value>true</value></property>
<property><name>yarn.timeline-service.store-class</name>
<value>org.apache.hadoop.yarn.server.timeline.EntityGroupFSTimelineStore</value></property>
<property><name>yarn.timeline-service.entity-group-fs-store.active-dir</name><value>/ats/active/</value></property>
<property><name>yarn.timeline-service.entity-group-fs-store.done-dir</name><value>/ats/done/</value></property>
<property><name>yarn.timeline-service.entity-group-fs-store.group-id-plugin-classes</name>
<value>org.apache.tez.dag.history.logging.ats.TimelineCachePluginImpl</value></property>
<property><name>yarn.timeline-service.leveldb-timeline-store.path</name><value>/data1/ats/leveldb/</value></property>
<!-- tez-site.xml -->
<property><name>tez.history.logging.service.class</name>
<value>org.apache.tez.dag.history.logging.ats.ATSV15HistoryLoggingService</value></property>
<property><name>tez.tez-ui.history-url.base</name><value>http://web01.example.com:9999/tez-ui/</value></property>Three details catch people. First, the Tez plugin class must be on the Timeline Server's classpath, not only on the clients', or v1.5 reads fail. Second, the active directory must be writable by every user who submits Tez jobs, because each AM writes its own files there as that user. Third, the UI itself is static files served by any web server; you point it at the timeline and RM addresses in its configuration file, and the browser, not the web server, makes the REST calls. That is why CORS must be on.
For Hive query entities, add HiveServer2's hook to the pre-, post- and failure-hook lists in hive-site.xml:
hive.exec.pre.hooks=org.apache.hadoop.hive.ql.hooks.ATSHook
hive.exec.post.hooks=org.apache.hadoop.hive.ql.hooks.ATSHook
hive.exec.failure.hooks=org.apache.hadoop.hive.ql.hooks.ATSHookAppend to existing values rather than replacing them; other hooks, such as lineage or audit hooks, often already live in these lists.
Querying ATS without the UI
Everything the UI shows is available from the v1 REST API at /ws/v1/timeline/{entityType}, with query parameters limit (default 100), primaryFilter (indexed), secondaryFilters (scans), windowStart, windowEnd and fields. This matters for automation: a nightly job can find every DAG that ran longer than its SLA without anyone opening a browser.
# The 20 most recent DAGs
curl -s "http://ats01:8188/ws/v1/timeline/TEZ_DAG_ID?limit=20&fields=primaryfilters,otherinfo"
# One DAG with its events
curl -s "http://ats01:8188/ws/v1/timeline/TEZ_DAG_ID/dag_1727840000000_0412_1"
# All vertices of that DAG, using the parent id as a primary filter
curl -s "http://ats01:8188/ws/v1/timeline/TEZ_VERTEX_ID?primaryFilter=TEZ_DAG_ID:dag_1727840000000_0412_1&limit=1000"The exact keys Tez puts in otherinfo vary across Tez versions, so inspect one vertex entity before you write a parser. The script below assumes start and end timestamps are present under the names shown and degrades gracefully if they are not:
import json, sys, urllib.request
ATS = "http://ats01:8188/ws/v1/timeline"
def get(path):
with urllib.request.urlopen(f"{ATS}/{path}", timeout=30) as r:
return json.load(r)
def slowest_vertices(dag_id, top=5):
data = get(f"TEZ_VERTEX_ID?primaryFilter=TEZ_DAG_ID:{dag_id}&limit=1000")
rows = []
for e in data.get("entities", []):
info = e.get("otherinfo", {})
start, end = info.get("startTime"), info.get("endTime")
if start and end:
name = info.get("vertexName", e["entity"])
rows.append((end - start, name, info.get("numTasks")))
if not rows:
sys.exit("no timing fields found; print one entity and adjust the key names")
for ms, name, tasks in sorted(rows, reverse=True)[:top]:
print(f"{name:<20} {ms/1000:8.1f}s tasks={tasks}")
slowest_vertices(sys.argv[1])Two warnings. Unbounded queries against a busy v1 server are expensive, because secondaryFilters scan; prefer primary filters and time windows. And ATS ACLs apply: with YARN ACLs enabled, an entity is visible only to its owner and to users granted view access, so a service account running reports needs that access explicitly (Tez exposes tez.am.view-acls for DAG data).
Reading a slow DAG: a worked example
Suppose a nightly Hive INSERT that usually takes 9 minutes took 41. Open the DAG in the Tez UI and work from the outside in.
- DAG details. Status SUCCEEDED, 41 minutes, no failed attempts that ended the job but 12 killed attempts. Note the queue and the AM host.
- Graphical view. Map 1 and Map 3 feed Reducer 2, which feeds Reducer 4 and the file sink. The vertex timeline shows Map 1 and Map 3 finished within 6 minutes. Reducer 2 ran for 33.
- Vertex swimlane. Reducer 2 has 200 tasks. 199 finish within 90 seconds; one runs for 31 minutes. That shape, a long single tail, is skew, not lack of capacity. If instead tasks started in waves with gaps between them, the problem would be containers, meaning queue capacity or preemption, which you would confirm in the ResourceManager.
- Task counters. Compare the slow task with a typical one. The slow task's
REDUCE_INPUT_RECORDSis 400 times the median, and itsADDITIONAL_SPILLS_BYTES_WRITTENis tens of GB while others show zero. One join key dominates, and its records overflowed the sort buffer and spilled repeatedly. - Hive query entity. The query joins events to accounts on account_id; a recent change started writing a placeholder account id for anonymous events.
The fix is in the data or the query: filter or salt the placeholder key, or enable Hive's skew-join handling for that join. Raising container memory would only make the one slow task spill a little less. The UI did not tell you the fix, but it narrowed 41 minutes of mystery to one key in one vertex in about five clicks. As for the killed attempts: if speculative execution is enabled, they are usually duplicates launched against the straggler, a symptom of skew rather than a separate problem; if it is not, check the RM for preemption.
Failure modes
- Blank UI or endless spinner: CORS disabled, the browser cannot resolve the ATS hostname, or mixed HTTP and HTTPS. Open the browser developer console before touching the server.
- DAG missing from the list: the job ran with a different
tez.history.logging.service.class(a per-session override, or the simple file logger), the AM died before flushing, or the TTL expired it. - A script sees only 100 entities: the REST API's default
limitapplies unless the caller passes a larger one; it is not data loss. - v1 Timeline Server overload: slow DAG completion across the cluster, AM logs complaining about posting timeline events, a fast-growing LevelDB directory.
- v1.5 permission errors: AMs fail to create files in the active directory, and history silently stops for some users only.
- Disk exhaustion: TTL disabled or very long on a busy cluster; LevelDB compaction then makes the problem worse before it gets better.
- Live data gone after completion: expected behaviour; if the finished DAG also lacks detail, the history service, not the UI, is the problem.
Operating the Timeline Server
The Hadoop documentation is blunt about v1's architecture: the single-server implementation limits the scalability of the service and prevents it from being a highly available part of YARN. Plan around that rather than against it.
- Retention.
yarn.timeline-service.ttl-enabledefaults to true andyarn.timeline-service.ttl-msto 604800000, seven days. Shorten it before you add disks; most teams need a week of DAG history, and anything older belongs in a summarised warehouse table. - Storage. Put the LevelDB path on a dedicated local disk, not the OS disk; with v1.5 the server still keeps local LevelDB state there.
- Heap. The v1.5 reader caches loaded entity groups; large DAGs inflate heap use. Watch GC on the Timeline Server the way you would on any JVM service.
- Monitoring. Alert on the Timeline Server process, its REST latency and the size of the active directory, which should drain as applications finish. Feed DAG durations into the dashboards described in Hive monitoring.
- Export what matters. A daily job that pulls DAG and vertex durations through REST into a Hive table gives you months of trend data at a tiny fraction of ATS storage.
Know the version boundary too. YARN also has Timeline Service v2, a distributed design backed by HBase, and the Tez documentation describes only the v1 and v1.5 setups covered here. If your platform moves to v2, confirm the Tez UI works against it in a test cluster before you rely on it.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| ATS v1 | Simple, one process, immediate reads | Server in every DAG's write path; scales poorly |
| ATS v1.5 | Writes scale with HDFS; server failure does not block DAGs | More moving parts; first read of a DAG is slower |
| Long TTL | Older DAGs inspectable | Disk, compaction time, slower queries |
| Hive ATSHook on | Query text and plan next to DAG data | Extra entities per query; another writer to ATS |
| REST export to a table | Long-term trends, SQL over history | A pipeline to maintain; counters need flattening |
What to do next
- Check
tez.history.logging.service.classon a real HiveServer2 session, not only in tez-site.xml, and confirm a test query's DAG appears in the UI within a minute. - If you run ATS v1 on a busy cluster, plan the move to v1.5 with the plugin class on the server classpath and an active directory every submitting user can write.
- Set the TTL deliberately and alert on Timeline Server disk, heap and REST latency.
- Run the vertex script above against yesterday's slowest DAG and practise the swimlane-then-counters reading on it.
- Add the Hive ATS hook if query text is missing from the UI, appending to existing hook lists.
- Build a small daily export of DAG and vertex durations into a table so trends survive the TTL, and review the LLAP page if your interactive queries run there instead.