The Spark UI is the first place to look when a job is slow, fails, or costs more than it should. It is also widely misread. Engineers stare at the Jobs tab, see a progress bar that is not moving, and restart the job with more executors, when the Stages tab would have shown one task out of two thousand reading forty times more data than the others. More hardware does not fix skew.

This article explains the UI as a system and as a diagnostic tool. It starts with where the numbers come from, because that explains why jobs sometimes vanish from the UI and why its numbers are occasionally wrong. It then walks through each tab with the questions it answers, works through a skew diagnosis end to end, covers event logs and the History Server for jobs that have already finished, and shows how to automate the same checks through the REST API. Configuration names and defaults were checked against the Spark 4.x documentation.

Advertisement

How the UI gets its numbers

Every Spark application's driver runs a web server, by default on port 4040 (spark.ui.port; if the port is taken Spark tries the next one, so a second application on the same host appears on 4041). The pages are not computed by querying executors. Instead, as the scheduler runs, it posts events such as job start, stage submitted, task end and executor added onto the listener bus, an asynchronous queue inside the driver. Executors contribute through task-completion results and periodic heartbeats that carry accumulator and metric updates.

Two consumers matter. A status listener turns events into an in-memory store of jobs, stages, tasks, executors and SQL executions, and the UI pages and REST API read from that store. An optional event-logging listener writes every event as a line of JSON to a file on shared storage, which the History Server can replay after the application ends into exactly the same store and the same pages.

Where Spark UI numbers come from: one event stream, two consumersExecutorstask end + metricsDriver schedulerDAGScheduler, TaskSchedulerheartbeatsListener busasync queues, bounded (10,000 default)eventsStatus listenerbuilds job, stage, task viewsEvent loggerJSON event log filesIn-memory storetrimmed by retained* limitsShared storageHDFS, S3, GCSLive UI on port 4040 + /api/v1exists only while the driver runsHistory Server on port 18080replays logs into a disk or hybrid storereplaySame pages + /api/v1 after the job endsonly if event logging was on
One event stream feeds both the live UI and the event log. The History Server rebuilds the same store from the log, which is why it shows the same pages, and why it shows nothing if event logging was off.

Two consequences follow. First, the live UI dies with the driver: when the application ends, port 4040 goes away. If you did not enable event logging, there is no record to look at afterwards. Second, the UI is only as complete as the event stream. The listener bus queues are bounded, 10,000 events by default (spark.scheduler.listenerbus.eventqueue.capacity). If a listener falls behind, for example on a job with hundreds of thousands of short tasks or with a slow custom listener, events are dropped and the driver log records a warning about dropped events. The UI then shows stages that never finish or task counts that do not add up. Fewer, larger tasks is usually a better fix than a bigger queue.

Retention limits: why your job disappeared

The in-memory store is trimmed so a long-running driver does not run out of heap. The limits, with their Spark 4.x defaults, are spark.ui.retainedJobs (1,000), spark.ui.retainedStages (1,000), spark.ui.retainedTasks (100,000 tasks per stage), spark.sql.ui.retainedExecutions (1,000) and spark.ui.retainedDeadExecutors (100). A streaming query or a notebook that has been running for days will have evicted its early jobs, so the job you are looking for may simply be gone. Raise the limits only with a matching increase in driver memory, and prefer the event log and History Server for investigations that need history.

Advertisement

Jobs tab: the timeline, not the progress bar

A job is created by each action such as count, write or collect. The Jobs tab lists jobs with their stages and a progress bar, which is the least useful thing on the page. Open the event timeline instead. It shows when executors were added and removed and when each job ran. Gaps between jobs where nothing is running mean the time is being spent on the driver: query planning, listing files on object storage, Python code between actions, or a collect pulling results back. Adding executors will not shorten those gaps.

The job detail page shows the DAG of stages. Stage boundaries fall at shuffles, so the number of stages tells you how many exchanges the job needs. The mechanics of jobs, stages and tasks are covered in Spark stages and tasks; here the question is simply which stage holds the time.

Stages tab: where most diagnoses happen

Sort stages by duration and open the slowest. The stage page has three parts worth reading in order. The summary metrics table shows the minimum, 25th percentile, median, 75th percentile and maximum of each task metric: duration, GC time, input size, shuffle read size and records, shuffle write, and spill. Compare the median with the maximum. When they are close, the stage is uniformly slow and needs more parallelism or cheaper work per row. When the maximum is many times the median, the stage has skew or a straggler.

The aggregated metrics by executor table separates data problems from machine problems. If the slow tasks all ran on one executor, suspect that host: a bad disk, a noisy neighbour or a failing node. If they are spread across executors but always process the most data, the problem is the data. The task table lists individual tasks; sort by duration, then look at the slow task's input or shuffle read size, its locality level, its GC time and its Shuffle Read Blocked Time, which measures how long it waited for remote shuffle blocks. High blocked time with low CPU points at the shuffle fetch path rather than at your code; see Spark shuffle architecture.

Spill columns deserve attention. Spill (Memory) is the deserialized size of data that was spilled, Spill (Disk) the serialized size written. Spill on one task is a skew symptom; spill on most tasks means partitions are too large for the execution memory available, which unified memory management explains in detail.

A worked example: diagnosing a slow join

A nightly join of 2 billion order events with a customer table usually takes 20 minutes; tonight it has run for 70. The Jobs tab shows one running job; its DAG has four stages, and stage 7, the join, has 1,999 of 2,000 tasks complete. The executors page shows most executors idle.

In stage 7's summary metrics, median task duration is 9 seconds and median shuffle read is 110 MB, but the maximum task has run for 58 minutes and has read 31 GB. The aggregated-by-executor table shows the slow task on an ordinary executor with no other failures, so this is not a host problem. The task has also spilled heavily. This is key skew: one join key, perhaps a default customer ID used for guest checkouts, carries a large share of the rows, and every row for that key lands in the same partition.

The SQL tab confirms it: the join is a sort-merge join. With adaptive query execution on, Spark can split skewed partitions automatically when a partition exceeds both a size threshold and a multiple of the median; if the final plan in the SQL tab shows no skew handling, check that AQE and its skew-join setting are enabled and that the thresholds suit this data, as described in Spark AQE. If the hot key is a placeholder value, filtering it out and handling it separately is often simpler and faster than any tuning.

SymptomWhere to lookWhat confirms itTypical fix
One task takes far longer than the restStages tab, summary metricsmax shuffle read or input size many times the medianAQE skew join, salting, better key
Spill to disk on many tasksStages tab, Spill (Memory) and Spill (Disk)spill on most tasks, not onemore partitions, more executor memory
Tasks slow but CPU idleStages tab, Shuffle Read Blocked Timelarge blocked timefetch bottleneck: network, shuffle service, disks
High GC timeExecutors tab and task tableGC time a large share of task timesmaller partitions, fewer cached objects, memory tuning
Plan did not change as expectedSQL / DataFrame tabfinal plan shows the same join or no pruningcheck statistics, AQE, filters
Executors lostExecutors tab, dead executorsexit reasons, container killsmemory overhead, preemption, OOM
Long gaps between jobsJobs tab event timelineno stages runningdriver-side work: collect, planning, Python

SQL / DataFrame tab: the plan with numbers on it

For DataFrame and SQL workloads, the SQL tab is often more useful than the Stages tab because it maps runtime metrics onto the physical plan. Each query execution shows a graph of operators, with per-operator metrics such as number of output rows, data size, time spent, spill size and, for scans, files and partitions read. You can see directly whether a filter was pushed into the scan, whether partition pruning removed partitions, and which join strategy was chosen.

With AQE, the plan can change while the query runs. The page shows the final plan once the query finishes; look for AQE shuffle reads, which report coalesced or split partitions, and for joins converted to broadcast at runtime. Row counts are the key check: an operator whose output is far larger than its input usually means an exploding join from duplicate keys. To read the operator names themselves, see Spark EXPLAIN plans.

Executors, Storage and Environment

The Executors tab shows, per executor, active and failed tasks, total task time, GC time, input, shuffle read and write, and storage memory used. Uneven task time across executors suggests skew or locality problems; a high share of GC time points at memory pressure; failed tasks concentrated on one executor point at a host. For live applications, the thread dump link catches tasks blocked on external calls. The tab also lists dead executors and their removal reasons, which is where container kills for exceeding memory overhead show up.

The Storage tab lists cached RDDs and DataFrames with their storage level, fraction cached and size in memory and on disk. A dataset that is only partly cached is being recomputed for the missing partitions every time it is used. The Environment tab shows the effective Spark configuration, JVM and classpath, and it is the authoritative answer to whether a setting you passed actually took effect. Values that match spark.redaction.regex, such as keys containing secret, password or token, are redacted.

Event logs and the History Server

For anything you need to investigate after the fact, enable event logging on every application: set spark.eventLog.enabled=true and point spark.eventLog.dir at durable shared storage. In Spark 4.x, rolling event logs and compression are on by default (spark.eventLog.rolling.enabled, 128 MB files by spark.eventLog.rolling.maxFileSize, and spark.eventLog.compress); check the defaults of your own version and set them explicitly if you run an older release. Rolling logs matter for long-running and streaming applications, whose single log file would otherwise grow without limit.

The History Server, on port 18080 by default, reads spark.history.fs.logDirectory and replays each application into a local store. Give it a disk store with spark.history.store.path so it does not re-parse every log on restart; the hybrid store (spark.history.store.hybridStore.enabled) builds the view in memory and then writes it to disk, which speeds up the first load of large applications. Enable its cleaner so old logs are removed, and size its heap for the largest application you expect to open.

# spark-defaults.conf for every application
spark.eventLog.enabled            true
spark.eventLog.dir                gs://analytics-spark-logs/events
spark.eventLog.rolling.enabled    true
spark.eventLog.rolling.maxFileSize 128m

# History Server
spark.history.fs.logDirectory     gs://analytics-spark-logs/events
spark.history.store.path          /var/lib/spark-history
spark.history.store.hybridStore.enabled true
spark.history.fs.cleaner.enabled  true

Automating checks with the REST API

Everything the UI shows is also available as JSON under /api/v1, from the live driver and from the History Server. That makes it possible to check every nightly job for skew or spill automatically instead of waiting for someone to open the UI. The script below lists completed stages of an application, fetches each stage's task-summary quantiles and flags stages whose slowest task took more than ten times the median.

import requests

BASE = "http://history-server:18080/api/v1"

def skewed_stages(app_id, ratio=10.0, min_seconds=60):
    stages = requests.get(f"{BASE}/applications/{app_id}/stages",
                          params={"status": "complete"}, timeout=30).json()
    for s in stages:
        url = (f"{BASE}/applications/{app_id}/stages/"
               f"{s['stageId']}/{s['attemptId']}/taskSummary")
        q = requests.get(url, params={"quantiles": "0.5,1.0"}, timeout=30).json()
        median_ms, max_ms = q["executorRunTime"]
        if max_ms / 1000 >= min_seconds and max_ms >= ratio * max(median_ms, 1):
            yield s["stageId"], s["name"], median_ms, max_ms

for stage_id, name, med, mx in skewed_stages("application_1727650000000_0042"):
    print(f"stage {stage_id} {name[:60]}: median {med/1000:.1f}s, max {mx/1000:.1f}s")

The same API lists executors, SQL executions with their plan descriptions, and environment settings, so a small nightly report can catch regressions in shuffle size, spill or duration before they become incidents.

Security and operational settings

The UI exposes the full configuration, SQL text and the ability to kill jobs. On shared clusters, put it behind authentication using servlet filters (spark.ui.filters) or the cluster manager's proxy, enable ACLs with spark.acls.enable and the view and modify ACL settings, and consider spark.ui.killEnabled=false for production pipelines.

What to do next

  1. Turn on event logging to durable storage for every application and run a History Server with a disk store and the cleaner enabled.
  2. For your slowest recurring job, open the slowest stage and compare median and maximum in the summary metrics.
  3. Check the aggregated-by-executor table to separate data skew from bad hosts.
  4. Open the SQL tab for the same job and confirm join strategies, pruning and AQE decisions in the final plan.
  5. Look for spill across most tasks and for GC time that is a large share of task time; tune partitions and memory accordingly.
  6. Search driver logs for dropped listener events and, if present, reduce task count before raising queue capacity.
  7. Schedule a REST API check that flags skewed or spilling stages after each nightly run.
Key takeaway: The Spark UI is a view over the driver's event stream: a status listener builds the live pages, an event logger writes the same events for the History Server, and bounded queues and retention limits explain missing jobs and odd numbers. Read the event timeline for driver-side gaps, the stage summary metrics for skew and spill, the executor breakdown for bad hosts, and the SQL tab for what the optimizer actually did. Enable event logging everywhere, and automate the same checks through the REST API.