Most writing about Jaeger is about getting spans in: deploying collectors, choosing storage, sampling. That is half of the system. The other half is getting answers out: finding the one slow trace among millions, seeing which span actually made a request slow, and spotting regressions before anyone opens a trace. That half has its own design, and its limits explain most of the frustrating moments engineers have with the Jaeger UI.

This article covers the read side of Jaeger v2: the data model it stores, how storage indexes it, why search works the way it does, how the critical path finds the span to blame, how to query traces programmatically, and how the Monitor tab turns spans into rate, error and duration metrics. Deployment, storage choice, Kafka buffering and sampling are covered in distributed tracing with Jaeger.

Advertisement

The data model Jaeger reads

Jaeger v2 receives OpenTelemetry data. A trace is the set of spans sharing a 16-byte trace ID. Each span has an 8-byte span ID, a parent span ID (empty for the root), a name that Jaeger calls the operation, a start time and duration, a span kind (server, client, producer, consumer or internal), a status, attributes, events and optional links to spans in other traces. Each batch of spans also carries a resource, whose service.name attribute is what Jaeger calls the service.

Two consequences shape everything else. The service is resource-level, so every search starts from a service name, and a service that does not set service.name appears as an unknown or default service and is effectively unsearchable. And a trace is assembled at read time from spans that arrived independently, from different hosts with different clocks and possibly minutes apart, so it can be incomplete or slightly inconsistent. Broken propagation shows up here as many small traces instead of one; trace context propagation covers the fix.

How storage indexes spans

Jaeger search is two-phase: an index lookup finds trace IDs, then traces are fetched by IDOTLP spansfrom SDKs or a Collectorjaeger_storage_exporterone write per spanspan storerows keyed by trace IDindex entriesservice, operation, tag, durationUI or api_v3 clientservice + time range + filtersphase 1: index lookupreturns up to N trace IDsphase 2: fetch by IDall spans of each traceadjustersclock skew, span linksscan index partitionsread rows by trace IDEach index lookup is capped, and the trace IDs from several lookups are intersected,so a search can return fewer traces than requested even though more matching traces exist.
Every span is written once as data and again as several index entries; search reads the indexes, then the data.

No storage backend can answer arbitrary questions about billions of spans quickly, so Jaeger writes each span once as data and once more for each index it maintains. The Cassandra schema makes this explicit. These are the tables in the repository's v004 schema with their primary keys:

TablePrimary keyAnswers
traces(trace_id, span_id, span_hash)All spans of a trace, by trace ID
service_names(service_name)The service drop-down
operation_names_v2((service_name), span_kind, operation_name)The operation drop-down
service_name_index((service_name, bucket), start_time)Recent traces of a service
service_operation_index((service_name, operation_name), start_time)Recent traces of one operation
duration_index((service_name, operation_name, bucket), duration, start_time, trace_id)Traces of an operation within a duration range
tag_index((service_name, tag_key, tag_value), start_time, trace_id, span_id)Traces where a service set an exact attribute value
dependencies_v2(ts_bucket, ts)The service dependency graph

Read the partition keys, the part in the inner parentheses, and the limits of search follow. Every index is partitioned by service, so there is no query for "any service where error=true". Tag search matches an exact key and value inside one service, so there are no ranges, prefixes or regular expressions. And when a query has several conditions, the Cassandra reader looks each one up separately, reading at most three times the requested trace count from each index, and intersects the resulting trace IDs. If the newest entries for the operation and the newest entries for the tag are different traces, the intersection is small or empty although older matches exist. A duration filter takes its own path through duration_index: in the current reader, tag filters are ignored when a duration is set, and older releases rejected the combination with an error.

Elasticsearch and OpenSearch store each span as a document in a time-based index, so the query is an Elasticsearch search and is more flexible with tags. The cost moves elsewhere: a search over seven days touches seven daily indices, and every attribute you add to spans grows the mapping and the index size. Whatever the backend, the search then fetches every trace ID it found in full, which is why a broad search over a long window is slow even when it returns only 20 traces.

Advertisement

Worked example: a fan-out where one child owns the latency

A product page's p99 rises to about 820 ms. In the UI you search service gateway, operation GET /product, minimum duration 700 ms, over the last hour. One trace has this shape, with times in milliseconds from the start of the root:

SpanParentStartEnd
gateway GET /product0820
authgateway540
aggregategateway45810
profileaggregate50180
recommendaggregate50790
adsaggregate50210
featuresrecommend60150
rankrecommend160780

The aggregator calls three services in parallel. Summing durations is misleading: profile, recommend and ads total 1,030 ms, more than the whole request. What matters is which work the request was waiting on at each moment. That is the critical path, and the Jaeger UI can highlight it in the trace view; the UI configuration has a criticalPathEnabled setting. The idea is simple: walk backwards from the end of the root, and at each point blame the child that finished last before that point, recursing into it.

def critical_path(spans, root):
    """spans: id -> (start_ms, end_ms, parent_id). Returns {span_id: ms it holds the critical path}."""
    children = {}
    for sid, (_, _, parent) in spans.items():
        if parent is not None:
            children.setdefault(parent, []).append(sid)
    blame = {}

    def walk(sid, until):
        start, end, _ = spans[sid]
        cursor = min(end, until)
        # The child that finishes last (before the cursor) is what the parent was waiting for.
        for kid in sorted(children.get(sid, []), key=lambda k: spans[k][1], reverse=True):
            k_start, k_end, _ = spans[kid]
            k_end = min(k_end, cursor)
            if k_start >= cursor or k_end <= start:
                continue  # fully hidden behind a later child, or outside the parent
            blame[sid] = blame.get(sid, 0) + (cursor - k_end)  # parent's own work after the child
            walk(kid, k_end)
            cursor = max(k_start, start)
        blame[sid] = blame.get(sid, 0) + (cursor - start)

    walk(root, spans[root][1])
    return {k: v for k, v in blame.items() if v > 0}

Run on the trace above, it blames rank for 620 ms, features for 90, auth for 35, recommend for 30, aggregate for 25 and gateway for 20, which sums exactly to the 820 ms of the request. Profile and ads are off the path entirely: making them faster changes nothing. Features and rank run in sequence inside recommend, so either parallelising them or speeding up rank would help, and rank is where 76 percent of the time goes. That is the answer a flame of span durations cannot give you. Once you know the culprit, add an attribute such as candidate count to rank's span so the next search can filter on it.

Querying traces programmatically

For anything you do more than twice, use the API rather than the UI. Jaeger's query service serves a gRPC API on port 16685, with jaeger.api_v3.QueryService using OTLP-shaped messages and a legacy api_v2 service, and an HTTP JSON gateway for api_v3 under /api/v3 on port 16686. The UI's own /api/traces endpoints are described in the documentation as intentionally undocumented and subject to change, so do not build tools on them.

The HTTP routes are /api/v3/services, /api/v3/operations, /api/v3/traces for search and /api/v3/traces/{trace_id} for one trace. In the Jaeger 2.0 gateway source the search parameters are query.service_name, query.operation_name, query.start_time_min and query.start_time_max (both required, RFC 3339), query.num_traces, and query.duration_min and query.duration_max as Go durations such as 700ms; that version's GET form did not support attribute filters. The main branch has since switched to camelCase names such as query.serviceName and query.searchDepth, added query.attributes, and still accepts every name above as a deprecated alias, so the script below uses them to work against both; check the API page for the version you run. Responses are OTLP JSON wrapped in a result field. This script pulls slow traces and ranks spans by critical-path time, treating IDs as opaque strings:

import requests, datetime as dt
from collections import Counter

BASE = "http://jaeger-query:16686/api/v3/traces"
now = dt.datetime.now(dt.timezone.utc)
params = {
    "query.service_name": "gateway",
    "query.operation_name": "GET /product",
    "query.start_time_min": (now - dt.timedelta(hours=1)).isoformat(),
    "query.start_time_max": now.isoformat(),
    "query.duration_min": "700ms",
    "query.num_traces": "50",
}
data = requests.get(BASE, params=params, timeout=30).json()["result"]

traces = {}
for rs in data.get("resourceSpans", []):
    svc = next(a["value"]["stringValue"] for a in rs["resource"]["attributes"] if a["key"] == "service.name")
    for ss in rs["scopeSpans"]:
        for s in ss["spans"]:
            traces.setdefault(s["traceId"], {})[s["spanId"]] = (
                int(s["startTimeUnixNano"]) / 1e6, int(s["endTimeUnixNano"]) / 1e6,
                s.get("parentSpanId") or None, f'{svc} {s["name"]}')

blame = Counter()
for spans in traces.values():
    shape = {k: v[:3] for k, v in spans.items()}
    root = next(k for k, v in shape.items() if v[2] is None or v[2] not in shape)
    for sid, ms in critical_path(shape, root).items():
        blame[spans[sid][3]] += ms
print(blame.most_common(10))

Aggregated over 50 slow traces, this tells you whether one span is consistently to blame or whether slowness is spread around, which a single trace cannot.

The Monitor tab: metrics derived from spans

Searching requires knowing which service and operation to look at. Service Performance Monitoring, the Monitor tab, answers that question first: it shows request rate, error rate and latency percentiles per service and operation, derived from spans. Jaeger does not compute them itself. The OpenTelemetry spanmetrics connector turns spans into metrics, a Prometheus-compatible store keeps them, and Jaeger's query extension reads them back. This is trimmed from the repository's config-spm.yaml:

service:
  extensions: [jaeger_storage, jaeger_query]
  pipelines:
    traces:
      receivers: [otlp]
      processors: [batch]
      exporters: [jaeger_storage_exporter, spanmetrics]
    metrics/spanmetrics:
      receivers: [spanmetrics]
      exporters: [prometheus]
extensions:
  jaeger_query:
    storage:
      traces: some_storage
      metrics: some_metrics_storage
  jaeger_storage:
    metric_backends:
      some_metrics_storage:
        prometheus:
          endpoint: http://prometheus:9090
connectors:
  spanmetrics:
    metrics_flush_interval: 60s

The connector produces a calls_total counter, with errors distinguished by the status_code label, and a duration histogram such as duration_milliseconds_bucket. Two cautions apply. The metrics are computed only from spans that reach the pipeline, so with head sampling in the SDKs the rates must be scaled by the sampling ratio; only sampling done later in a Collector, such as tail sampling, can run after the connector has counted everything. And cardinality multiplies: the Jaeger documentation estimates about 72 series per operation with default settings, so a service that puts IDs into span names can overwhelm Prometheus. The pattern is covered in span metrics.

Operating the query side

  • Bound every search. A wide time range with no operation scans many index partitions or daily indices. Default the UI to the last hour and teach people to add an operation and a minimum duration.
  • Index what people search for. Tag search works only on exact values. Put stable, low-cardinality facts in attributes, such as tenant, region, route template and error type, and keep raw IDs out of span names.
  • Watch huge traces. Batch jobs and retry loops can produce traces with tens of thousands of spans that are slow to fetch and render. Split long-running work into linked traces.
  • Expect clock skew. Spans from different hosts are timed by different clocks. The query service applies adjusters before returning a trace, including one for clock skew, but a child that appears to start before its parent usually means a host clock is off.
  • Protect it. Traces contain URLs, user identifiers and sometimes payloads. Put the UI and APIs behind authentication and scrub sensitive attributes in a Collector before storage; see the OpenTelemetry Collector.
  • Keep what you need longer. Retention is short to control cost; save traces cited in incident reviews to archive storage, not by raising retention for everything.

Failure modes and trade-offs

  • Search returns too few traces or none although matches exist: the capped index lookups did not overlap, a duration filter silently dropped the tag filter, the time range is in the wrong time zone, or the span carrying the tag belongs to a different service than the one selected. Narrow the time window, or search on one condition at a time.
  • The slow trace was never kept: head sampling decided at the root before latency was known. Tail sampling keeps slow and failed traces at the cost of buffering; see tail sampling.
  • Dependency graph is empty or stale: for large backends the graph is computed by a separate batch job over stored spans, so check that it runs for your storage and Jaeger version.
  • Monitor tab is blank: the metrics backend is not configured on the query extension, or the metric names in Prometheus do not match what Jaeger expects.
  • The trade-off: every index makes a search possible and every span more expensive to write. Index the few dimensions people actually search, and answer the rest with span metrics.

What to do next

  1. Check that every service sets service.name and that one request produces one trace end to end.
  2. Write down the five searches your on-call engineers run most, and confirm each is served by an index.
  3. Turn on the critical path view and use it in the next latency investigation before reading span durations.
  4. Script one api_v3 query for your slowest endpoint and run it on a schedule against the version you deploy.
  5. Enable span metrics and the Monitor tab, and check series count per operation after a week.
  6. Put the query service behind authentication and add attribute scrubbing to the Collector.
Key takeaway: Jaeger stores every span once as data and again in a few per-service indexes, and search is an index lookup followed by fetching whole traces, so what you can find is decided by what is indexed when it is written. Use the critical path rather than raw durations to find the span to blame, use api_v3 rather than the UI's internal endpoints for tools, and let span metrics tell you where to look before you search.