A latency histogram tells you that the 99th percentile of checkout requests jumped from 300 ms to 2.1 s at 14:05. It cannot tell you why, because a histogram is an aggregate: thousands of requests were folded into a handful of bucket counters and the individual requests are gone. Traces have the opposite shape. Each trace explains one request in detail, but there are millions of them and you do not know which one to open. An exemplar joins the two. It is a single example observation, stored next to a metric sample, that carries the trace id of a real request that landed in that bucket.
This article explains the mechanism from the instrumentation call to the click in a dashboard: the wire format, how client libraries and the OpenTelemetry SDK choose which observation becomes the exemplar, how Prometheus stores and serves exemplars, how Grafana turns them into links, and the operational traps, especially sampling, that make exemplars point at traces that no longer exist. Background on the two signals it joins is in histograms and distributed traces.
What an exemplar is, precisely
An exemplar has four parts: a small label set (almost always just a trace id, sometimes a span id), the observed value, an optional timestamp, and the metric sample it is attached to. If a request took 2.31 s and fell into the histogram bucket with upper bound 2.5, the exemplar on that bucket says: "one of the observations counted here was 2.31 s, at this time, and its trace id is 4bf92f35...". The exemplar does not change the metric. The bucket counter increments exactly as it would without it.
Two properties follow. First, an exemplar is a sample, not an index. A bucket that counted 40,000 requests in a scrape interval carries one exemplar, so you get a representative, not a lookup of every slow request. Second, exemplar labels are not series labels. Putting a trace id into a normal label would create a new time series per request and destroy the metrics backend. Exemplar labels are stored on the side, with a fixed memory budget, which is the whole reason the feature exists.
The architecture end to end
There are two common paths, and both end in the same place. In the Prometheus-native path, application code calls the client library with the observed value and the current trace id; the library keeps the most recent exemplar per bucket or counter; the metrics endpoint renders it in the OpenMetrics text format; Prometheus scrapes it and keeps it in a separate exemplar store. In the OpenTelemetry path, the SDK attaches exemplars automatically from the active span context, sends them over OTLP to a Collector, and the Collector forwards them to a backend through remote write or a Prometheus exporter.
The tracing backend is a separate system. The exemplar holds only an id, so the link works only if the trace with that id was exported, kept by sampling, and is still within the backend's retention window.
The wire format
In the OpenMetrics text format an exemplar follows a sample line after a # separator: a label set in braces, the value, and an optional Unix timestamp. Exemplars are allowed on histogram _bucket lines and on counter _total lines. The classic Prometheus text format (version 0.0.4) has no exemplar syntax, so a target that only speaks that format cannot ship them.
# TYPE http_request_duration_seconds histogram
http_request_duration_seconds_bucket{route="/checkout",le="0.5"} 9810 # {trace_id="a1f3c09e7d4b2e61"} 0.41 1759327500.112
http_request_duration_seconds_bucket{route="/checkout",le="2.5"} 9944 # {trace_id="4bf92f3577b34da6"} 2.31 1759327502.870
http_request_duration_seconds_bucket{route="/checkout",le="+Inf"} 9951
http_request_duration_seconds_sum{route="/checkout"} 2817.4
http_request_duration_seconds_count{route="/checkout"} 9951
# TYPE checkout_errors counter
checkout_errors_total{reason="card_declined"} 37 # {trace_id="9c0e11d2aa7f4410"} 1 1759327499.004
# EOFThe OpenMetrics specification caps the combined length of an exemplar's label names and values at 128 UTF-8 characters. A 32-hex-character trace id plus a 16-character span id fits comfortably; a user id, URL and tenant name does not. Client libraries enforce the limit loudly: the Go client panics in the observe call and the Python client raises a ValueError, so an oversized label set breaks the request path rather than just losing the exemplar. Keep exemplar labels to identifiers that let you find the trace and nothing more.
Attaching exemplars in a Prometheus client library
In the Go client the histogram and counter types implement the ExemplarObserver and ExemplarAdder interfaces. You attach the trace id only when the span is sampled, so the exemplar never names a trace that the tracer already decided to drop. The handler must have OpenMetrics enabled, or the exemplars are silently left out of the response.
var latency = prometheus.NewHistogramVec(prometheus.HistogramOpts{
Name: "http_request_duration_seconds",
Help: "Request latency.",
Buckets: []float64{.025, .05, .1, .25, .5, 1, 2.5, 5},
}, []string{"route"})
func record(ctx context.Context, route string, d time.Duration) {
obs := latency.WithLabelValues(route)
sc := trace.SpanContextFromContext(ctx) // OpenTelemetry span context
if sc.IsSampled() {
obs.(prometheus.ExemplarObserver).ObserveWithExemplar(
d.Seconds(), prometheus.Labels{"trace_id": sc.TraceID().String()})
return
}
obs.Observe(d.Seconds())
}
// Exemplars are only rendered when the scraper negotiates OpenMetrics.
http.Handle("/metrics", promhttp.HandlerFor(reg,
promhttp.HandlerOpts{EnableOpenMetrics: true}))The Python client takes the exemplar as an argument to observe and inc. It renders exemplars only in the OpenMetrics output, which its HTTP handlers choose when the scraper's Accept header asks for it.
from prometheus_client import Histogram
from opentelemetry import trace
LATENCY = Histogram("http_request_duration_seconds", "Request latency.", ["route"])
def record(route: str, seconds: float) -> None:
ctx = trace.get_current_span().get_span_context()
exemplar = {"trace_id": format(ctx.trace_id, "032x")} if ctx.trace_flags.sampled else None
LATENCY.labels(route).observe(seconds, exemplar=exemplar)Both libraries keep the latest exemplar per bucket and overwrite it on every new one. Between two scrapes, the exemplar you see on the 2.5 s bucket is simply the last request that landed there, which is usually what you want: it is recent and it is genuinely slow.
Exemplars in OpenTelemetry: filters and reservoirs
The OpenTelemetry metrics SDK makes exemplars automatic. Two pluggable pieces decide which measurements become exemplars. The exemplar filter decides which measurements are eligible: always_on, always_off, or trace_based, which admits only measurements recorded while a sampled span is active. trace_based is the default and can be set with the OTEL_METRICS_EXEMPLAR_FILTER environment variable. The exemplar reservoir decides which eligible measurements are kept until the next export. For explicit-bucket histograms the default reservoir keeps one exemplar per bucket, aligned with the buckets; for other instruments a small fixed-size reservoir samples uniformly.
On the wire, an OTLP exemplar carries the observed value, a timestamp, the trace id and span id as dedicated fields, and any measurement attributes that the view filtered out of the metric's own attributes. That last field is useful: if a view drops a high-cardinality attribute such as a customer id to keep series counts down, the dropped value can still survive on the exemplar.
From the Collector, exemplars reach Prometheus-compatible backends through the Prometheus remote-write exporter, or through the Prometheus exporter when it is configured to serve OpenMetrics. The spanmetrics connector, which derives RED metrics from spans, can also attach exemplars so its latency histograms link back to the spans that produced them; span metrics covers that connector in depth.
Prometheus storage and the query API
Prometheus keeps exemplars outside the time-series database. With exemplar storage enabled (the --enable-feature=exemplar-storage flag; check your version's feature-flag page), exemplars go into a fixed-size in-memory circular buffer shared by all series and are also written to the WAL so a restart does not lose them. The buffer size is set in the configuration file, and the documentation estimates roughly 100 bytes per exemplar that carries only a trace id, so 100,000 exemplars cost on the order of 10 MB.
# prometheus.yml
storage:
exemplars:
max_exemplars: 200000
remote_write:
- url: https://metrics.example.internal/api/v1/write
send_exemplars: trueBecause the buffer is circular, retention is measured in count, not time. A busy server with thousands of histogram buckets overwrites old exemplars quickly, so a spike from yesterday may have none left even though its series data is still there. Long-term stores that accept remote write with exemplars hold them according to their own limits.
Exemplars are queried through their own endpoint, /api/v1/query_exemplars, which takes a PromQL expression and a time range and returns the exemplars of every series the expression selects. Dashboards call it next to the normal range query.
curl -s 'http://prometheus:9090/api/v1/query_exemplars' \
--data-urlencode 'query=http_request_duration_seconds_bucket{route="/checkout"}' \
--data-urlencode 'start=2026-10-01T14:00:00Z' \
--data-urlencode 'end=2026-10-01T14:15:00Z'
Grafana: turning ids into links
In Grafana the Prometheus data source has an Exemplars setting where you map an exemplar label, usually trace_id, to an internal link into a tracing data source such as Tempo or Jaeger, or to an external URL template. Each query in a time-series panel has an Exemplars toggle. When it is on, Grafana fetches exemplars for the same selector and draws them as points at the exemplar's time and value. Hovering shows the labels; clicking the trace id opens the trace. Dashboard practice is covered in the Grafana deep dive.
Worked example: from a p99 spike to a root cause
At 14:05 the checkout p99, computed with histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket{route="/checkout"}[5m]))), rises from 300 ms to 2.1 s. With exemplars turned on, the panel shows a cluster of dots between 1.8 s and 2.4 s starting at 14:04. You click one; it opens trace 4bf92f3577b34da6. The trace shows the checkout service spending 1.9 s in a span named inventory.reserve, and inside it a database span waiting on a lock. Two more dots from the same window show the same shape. You now have a hypothesis, lock contention in the inventory database, within a minute, and you can confirm it with the database's lock metrics.
Note what the exemplar did not do. It did not tell you that every slow request had this cause; it gave you three real examples. Check a few dots from different buckets and different pods before concluding. A single exemplar can be an outlier of the outliers.
The sampling trap
Most production tracing samples. With head sampling, the decision is made when the trace starts, and the trace_based filter or an IsSampled check keeps exemplars consistent with it: only sampled traces become exemplars. The cost is coverage. At a 1% sampling rate, a bucket that sees ten slow requests per scrape interval will often have no sampled one, and the slow bucket shows no dot exactly when you need it.
With tail sampling the problem inverts. The SDK sees the span as sampled, attaches the exemplar, and then the Collector's tail sampler drops the trace a few seconds later because it looked normal. The dashboard shows a dot that leads to "trace not found". The fix is to make the tail-sampling policy keep what exemplars are likely to point at: keep all errors and all traces above a latency threshold close to your slowest buckets. Tail sampling explains the policy types. Also align retention: exemplars older than the tracing backend's retention are dead links by construction.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| No exemplars at all | Target serves the classic text format, or the handler does not enable OpenMetrics | Enable OpenMetrics in the handler and confirm the Accept negotiation with curl |
| Scraped but not queryable | Exemplar storage not enabled in Prometheus | Enable the feature and set max_exemplars |
| Dots only for recent minutes | Circular buffer too small for the series count | Raise max_exemplars or send exemplars to a long-term store |
| Click leads to trace not found | Tail sampling or retention dropped the trace | Keep slow and error traces; align retention |
| Panic or exception in the instrumentation call | Exemplar label set over the 128-character limit | Carry only trace and span ids |
| Dots missing in a remote backend | send_exemplars not set on remote write | Enable it and check the receiver accepts exemplars |
Trade-offs and cost
Exemplars are cheap compared with the alternatives. One exemplar per bucket per scrape adds a few dozen bytes to each scrape response and about 100 bytes per stored exemplar, against the alternative of correlating logs and traces by time window and guessing. The costs to weigh are memory in the exemplar buffer, a small per-observation overhead to read the span context, and the coupling to tracing: exemplars are only as good as your trace retention and sampling. They also do not help for metrics that have no request context, such as queue depth gauges, and gauges are not an exemplar target in the OpenMetrics format.
Do not use exemplars as a cardinality escape hatch. Putting a user id on every exemplar does not create series, but it does make the 128-character budget tight and can leak personal data into a system that was not designed to hold it. For per-entity analysis, use traces or logs, and keep series cardinality under control separately.
What to do next
- Pick one latency histogram on a service that is already traced and attach trace ids to observations only when the span is sampled.
- Enable OpenMetrics on the metrics handler and confirm with
curl -H 'Accept: application/openmetrics-text' host:port/metricsthat bucket lines carry a# {trace_id=...}suffix. - Enable exemplar storage in Prometheus, size max_exemplars from your bucket count and scrape rate, and set send_exemplars on remote write if you use a long-term store.
- Map the trace_id label to your tracing data source in Grafana and turn on the Exemplars toggle in the latency panel.
- Change tail-sampling policies to keep all error traces and all traces above your slowest bucket boundary, and check that tracing retention covers the exemplar retention.
- Run a drill: inject latency in staging, follow a dot to its trace, and note every link that did not resolve.