"The three pillars of observability" names the three kinds of telemetry most systems emit: metrics, logs and traces. The phrase is useful shorthand and misleading as a design. Each pillar answers a different question at a very different cost, and their value comes almost entirely from how easily you can move from one to another during an investigation. Three disconnected tools with three query languages and no shared identifiers give you three partial views of an outage, not observability.

This page treats them as one system. It explains what each signal is good and bad at, follows a single request as it produces all three, shows the identifiers that join them, walks through an incident, estimates what each costs at a realistic request rate, and ends with the critiques of the pillar model and a checklist. Each signal has its own in-depth page on this site; this one is about the connections between them.

Advertisement

Three signals, three questions

MetricsLogsTraces
UnitA numeric time series: name plus labels, sampled or aggregated over timeA timestamped event, ideally structured key-value fieldsA tree of spans for one request, each with timing, attributes and a parent
AnswersIs something wrong, how much, since when?What exactly happened in this component?Where in the call path did time go or the error start?
Cost driverNumber of distinct series (cardinality)Bytes ingested and indexedSpans kept after sampling
Typical retentionMonths to years, downsampledDays to weeks hot, longer in cheap storageDays
Weak atExplaining individual requestsAggregates and cross-service causalityLong-term trends and rare events lost to sampling

The table explains the shape of most observability systems. Metrics are pre-aggregated, so they are cheap to store for a long time and fast to query, which makes them the right source for alerts and SLOs. Aggregation also throws away the identity of individual requests, so a metric can tell you that p99 latency tripled but not which requests were slow. Logs and traces keep per-event detail at a per-event price, which is why they are retained for shorter periods and often sampled.

One request emits all three

Consider a checkout request that takes 2.4 seconds. If the service is instrumented with OpenTelemetry, three things happen. The tracing SDK creates a span named checkout with a 128-bit trace id, and the instrumented HTTP and database clients create child spans that carry the same trace id and propagate it to downstream services in the W3C traceparent header. The metrics SDK records 2,400 ms into a histogram labelled only with low-cardinality attributes such as outcome. The logger writes one structured event containing the order id, the outcome, the duration, and the current trace and span ids.

import logging
from opentelemetry import metrics, trace

tracer = trace.get_tracer("checkout")
meter = metrics.get_meter("checkout")
latency = meter.create_histogram("checkout.duration", unit="ms")
log = logging.getLogger("checkout")

def checkout(order):
    with tracer.start_as_current_span("checkout") as span:
        span.set_attribute("order.items", len(order.items))
        t0 = now_ms()
        outcome = "error"
        try:
            charge(order)                   # child spans come from instrumented clients
            outcome = "ok"
        except PaymentDeclined:
            outcome = "declined"
            raise
        finally:
            ms = now_ms() - t0
            # Low-cardinality attributes only: never the order id or user id here.
            latency.record(ms, {"outcome": outcome})
            ctx = span.get_span_context()
            log.info("checkout finished", extra={
                "trace_id": format(ctx.trace_id, "032x"),
                "span_id": format(ctx.span_id, "016x"),
                "order_id": order.id,       # high cardinality belongs in logs and spans
                "outcome": outcome,
                "duration_ms": ms,
            })

The split of attributes is the important design decision. The order id goes into the log and could go on the span, where high cardinality is affordable because each record is stored once. It must never become a metric label: one label with a million values produces a million time series. The outcome label, with a handful of values, is safe everywhere.

Advertisement

The joins that make it one system

Three kinds of shared key let an investigation move between signals.

  • Resource attributes. service.name, deployment environment, version, host or pod, attached identically to every signal by the SDK or the collector. Without them you cannot even say that a metric and a log came from the same process. Enforce one naming convention in the collector rather than trusting each team.
  • Trace and span ids in logs. Writing the current ids into every log event lets you jump from a span to the exact log lines it produced, and from an error log to its full trace. OpenTelemetry's log data model has dedicated trace and span id fields, and log bridges for common logging libraries can fill them automatically.
  • Exemplars on metrics. An exemplar is a sample measurement attached to a metric point together with the trace id of a request that produced it. A latency histogram bucket can thus point to one concrete slow trace. In Prometheus, exemplar storage is a feature flag (--enable-feature=exemplar-storage) backed by a fixed-size in-memory buffer, and scraped targets must also expose exemplars, for example in the OpenMetrics format.
# Prometheus stores exemplars only with this feature flag; scraped targets must
# also expose them, for example in the OpenMetrics format.
prometheus --enable-feature=exemplar-storage --config.file=prometheus.yml

# PromQL: p99 checkout latency; a dashboard can show exemplar trace ids on this panel.
histogram_quantile(0.99, sum by (le) (rate(checkout_duration_milliseconds_bucket[5m])))
One request, three signals, joined by shared keyscheckout serviceOpenTelemetry SDKCollectorbatch, sample, enrichMetrics storeaggregates, exemplarsLog storeevents with trace_idTrace storespans by trace_idOTLP1. Alert on a metricp99 latency, error rate2. Exemplar to a tracewhich span was slow3. trace_id to logswhat the code saidShared keys: service.name and other resource attributes on all three; trace_id on spans, logs and exemplars.
Telemetry flows through a collector into three stores; investigations flow back across them through shared keys.

An incident, end to end

An SLO burn-rate alert fires: checkout p99 latency has risen from 400 ms to 2.5 s over 20 minutes, while the error rate is flat. The metric tells you the scope: all regions, starting at 14:05, coinciding with no deploy. That is all a metric can say.

Opening the latency panel, the on-call engineer clicks an exemplar in the slowest bucket and lands on a trace. The waterfall shows the checkout span at 2.4 s, and inside it a call to the inventory service of 2.1 s, almost all of it a single database query span. Other exemplars show the same pattern. The trace has localised the problem to one dependency and one query, which no amount of staring at service-level metrics would have done quickly.

Filtering logs by the inventory service and that trace id shows a warning that the query plan changed after a statistics refresh at 14:04. The logs supplied the why. The fix is to restore the previous plan, the metric confirms recovery, and the postmortem adds a metric for query plan changes. The sequence, metric to trace to log, is the normal order: metrics detect, traces locate, logs explain.

What each signal costs: a worked estimate

Take a service handling 2,000 requests per second. These numbers are illustrative assumptions, not benchmarks; substitute your own.

SignalAssumptionRaw volume per day
Logs5 events per request, 1 KB each2,000 x 5 x 1 KB = 10 MB/s, about 864 GB
Traces at 100 percent20 spans per request, 500 bytes each20 MB/s, about 1.7 TB
Traces at 10 percent head samplingSame spans, 1 in 10 traces keptAbout 173 GB
Metrics20,000 series, one sample every 15 s, about 2 bytes per compressed sample20,000 x 5,760 x 2 bytes, about 230 MB

Two conclusions follow. Metrics cost is driven by series count, not by traffic: double the requests and the metric bill barely moves, but add a label with 1,000 values and it multiplies by up to 1,000. Log and trace cost scales with traffic, so it is controlled by verbosity, sampling and retention. Most observability cost problems are either a cardinality explosion in metrics or unbounded debug logging, and the estimate above tells you which to look for first.

Sampling and cardinality trade-offs

Head sampling decides at the root span whether to keep a trace, so it is cheap and consistent across services, but it discards the rare slow or failed request as readily as a normal one. Tail sampling buffers complete traces in a collector tier and keeps those with errors or high latency, at the cost of memory and a stateful routing layer. A common compromise keeps all error and slow traces plus a small random share of the rest; whatever you choose, record the sampling rate so counts derived from traces can be corrected.

Because exemplars and log trace ids point to traces, sampling breaks links: an exemplar can reference a trace that was dropped. Prefer tail sampling for the traces you expect to reach from alerts, or accept and document the dead links. On the metrics side, keep a reviewed allowlist of labels and drop or aggregate the rest in the collector before they reach storage.

Where the pillar model falls short

The strongest critique is that the pillars describe storage formats, not goals. An engineer debugging an outage wants to ask arbitrary questions of rich per-request data, and three separately indexed copies of partial data make that hard. The wide-event approach records one structured event per request per service with many attributes, from which metrics can be derived and traces assembled. It trades cheaper pre-aggregation for a storage engine able to aggregate high-cardinality events quickly. Many teams end up in between: traces and structured logs as the rich source, with metrics derived from spans for dashboards and alerts.

The model is also incomplete. Continuous profiling, sampled stack traces showing which functions use CPU and memory, answers questions none of the three can, such as why a span was slow when it called nothing. OpenTelemetry announced a public alpha of its profiles signal in March 2026; as an alpha, expect the format and tooling to change, and check its current status before depending on it.

Failure modes

  • Unjoinable signals. Different service names per signal, or logs without trace ids. Symptom: every investigation becomes a manual timestamp match. Fix: set resource attributes once in the SDK and enforce them in the collector.
  • Cardinality explosions. A user id or URL path with ids used as a metric label. Fix: label allowlists and limits on the number of series per metric.
  • Broken propagation. A proxy, queue or thread pool that drops the trace context, splitting one request into several unrelated traces. Fix: test propagation across each hop, including asynchronous ones.
  • Signals that disagree. Error rates computed differently from metrics and from logs. Fix: define the SLI once and derive the other views from it.
  • Clock skew. Child spans that appear to start before their parent. Fix: time synchronisation on every host, and treat small negative offsets as skew, not causality.

What to do next

  1. For your most important service, write down which questions you answer with each signal today, and where the hand-off between them is manual.
  2. Standardise service.name and environment across metrics, logs and traces, enforced in the collector.
  3. Add trace and span ids to every structured log event, and verify that one click goes from a trace to its logs.
  4. Enable exemplars on your latency histograms and confirm a dashboard panel links to a real trace.
  5. Estimate daily volume per signal with the method above, then set a label allowlist and a log retention policy.
  6. Choose a sampling policy that keeps error and slow traces, and record the rate.
  7. Go deeper with metrics architecture, logs architecture, tracing architecture, structured logging, trace sampling strategies and the OpenTelemetry Collector.
Key takeaway: Metrics, logs and traces answer different questions at different costs: metrics detect and trend cheaply but lose individual requests, traces locate where time and errors arise within one request, and logs explain what the code did. Their value comes from joining them. Use the same resource attributes on every signal, write trace and span ids into logs, and attach exemplars to latency histograms so an alert leads to a trace and the trace leads to logs. Metric cost follows cardinality and log and trace cost follows traffic, so control them with label allowlists, sampling and retention. Treat the pillars as storage formats, not goals, and consider wide events and profiling where they answer questions the three cannot.