Every distributed tracing system eventually produces more spans than anyone wants to pay to store. Sampling is the answer, but sampling is not one technique: it is a choice of where in the pipeline a keep-or-drop decision is made, what information is available at that point, and how the decisions made at different points agree with each other. Getting that choice wrong gives you broken traces with missing children, dashboards whose error rates are wrong by a factor of ten, or a tail sampler that runs out of memory at peak.

This article is the map that ties the individual mechanisms together. Head sampling, tail sampling and adaptive sampling each have their own article on this site; here the focus is choosing between them, composing them into one pipeline, and keeping the numbers you derive from sampled data honest. Details of the OpenTelemetry specification and Collector components were checked against their documentation in October 2026.

Advertisement

Why sample, and what it costs

The cost of tracing scales with spans, not requests. A service handling 4,000 requests per second, where each request produces 25 spans across the services it touches, emits 100,000 spans per second. At roughly one kilobyte per span on the wire, that is about 100 MB per second, or more than 8 TB per day, before indexing and replication. Most of those traces are successful, fast and identical in shape to thousands of others.

Sampling keeps a subset. What you lose depends on how the subset is chosen. A uniform random sample preserves statistics such as latency percentiles but keeps rare events, like a one-in-ten-thousand error, only rarely. A sample biased towards errors and slow requests keeps the interesting traces but can no longer be used to compute rates. Every strategy below is a different trade between those two properties, plus cost and complexity.

The decision points

Where the decision is madeWhat it can seeCost of dropped spansTrace completeness
Head, in the root service's SDKTrace ID, root span name and attributesNever created or exportedComplete if children follow the parent's flag
Probabilistic, in a collectorTrace ID, span attributes, tracestateExported, then droppedComplete if every collector uses the same rule
Tail, after the trace finishesEvery span: status, duration, attributesExported, buffered, then droppedComplete only if one instance sees the whole trace
Adaptive, any of the aboveRecent traffic rates per keyDepends on where it runsDepends on where it runs

The rows are not alternatives so much as layers. Head sampling is cheapest because unsampled spans are never serialised, but it decides before anything interesting has happened. Tail sampling decides with full knowledge but pays for every span to cross the network and sit in memory. Mature setups use a cheap early layer to control volume and a smart late layer to choose what to keep.

A hybrid sampling pipeline: decide cheaply early, decide well lateService SDKsParentBased samplerSpan metricsRED counts, unsampledAgent collectorsbatch, enrichspansLoad-balancing tierroute by trace IDTail sampler 1Tail sampler 2Tail sampler Nbuffer decision_wait,apply policiesTrace backenderrors, slow, baseline sampleMetrics backendcounts never sampledConsistent probability sampling on the 56-bit randomness value R02^56T for p = 0.5 (th:8)T for p = 0.25 (th:c)R below T: dropped at 50%R >= T: kept at both ratesA trace kept at 25% is always also kept at 50%: lower-rate samples are subsets
Top: a hybrid pipeline in which span metrics are computed before any sampling and tail samplers receive whole traces through a trace-ID-routed load-balancing tier. Bottom: how consistent probability thresholds make lower-rate samples subsets of higher-rate ones.
Advertisement

Head sampling and the parent's flag

In OpenTelemetry SDKs the usual head configuration is a ParentBased sampler wrapping a TraceIdRatioBased root sampler. The root span is sampled with probability p using the trace ID, and every downstream span copies its parent's decision, which travels in the sampled bit of the W3C traceparent header. The standard environment variables are OTEL_TRACES_SAMPLER=parentbased_traceidratio and OTEL_TRACES_SAMPLER_ARG=0.1 for ten percent.

The parent-based rule is what keeps traces whole: without it each service would sample independently and a trace would be missing random subtrees. The weakness is that a misconfigured service, one that ignores the incoming flag or does not propagate context, silently breaks traces. Head-based sampling architecture and trace context propagation cover both sides of that contract.

Consistent probability sampling

Parent-based sampling forces every service to accept the root's rate. Sometimes you want different rates at different places, for example a collector that further reduces a noisy service, without breaking traces. Consistent probability sampling makes that possible by deriving the decision from a value that every participant sees identically.

The OpenTelemetry specification for this, TraceState: Probability Sampling, has status Development at the time of writing, so check your SDK and Collector versions before relying on it. It defines a 56-bit randomness value R and a 56-bit rejection threshold T. R is normally the least significant 56 bits of the trace ID, which W3C Trace Context Level 2 allows a producer to declare random with a flag, or it can be carried explicitly in the rv sub-key of the ot entry in tracestate. The rule is: keep the span if R is greater than or equal to T. For a sampling probability p, T is (1 - p) times 2 to the 56.

T is written into tracestate as the th sub-key in hexadecimal with trailing zeros removed. The specification's examples are th:0 for keep everything, th:8 for 50 percent, th:c for 25 percent and th:fd70a4 for one percent.

MAX = 1 << 56

def threshold(p, digits=14):
    """Rejection threshold for probability p, encoded as the OTel th value."""
    t = round((1 - p) * MAX)
    step = 16 ** (14 - digits)                  # optional precision reduction, rounded
    t = min(round(t / step) * step, MAX - step)
    return t, (format(t, "014x").rstrip("0") or "0")

def randomness(trace_id_hex, rv=None):
    return int(rv, 16) if rv else int(trace_id_hex[-14:], 16)   # low 56 bits

def keep(trace_id_hex, p, rv=None):
    t, th = threshold(p)
    return randomness(trace_id_hex, rv) >= t, th

def adjusted_count(t):
    return MAX / (MAX - t)                       # equals 1 / p

print(threshold(0.5)[1], threshold(0.01, digits=6)[1])   # 8 fd70a4

Two properties make this useful. Because every sampler compares the same R, a trace kept at 25 percent is always kept at 50 percent: lower-rate samples are subsets of higher-rate ones, so traces stay whole even when rates differ by service. And because T is recorded on the span, any later consumer can compute the adjusted count, the number of original spans each kept span represents, which is the reciprocal of the probability.

In the Collector, the probabilistic sampler processor supports this through its proportional and equalizing modes; its default mode is hash_seed, an older FNV-hash scheme that is consistent only between collectors configured with the same seed. Proportional mode multiplies with upstream sampling, so 50 percent applied after an SDK's 50 percent leaves 25 percent of the original traffic.

Tail sampling and its buffer

A tail sampler waits until a trace is probably complete, then evaluates policies over all of its spans. The OpenTelemetry Collector's tail sampling processor, at beta stability for traces, supports policies including status_code, latency, probabilistic, rate_limiting, string_attribute, ottl_condition, and, drop and composite. A trace is kept if any policy votes to keep it, unless a drop policy matches.

Its documentation is explicit that all spans of a trace must reach the same collector instance. That is why the pipeline has two tiers: stateless collectors with the load-balancing exporter, configured with routing_key: traceID, in front of stateful tail samplers.

processors:
  tail_sampling:
    decision_wait: 10s
    num_traces: 60000
    expected_new_traces_per_sec: 1000
    decision_cache:
      sampled_cache_size: 200000
    policies:
      - name: errors
        type: status_code
        status_code: {status_codes: [ERROR]}
      - name: slow
        type: latency
        latency: {threshold_ms: 1500}
      - name: baseline
        type: probabilistic
        probabilistic: {sampling_percentage: 5}

The buffer is the operational risk. With the default decision_wait of 30 seconds, every trace started in the last 30 seconds is held in memory. num_traces caps how many, with a default of 50,000; when it is exceeded the oldest traces are evicted before a decision, which silently loses exactly the slow traces you wanted. The sampled decision cache lets spans that arrive after the decision follow it instead of being judged as a new, incomplete trace. Tail-based sampling architecture goes deeper into routing and decision windows.

Worked example: sizing the tail tier

Take the service above: 4,000 traces per second, 25 spans each. With a 30 second decision window, 120,000 traces and 3 million spans are in flight at once. Spread over four tail samplers, each holds 30,000 traces, already more than half the default num_traces, with no headroom for a traffic spike. In-memory spans are larger than their wire size, so budget well above 3 GB across the tier.

The fixes, in order: measure how long your traces really take, and set decision_wait a little above the 99th percentile trace duration, perhaps 10 seconds, which cuts in-flight state by two thirds. Set num_traces per instance to at least twice the expected in-flight count. Add a head sampler at 50 percent to halve the input, accepting that rare errors are then kept only half the time. Then alert on the processor's eviction and late-span metrics, because both mean decisions are being made on partial traces.

With the policies above, the backend receives every error trace, every trace slower than 1.5 seconds and 5 percent of the rest, typically a small fraction of the input. If errors are 0.5 percent and slow traces 1 percent, about 6.4 percent of traces are kept.

Keeping the numbers honest

Sampled data answers two kinds of question differently. Questions about individual traces, such as why one checkout was slow, need the trace and nothing else. Questions about rates and counts, such as error rate or requests per second, need either unsampled counts or correct weights.

The robust answer is to compute request, error and duration metrics before sampling, for example with the Collector's span metrics connector placed in the agent tier, and to use traces only for exemplars and drill-down. Where you must estimate from sampled spans, weight each span by its adjusted count. A uniform 10 percent sample with weights of 10 gives unbiased counts. A tail sample does not: error traces are kept at 100 percent and the rest at 5 percent, so an unweighted error rate would be grossly inflated. Record the effective probability on each kept trace, for example as an attribute, so weights survive tail decisions too.

def estimated_counts(spans):
    total = errors = 0.0
    for s in spans:
        w = s.attributes.get("sampling.adjusted_count", 1.0)   # 1 / probability
        total += w
        errors += w if s.status == "ERROR" else 0.0
    return total, errors, errors / total if total else 0.0

The attribute name in that snippet is a local convention, not a standard field; what matters is that every stage that changes probability updates it.

Failure modes

  • Broken traces from mixed decisions. One service samples independently instead of following the parent, or two collectors hash with different seeds. Detect with a job that counts traces whose spans reference missing parents.
  • Eviction before decision. num_traces is too small for the window, so long traces are dropped first. Watch eviction counters.
  • Split traces on scale-out. When the tail tier scales, the load-balancing ring changes and in-flight traces are split across two instances. Scale during quiet periods and keep the late-span decision cache.
  • Late and asynchronous spans. Queue consumers and batch jobs finish after the decision window. Either link them as separate traces or accept that they follow the cached decision.
  • Biased dashboards. Rates computed from tail-sampled traces without weights.
  • Mixed SDK versions. Older SDKs that do not set the random flag or the th value break consistent sampling assumptions downstream.

Choosing a strategy

For a small system, head sampling with a ParentBased ratio sampler and metrics computed before sampling is enough, and it costs almost nothing. Add tail sampling when you need every error and slow trace and can run a stateful tier. Add consistent probability sampling when different teams need different rates without breaking traces. Add adaptive rate control, described in adaptive trace sampling, when traffic varies so much that a fixed rate either overspends at peak or keeps too little at night.

What to do next

  1. Measure spans per second, bytes per span and the p99 trace duration for your busiest services.
  2. Move request, error and duration metrics ahead of any sampling so dashboards never depend on sampled traces.
  3. Configure ParentBased head sampling everywhere and verify propagation with a broken-trace detector.
  4. If you add tail sampling, deploy the load-balancing tier first, size num_traces from your measurements and alert on evictions.
  5. Record the effective sampling probability on every kept trace and weight any estimates by it.
  6. Check whether your SDKs and Collector support the th and rv keys before adopting consistent probability sampling across teams.
Key takeaway: Trace sampling is a set of decision points, not a single setting. Head sampling is cheap and keeps traces whole through the parent's flag but decides blind; tail sampling decides with full knowledge but needs trace-ID routing and a carefully sized buffer; consistent probability sampling lets rates differ by service while keeping lower-rate samples subsets of higher ones. Compute your metrics before any sampling, carry the effective probability with every kept trace, and treat evictions and broken traces as alerts.