Every distributed tracing system eventually produces more spans than anyone wants to pay to store. Sampling is the answer, but sampling is not one technique: it is a choice of where in the pipeline a keep-or-drop decision is made, what information is available at that point, and how the decisions made at different points agree with each other. Getting that choice wrong gives you broken traces with missing children, dashboards whose error rates are wrong by a factor of ten, or a tail sampler that runs out of memory at peak.
This article is the map that ties the individual mechanisms together. Head sampling, tail sampling and adaptive sampling each have their own article on this site; here the focus is choosing between them, composing them into one pipeline, and keeping the numbers you derive from sampled data honest. Details of the OpenTelemetry specification and Collector components were checked against their documentation in October 2026.
Why sample, and what it costs
The cost of tracing scales with spans, not requests. A service handling 4,000 requests per second, where each request produces 25 spans across the services it touches, emits 100,000 spans per second. At roughly one kilobyte per span on the wire, that is about 100 MB per second, or more than 8 TB per day, before indexing and replication. Most of those traces are successful, fast and identical in shape to thousands of others.
Sampling keeps a subset. What you lose depends on how the subset is chosen. A uniform random sample preserves statistics such as latency percentiles but keeps rare events, like a one-in-ten-thousand error, only rarely. A sample biased towards errors and slow requests keeps the interesting traces but can no longer be used to compute rates. Every strategy below is a different trade between those two properties, plus cost and complexity.
The decision points
| Where the decision is made | What it can see | Cost of dropped spans | Trace completeness |
|---|---|---|---|
| Head, in the root service's SDK | Trace ID, root span name and attributes | Never created or exported | Complete if children follow the parent's flag |
| Probabilistic, in a collector | Trace ID, span attributes, tracestate | Exported, then dropped | Complete if every collector uses the same rule |
| Tail, after the trace finishes | Every span: status, duration, attributes | Exported, buffered, then dropped | Complete only if one instance sees the whole trace |
| Adaptive, any of the above | Recent traffic rates per key | Depends on where it runs | Depends on where it runs |
The rows are not alternatives so much as layers. Head sampling is cheapest because unsampled spans are never serialised, but it decides before anything interesting has happened. Tail sampling decides with full knowledge but pays for every span to cross the network and sit in memory. Mature setups use a cheap early layer to control volume and a smart late layer to choose what to keep.
Head sampling and the parent's flag
In OpenTelemetry SDKs the usual head configuration is a ParentBased sampler wrapping a TraceIdRatioBased root sampler. The root span is sampled with probability p using the trace ID, and every downstream span copies its parent's decision, which travels in the sampled bit of the W3C traceparent header. The standard environment variables are OTEL_TRACES_SAMPLER=parentbased_traceidratio and OTEL_TRACES_SAMPLER_ARG=0.1 for ten percent.
The parent-based rule is what keeps traces whole: without it each service would sample independently and a trace would be missing random subtrees. The weakness is that a misconfigured service, one that ignores the incoming flag or does not propagate context, silently breaks traces. Head-based sampling architecture and trace context propagation cover both sides of that contract.
Consistent probability sampling
Parent-based sampling forces every service to accept the root's rate. Sometimes you want different rates at different places, for example a collector that further reduces a noisy service, without breaking traces. Consistent probability sampling makes that possible by deriving the decision from a value that every participant sees identically.
The OpenTelemetry specification for this, TraceState: Probability Sampling, has status Development at the time of writing, so check your SDK and Collector versions before relying on it. It defines a 56-bit randomness value R and a 56-bit rejection threshold T. R is normally the least significant 56 bits of the trace ID, which W3C Trace Context Level 2 allows a producer to declare random with a flag, or it can be carried explicitly in the rv sub-key of the ot entry in tracestate. The rule is: keep the span if R is greater than or equal to T. For a sampling probability p, T is (1 - p) times 2 to the 56.
T is written into tracestate as the th sub-key in hexadecimal with trailing zeros removed. The specification's examples are th:0 for keep everything, th:8 for 50 percent, th:c for 25 percent and th:fd70a4 for one percent.
MAX = 1 << 56
def threshold(p, digits=14):
"""Rejection threshold for probability p, encoded as the OTel th value."""
t = round((1 - p) * MAX)
step = 16 ** (14 - digits) # optional precision reduction, rounded
t = min(round(t / step) * step, MAX - step)
return t, (format(t, "014x").rstrip("0") or "0")
def randomness(trace_id_hex, rv=None):
return int(rv, 16) if rv else int(trace_id_hex[-14:], 16) # low 56 bits
def keep(trace_id_hex, p, rv=None):
t, th = threshold(p)
return randomness(trace_id_hex, rv) >= t, th
def adjusted_count(t):
return MAX / (MAX - t) # equals 1 / p
print(threshold(0.5)[1], threshold(0.01, digits=6)[1]) # 8 fd70a4Two properties make this useful. Because every sampler compares the same R, a trace kept at 25 percent is always kept at 50 percent: lower-rate samples are subsets of higher-rate ones, so traces stay whole even when rates differ by service. And because T is recorded on the span, any later consumer can compute the adjusted count, the number of original spans each kept span represents, which is the reciprocal of the probability.
In the Collector, the probabilistic sampler processor supports this through its proportional and equalizing modes; its default mode is hash_seed, an older FNV-hash scheme that is consistent only between collectors configured with the same seed. Proportional mode multiplies with upstream sampling, so 50 percent applied after an SDK's 50 percent leaves 25 percent of the original traffic.
Tail sampling and its buffer
A tail sampler waits until a trace is probably complete, then evaluates policies over all of its spans. The OpenTelemetry Collector's tail sampling processor, at beta stability for traces, supports policies including status_code, latency, probabilistic, rate_limiting, string_attribute, ottl_condition, and, drop and composite. A trace is kept if any policy votes to keep it, unless a drop policy matches.
Its documentation is explicit that all spans of a trace must reach the same collector instance. That is why the pipeline has two tiers: stateless collectors with the load-balancing exporter, configured with routing_key: traceID, in front of stateful tail samplers.
processors:
tail_sampling:
decision_wait: 10s
num_traces: 60000
expected_new_traces_per_sec: 1000
decision_cache:
sampled_cache_size: 200000
policies:
- name: errors
type: status_code
status_code: {status_codes: [ERROR]}
- name: slow
type: latency
latency: {threshold_ms: 1500}
- name: baseline
type: probabilistic
probabilistic: {sampling_percentage: 5}The buffer is the operational risk. With the default decision_wait of 30 seconds, every trace started in the last 30 seconds is held in memory. num_traces caps how many, with a default of 50,000; when it is exceeded the oldest traces are evicted before a decision, which silently loses exactly the slow traces you wanted. The sampled decision cache lets spans that arrive after the decision follow it instead of being judged as a new, incomplete trace. Tail-based sampling architecture goes deeper into routing and decision windows.
Worked example: sizing the tail tier
Take the service above: 4,000 traces per second, 25 spans each. With a 30 second decision window, 120,000 traces and 3 million spans are in flight at once. Spread over four tail samplers, each holds 30,000 traces, already more than half the default num_traces, with no headroom for a traffic spike. In-memory spans are larger than their wire size, so budget well above 3 GB across the tier.
The fixes, in order: measure how long your traces really take, and set decision_wait a little above the 99th percentile trace duration, perhaps 10 seconds, which cuts in-flight state by two thirds. Set num_traces per instance to at least twice the expected in-flight count. Add a head sampler at 50 percent to halve the input, accepting that rare errors are then kept only half the time. Then alert on the processor's eviction and late-span metrics, because both mean decisions are being made on partial traces.
With the policies above, the backend receives every error trace, every trace slower than 1.5 seconds and 5 percent of the rest, typically a small fraction of the input. If errors are 0.5 percent and slow traces 1 percent, about 6.4 percent of traces are kept.
Keeping the numbers honest
Sampled data answers two kinds of question differently. Questions about individual traces, such as why one checkout was slow, need the trace and nothing else. Questions about rates and counts, such as error rate or requests per second, need either unsampled counts or correct weights.
The robust answer is to compute request, error and duration metrics before sampling, for example with the Collector's span metrics connector placed in the agent tier, and to use traces only for exemplars and drill-down. Where you must estimate from sampled spans, weight each span by its adjusted count. A uniform 10 percent sample with weights of 10 gives unbiased counts. A tail sample does not: error traces are kept at 100 percent and the rest at 5 percent, so an unweighted error rate would be grossly inflated. Record the effective probability on each kept trace, for example as an attribute, so weights survive tail decisions too.
def estimated_counts(spans):
total = errors = 0.0
for s in spans:
w = s.attributes.get("sampling.adjusted_count", 1.0) # 1 / probability
total += w
errors += w if s.status == "ERROR" else 0.0
return total, errors, errors / total if total else 0.0The attribute name in that snippet is a local convention, not a standard field; what matters is that every stage that changes probability updates it.
Failure modes
- Broken traces from mixed decisions. One service samples independently instead of following the parent, or two collectors hash with different seeds. Detect with a job that counts traces whose spans reference missing parents.
- Eviction before decision.
num_tracesis too small for the window, so long traces are dropped first. Watch eviction counters. - Split traces on scale-out. When the tail tier scales, the load-balancing ring changes and in-flight traces are split across two instances. Scale during quiet periods and keep the late-span decision cache.
- Late and asynchronous spans. Queue consumers and batch jobs finish after the decision window. Either link them as separate traces or accept that they follow the cached decision.
- Biased dashboards. Rates computed from tail-sampled traces without weights.
- Mixed SDK versions. Older SDKs that do not set the random flag or the
thvalue break consistent sampling assumptions downstream.
Choosing a strategy
For a small system, head sampling with a ParentBased ratio sampler and metrics computed before sampling is enough, and it costs almost nothing. Add tail sampling when you need every error and slow trace and can run a stateful tier. Add consistent probability sampling when different teams need different rates without breaking traces. Add adaptive rate control, described in adaptive trace sampling, when traffic varies so much that a fixed rate either overspends at peak or keeps too little at night.
What to do next
- Measure spans per second, bytes per span and the p99 trace duration for your busiest services.
- Move request, error and duration metrics ahead of any sampling so dashboards never depend on sampled traces.
- Configure ParentBased head sampling everywhere and verify propagation with a broken-trace detector.
- If you add tail sampling, deploy the load-balancing tier first, size
num_tracesfrom your measurements and alert on evictions. - Record the effective sampling probability on every kept trace and weight any estimates by it.
- Check whether your SDKs and Collector support the th and rv keys before adopting consistent probability sampling across teams.