A distributed trace is a record of causality. It says that this database call happened because of that service call, which happened because of a user's click, even though the three ran in different processes on different machines with different clocks. OpenTelemetry supplies the vocabulary and the libraries, but whether a trace is complete and truthful depends on distributed-systems properties: how context crosses process boundaries, what happens at trust boundaries, how work that resumes hours later stays connected, and how much to believe span timestamps taken on different hosts.
This article treats tracing from that angle. It explains the causal model, the W3C propagation contract and where it breaks, shows code for carrying context through durable work and across untrusted ingress, explains clock skew in span timing, and shows how to measure trace completeness as a number. The SDK pipeline, the Collector and sampling strategies each have their own articles, linked at the end. Specification details were checked on 2026-10-03.
A trace as a causal graph
The span model: causality, not time
Each span records one operation: a name, a start and end time, attributes, a status, and three identifiers. The trace id, 16 bytes, is shared by every span in the trace. The span id, 8 bytes, names this span. The parent span id names the span that caused it. Parent edges form a tree, and that tree is the trace. A span can also carry links to spans in the same or other traces. A link says that one operation is related to another without being its child, which is how you model fan-in, where one batch job serves many requests, and deferred work, where the causing request finished long ago.
Two consequences follow. First, the tree encodes happened-because, not happened-before. Ordering spans by timestamp across hosts is unreliable, which this article returns to below, but ordering them by parent edges is exact. Second, a missing parent edge cannot be repaired afterwards. If a service fails to pass context on, its spans start a new trace, and no backend can reliably stitch the two halves together. Tracing quality is decided at propagation time.
The propagation contract
Context crosses a process boundary as headers defined by the W3C Trace Context specification. The traceparent header has four dash-separated fields:
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
| | | |
| trace id, 32 hex (16 bytes) parent id, flags, 8 bits:
version 16 hex (8 bytes) 0x01 sampled, 0x02 random
tracestate: ot=th:c,vendor1=opaque-value
baggage: tenant.id=t-42,request.class=refundThe sampled flag tells downstream services that the caller may have recorded its span, so they should record theirs to keep the trace whole. The random flag, added in Trace Context Level 2 (a W3C Candidate Recommendation Draft dated March 2024), promises that at least the right-most 7 bytes of the trace id are random; samplers can then decide from the trace id alone and reach the same decision everywhere. tracestate carries vendor-specific key-value pairs, at most 32 members, and implementations should propagate at least 512 characters of it. baggage is a separate W3C header for application key-value pairs that every downstream service can read.
In OpenTelemetry the code that reads and writes these headers is a propagator, configured with OTEL_PROPAGATORS, where tracecontext,baggage is the usual default. Instrumentation libraries call it on every inbound and outbound request. The contract is only as strong as its weakest hop: every proxy, queue, scheduler and SDK on the path must carry these headers through unchanged.
Where propagation breaks
In practice, traces break at a predictable set of places.
- Infrastructure that drops headers. Some proxies, API gateways and message brokers forward only allow-listed headers. One misconfigured hop splits every trace into two.
- Mixed propagation formats. One service speaks B3, another W3C, and without a composite propagator neither recognises the other's context.
- Work that outlives the request. Jobs written to a table, timers in a workflow engine and retries scheduled for later have no live context unless you store it with the work.
- Thread and task hand-offs. Context lives in a thread-local or async context variable; work submitted to a pool without wrapping loses it.
- Trust boundaries. Context from outside your organisation cannot be trusted: a caller can choose trace ids, force sampling on, or send baggage that leaks into your logs.
Durable and delayed work
When work is persisted and resumed later, store the context with it. The question is then whether the resumed work should be a child of the original span or a new trace that links to it. Use a child when the work is part of the same user-visible operation and runs soon, within a timeout you would accept in one trace view. Use a new trace with a link when it runs much later, runs many times, or serves many requests at once; otherwise one trace grows for hours, never completes in the backend, and confuses latency analysis.
import json
from opentelemetry import trace, propagate
from opentelemetry.trace import Link, SpanKind
tracer = trace.get_tracer("refunds")
def schedule_refund_check(db, refund_id, run_at):
carrier = {}
propagate.inject(carrier) # current traceparent, tracestate, baggage
db.execute("INSERT INTO jobs (refund_id, run_at, trace_ctx) VALUES (%s, %s, %s)",
(refund_id, run_at, json.dumps(carrier)))
def run_refund_check(job):
origin = trace.get_current_span(propagate.extract(json.loads(job.trace_ctx)))
# Hours later: start a fresh trace, but keep the causal edge as a link.
with tracer.start_as_current_span(
"refund.check", kind=SpanKind.CONSUMER,
links=[Link(origin.get_span_context(), {"app.link.reason": "scheduled"})]):
check_settlement(job.refund_id)Storing the full carrier, not just the trace id, preserves tracestate and baggage, so sampling and tenant attributes survive the delay. Columns like trace_ctx are small, but treat baggage inside them as data that may be read later by other tools.
Trust boundaries
At an ingress from outside your trust domain, do not continue the caller's trace. Start a new root span, record the caller's context as a link so the relationship is still visible, and discard inbound baggage. This prevents external callers from forcing every request to be sampled, from colliding trace ids, or from injecting values that your services copy into logs.
from opentelemetry import context as otel_context
TRUSTED = {"internal-mesh", "batch-runner"}
def handle_ingress(headers, caller_identity, handler):
incoming = trace.get_current_span(propagate.extract(headers)).get_span_context()
if caller_identity in TRUSTED:
ctx, links = propagate.extract(headers), []
else:
ctx = otel_context.Context() # empty: no parent, no baggage
links = [Link(incoming, {"app.link.reason": "untrusted-ingress"})] if incoming.is_valid else []
with tracer.start_as_current_span("ingress", context=ctx, kind=SpanKind.SERVER, links=links):
return handler()Apply the same rule in the other direction. Calls to a partner's API carry your traceparent only if you are content for them to see your trace ids and sampling decisions; many teams strip baggage on egress, because it often contains tenant or user identifiers.
Clock skew and span timing
Span start and end times come from each host's wall clock. Within one process, a span's duration is accurate because both ends use the same clock. Across hosts, comparing a child's start to its parent's start mixes two clocks, and the result includes their offset. If the server's clock runs 30 ms ahead of the client's, a 10 ms server span appears to start 30 ms after the client began waiting, and a server whose clock runs behind can appear to start before its caller.
The useful rule is to trust durations within a span, trust parent edges for ordering, and treat cross-host gaps smaller than your clock synchronisation error as noise. If your fleet keeps clocks within 5 ms, a 3 ms gap between a client span and its server span means nothing, while a 300 ms gap is real time spent in a network, a proxy queue or a connection pool. Some backends adjust skew for display; the stored timestamps are unchanged. A child whose start precedes its parent's start by more than your error bound is itself a signal that a host's clock is wrong, which the completeness job below can count.
Sampling as a distributed decision
Sampling is a distributed decision too: a trace is only whole if every service makes the same choice. Head sampling achieves that by deciding once at the root and passing the sampled flag downstream. Probability samplers can decide independently and still agree when they derive the decision from the trace id: OpenTelemetry's probability sampling uses the low 56 bits of the trace id as randomness, which is why the random flag matters, and records the threshold in tracestate as ot=th:.... Tail sampling decides after the trace ends, so every span of a trace must reach the same Collector instance; the Collector's load-balancing exporter routes by trace id for exactly this reason. The full treatment, including keeping counts honest, is in the sampling article linked below.
Measuring trace completeness
Trace completeness can be measured. Export a sample of spans to a store you can query and count three things per service pair: orphans, spans whose parent span id never arrives; roots per request, where more than one root for one inbound request means a propagation break; and skew violations, children starting before their parents by more than the error bound.
def completeness(spans, skew_bound_ms=5):
by_id = {s["span_id"]: s for s in spans}
orphans, skew, total = 0, 0, 0
for s in spans:
if not s["parent_span_id"]:
continue # a root
total += 1
parent = by_id.get(s["parent_span_id"])
if parent is None:
orphans += 1 # parent lost, dropped or never sent
elif (parent["start_ms"] - s["start_ms"]) > skew_bound_ms:
skew += 1 # child before parent: clock problem
return {"orphan_ratio": orphans / max(total, 1), "skew_ratio": skew / max(total, 1)}Run it over complete traces only, after the backend's assembly delay, and group the results by the orphan's service and the parent's expected service. A service pair with a sudden rise in orphans points at one hop that stopped propagating, usually after a proxy, SDK or broker upgrade.
Worked example: the broken refund trace
A refund flow crosses an edge gateway, a refund API, a payments service, a partner bank and a scheduler that re-checks settlement two hours later. Support engineers complain that refund traces are useless: some end at the API, some are hours long, and some show the payments call starting before the API called it.
Measure. The completeness job shows 40 percent orphans for spans whose parents should be in the refund API, a single multi-hour trace per refund, and a 2 percent skew rate concentrated on one payments host.
Fix propagation. The orphans trace back to an internal API gateway that forwards an allow-list of headers; adding traceparent, tracestate and baggage drops orphans below 1 percent.
Fix durable work. The scheduler continued the original trace as a child, so every refund produced one trace lasting two hours. Switching to the stored-context-plus-link pattern gives a short request trace and a separate settlement trace linked to it.
Fix trust. The gateway continued traces from the partner's callbacks, and the partner sets the sampled flag on everything, inflating span volume. The ingress pattern starts new roots for partner callbacks with a link.
Fix time. The skewed host's time synchronisation daemon had stopped; restarting it and alerting on clock offset clears the skew violations.
Failure modes
- Split traces from a header-dropping hop. Detect with orphan ratios per service pair.
- Endless traces from continuing context into delayed or repeated work. Use links.
- Forced sampling from untrusted callers setting the sampled flag. Restart traces at the trust boundary.
- Baggage leaks of user or tenant identifiers to partners or logs. Strip on egress and review what you put in it.
- Misleading timelines from clock skew. Trust durations and parent edges; alert on clock offset.
- Incomplete tail decisions when spans of one trace reach different Collector instances. Route by trace id.
Trade-offs
Every choice here trades connection against control. Continuing context everywhere gives the longest, most connected traces and the most exposure to bad inputs and unbounded trace lengths. Restarting at boundaries and linking keeps traces bounded and safe, at the cost of following links across traces in the backend, which not every backend renders well. Storing context with durable work costs a column and a little discipline. Measuring completeness costs a batch job and a span export, and it turns tracing quality from an impression into a number you can alert on.
Related reading: OpenTelemetry in depth for the API, SDK and exporters, trace sampling strategies for head, tail and consistent probability sampling, OpenTelemetry pipeline design for Collector tiers, OpenTelemetry for Java for context across threads, and clocks in distributed systems for why timestamps cannot order events.
What to do next
- List every hop on your three most important request paths, including proxies, queues and schedulers, and confirm each forwards
traceparent,tracestateandbaggage. - Set
OTEL_PROPAGATORSconsistently across services, with a composite propagator wherever legacy formats remain. - Store the full propagation carrier with any persisted work, and choose child or link by delay and fan-in.
- Restart traces at untrusted ingress with a link, and strip baggage on egress to partners.
- Export a span sample and compute orphan and skew ratios per service pair; alert on changes.
- Alert on host clock offset, and treat cross-host gaps below that bound as noise.
- If you tail-sample, route spans to Collector instances by trace id.