A trace is a tree. Every span has at most one parent, and that single edge is what lets a tracing UI draw a waterfall. Real systems are not always trees. A consumer reads fifty messages that were produced by fifty different requests and handles them in one database transaction. A nightly job processes rows that thousands of users created during the day. A workflow is retried an hour after it first failed. None of these has one honest parent.

Span links are the second kind of edge. A link is a pointer from one span to another span's context, usually in a different trace, with its own attributes. It does not change the tree, it does not move timing, and it does not make the two traces one. This article explains the data model, shows when to choose a link over a parent, builds a batch consumer that links to every message's producer, and covers what happens to links in sampling, storage and queries. It assumes you know spans and context propagation from distributed tracing architecture and trace context propagation.

Advertisement

The Link data model

In OpenTelemetry a link has two parts. The first is a span context: the 16-byte trace ID, the 8-byte span ID, the trace flags (including the sampled bit) and the trace state of the target span. The second is a set of attributes that describe why the link exists. On the wire, OTLP stores links as a repeated field on the span, next to a dropped_links_count that records how many were discarded by limits.

Three properties follow from this shape, and most design mistakes come from forgetting one of them.

  • Links are one-way. The link lives on the span that declares it. The target span is not modified and is usually already exported by the time the link is created. Walking from a producer to its consumers therefore needs an index or a query, not a pointer.
  • Links carry no timing. A child span is expected to sit inside its parent's lifetime in a waterfall. A link target can be minutes or days older, or in a trace that was never stored.
  • Links do not merge traces. The linking span keeps its own trace ID. Sampling, retention and trace-level aggregation all treat the two traces separately.
QuestionParent edgeLink
How many per span?Zero or oneZero or more, bounded by a limit
Same trace?AlwaysUsually not, though it can be
Drawn asNesting in a waterfallA clickable reference to another trace
Used by parent-based sampling?YesNo
When can it be set?Only at span startAt start (preferred) or later

When a parent is the wrong edge

Use a parent when the new work is a causal, synchronous-enough part of one request and the waterfall should show it. Use a link when any of these is true.

  • Fan-in. One unit of work depends on many upstream contexts: a batch consumer, an aggregation window, a compaction job, a bulk API call that carries items from several callers. Picking one of them as parent is arbitrary and hides the rest.
  • Long or unbounded gaps. A message that sits in a queue for six hours would stretch the producer's trace to six hours and break duration statistics. Start a new trace on the consumer and link back.
  • Retries and resumptions. A retried workflow step or a resumed saga gets a fresh trace that links to the failed attempt, so each attempt has a clean timeline but the history is one click away.
  • Trust boundaries. At a public edge you may not want an outside caller to choose your trace ID or sampling decision. Start a new root and link to the inbound context, so it is recorded but not obeyed.

The OpenTelemetry messaging semantic conventions encode the first two cases. They note that a span can have only one parent, make links the default way to correlate producers and consumers for batches, and say a process or receive span for a batch SHOULD link to each message's creation context.

Advertisement

Architecture of a linked pipeline

The diagram shows the standard fan-in shape. Each producer runs inside its own trace and injects a traceparent header into the message. The broker carries the header unchanged. The consumer polls a batch, extracts one context per message, and starts a single processing span in a new trace with one link per extracted context. Child spans for the database write and any downstream publish hang off the processing span as normal children.

Span links: one consumer span, many producer traces, no shared parentProducer Atrace 4bf9...Producer Btrace 9c1e...Producer Ctrace 07ad...Brokertraceparent in headerssendsendsendprocess orders (batch)new trace, 3 linkspoll 3Child spansdb write, publishparentTrace backendlinks stored on consumerexport OTLPA link points backwards only: the consumer span knows its producers, the producers do not know it.Finding consumers from a producer needs a query over link fields, not a tree walk.
Three producer traces send messages with traceparent headers. The consumer starts one span in a new trace with three links, and its children are ordinary parent-child spans. The backend stores the links on the consumer span only.

The conventions also say the producer SHOULD attach a creation context to each message, ideally where intermediaries cannot change it, and that if a create span exists its context is the one to inject; otherwise the send span's context is used. That matters when a producer batches sends: inject per message at creation time, not once per network request, or every message in the batch will point at the same span.

Worked example: a Kafka batch consumer in Python

The producer injects the current context into a dictionary and copies it into Kafka headers. The consumer reverses that per message, collects valid span contexts, and passes them as links when it starts the processing span. Passing links at start is deliberate, for reasons covered in the sampling section.

from confluent_kafka import Consumer, Producer
from opentelemetry import propagate, trace
from opentelemetry.trace import Link, SpanKind

tracer = trace.get_tracer("orders")

def send(producer: Producer, order: bytes) -> None:
    with tracer.start_as_current_span("send orders", kind=SpanKind.PRODUCER):
        carrier: dict[str, str] = {}
        propagate.inject(carrier)               # writes traceparent (+ tracestate)
        producer.produce("orders", order,
                         headers=[(k, v.encode()) for k, v in carrier.items()])

def links_for(batch) -> list[Link]:
    links = []
    for msg in batch:
        carrier = {k: v.decode() for k, v in (msg.headers() or [])}
        ctx = propagate.extract(carrier)
        sc = trace.get_current_span(ctx).get_span_context()
        if sc.is_valid:                          # skip messages with no or bad header
            links.append(Link(sc, {"messaging.kafka.offset": msg.offset()}))
    return links

def run(consumer: Consumer) -> None:
    while True:
        batch = consumer.consume(num_messages=50, timeout=1.0)
        batch = [m for m in batch if m.error() is None]
        if not batch:
            continue
        with tracer.start_as_current_span(
            "process orders",
            kind=SpanKind.CONSUMER,
            links=links_for(batch),              # links at creation: visible to samplers
            attributes={
                "messaging.system": "kafka",
                "messaging.operation.type": "process",
                "messaging.destination.name": "orders",
                "messaging.batch.message_count": len(batch),
            },
        ):
            write_batch_to_db(batch)             # children of the process span
            consumer.commit(asynchronous=False)

Notice what is not done: the consumer does not attach the extracted context, so the processing span is a new root. The link attribute records the offset so a reader can tell which message each link belongs to. The batch count is set only because this span covers a batch; the conventions say not to set it on single-message spans.

Go and Java have the same shape. In Go you pass trace.WithLinks(trace.LinkFromContext(ctx)) to tracer.Start for each extracted context. In Java you call spanBuilder.addLink(spanContext, attributes) before startSpan(). Instrumentation libraries for Kafka, SQS and similar clients increasingly create these links for you; check what yours emits before adding a second set by hand.

Links and sampling

Sampling is where links surprise people. The tracing API specification requires that links can be recorded at span creation and that links can be added after creation, and it states that links added at creation may be considered by samplers while links added later may not. The default ParentBased sampler looks only at the parent, so a new consumer root is sampled by its own root sampler regardless of whether the producers were sampled.

With 10% head sampling on both sides, a batch of 50 messages has about 5 sampled producer traces, and the consumer trace itself is sampled 10% of the time. When you open a sampled consumer trace, most of its links point at traces that were never stored. That is not a bug, but it should shape your UI expectations and your sampler.

If following links matters, a custom sampler can keep a consumer span whenever any linked context was sampled. It only works for links passed at creation, which is another reason to collect them before starting the span.

# Pseudocode for a link-aware root sampler (wraps your normal sampler)
def should_sample(parent_ctx, trace_id, name, kind, attributes, links):
    if parent_span_context(parent_ctx).is_valid:
        return delegate.should_sample(...)        # keep normal parent-based behaviour
    if any(link.context.trace_flags.sampled for link in links):
        return RECORD_AND_SAMPLE                  # keep the consumer if any producer was kept
    return ratio_sampler.should_sample(...)

Two caveats apply. First, with large batches the chance that at least one link is sampled is high, so this sampler can raise consumer volume sharply; with 50 links and 10% producer sampling, roughly 99% of batches have a sampled link. Cap it, for example by requiring a sampled link and a coin flip, or by only honouring links that carry a priority attribute. Second, tail sampling in a collector groups spans by trace ID, so a tail policy for the consumer trace cannot see whether the producer trace was kept. The tail sampling architecture article covers routing by trace ID; links cross that boundary by design. For how head decisions are made and propagated, see head-based sampling.

Limits, storage and querying

SDKs bound links per span. The specification's environment variables are OTEL_SPAN_LINK_COUNT_LIMIT and OTEL_LINK_ATTRIBUTE_COUNT_LIMIT, each defaulting to 128. A consumer that polls 500 messages and links them all will silently keep 128 and count the rest in dropped_links_count. Either raise the limit deliberately, split the work into smaller processing spans, or link a representative subset and record the full count as an attribute.

Each link costs roughly the size of a span context plus its attributes, and it is stored with the consumer span. At high batch rates links can be a noticeable share of span bytes, so treat link attributes like span attributes: a few low-cardinality keys plus the one identifier that distinguishes the message.

Backends differ in what they do with links. Jaeger's model has span references, and the OpenTracing-era FOLLOWS_FROM reference served the same purpose before OpenTelemetry. Grafana Tempo exposes links in TraceQL through the link:traceID and link:spanID intrinsics plus link.-scoped attributes, which is what makes the reverse question answerable: given a producer trace, which consumer spans point at it?

{ link:traceID = "4bf92f3577b34da6a3ce929d0e0e4736" }

Without a query like that, a producer trace looks like a dead end. Make sure the on-call runbook says how to find the consumer side, because the waterfall will not show it.

Failure modes

  • Accidental parenting. The consumer attaches the first message's extracted context and starts the batch span under it. One producer trace now contains a long consumer subtree and the other 49 are invisible. Start a new root and link instead.
  • Links added after start. Code that starts the span first and loops add_link over messages produces correct data but the sampler never saw the links, so a link-aware sampler silently does nothing.
  • Silent truncation. Batches above the link limit lose links without an error. Alert on non-zero dropped_links_count if your backend exposes it.
  • Lost headers. A proxy, a connector or a re-publishing step drops message headers, so is_valid is false and the link is skipped. Count messages without a valid context as a metric; a sudden rise is a broken hop.
  • Same span for every message. The producer injects once per network batch, so all links point at one send span. Inject at message creation.

Trade-offs: one batch span or one span per message

There are two reasonable shapes for a batch consumer, and the messaging conventions allow both.

ShapeWhat you getWhat it costs
One process span with N linksOne span per batch, accurate batch duration, cheapPer-message errors and latency must be attributes or events; link limits apply
Receive span with links, plus a process span per message parented to each producerProducer traces show their own consumption; per-message errors are visibleN extra spans per batch; producer traces stretch across queue time
Hybrid: batch span with links, per-message child spansPer-message detail inside one consumer traceN spans per batch, and still one-way navigation to producers

A good default is the first shape for high-volume pipelines and the hybrid for low-volume, high-value ones such as payments, where each message's outcome matters enough to pay for a span. Avoid parenting per-message spans into producer traces when queue time can be long; it distorts latency statistics for the producer service.

Whatever you choose, keep attribute names consistent with the semantic conventions so dashboards and queries work across services.

Operating linked traces

  • Emit a metric for messages consumed without a valid creation context, per topic and per producer service. It is the cheapest early warning that propagation broke.
  • Document the reverse lookup query for your backend in the runbook, next to the service's dashboards.
  • Review link limits whenever batch sizes change; a configuration change from 100 to 500 messages per poll quietly drops three quarters of the links.
  • Test propagation end to end in CI: produce a message inside a known trace, consume it, and assert the consumer span has a link with that trace ID.

What to do next

  1. List every place in your system where work fans in, waits in a queue, is retried or crosses a trust boundary, and mark which edge each one uses today.
  2. Change those spans to start a new root and pass one link per upstream context at creation time, with a small identifying attribute on each link.
  3. Check your instrumentation libraries for links they already create, so you do not duplicate them.
  4. Decide your sampling policy for linked spans, and if you add a link-aware sampler, cap how much it can raise volume.
  5. Set OTEL_SPAN_LINK_COUNT_LIMIT to match your largest batch, or split batches, and alert on dropped links.
  6. Write and test the backend query that finds consumers from a producer trace, and put it in the runbook.
Key takeaway: A span link records that one span depends on another span's context without making it the parent. Use links for fan-in, long queue gaps, retries, trust boundaries and scheduled work; use parents for synchronous causal work. Pass links when the span starts so samplers can see them, keep link attributes small, respect the default limit of 128 links per span, and give on-call engineers a query that walks links backwards, because the waterfall never will.