OpenTelemetry is a vendor-neutral standard for producing telemetry: APIs and SDKs that create traces, metrics and logs, a wire protocol (OTLP) that carries them, and a Collector that receives, processes and forwards them. Each piece is easy to adopt on its own. The value only appears when every tier of a system takes part, because the question you need answered during an incident is about a whole request: the user clicked Pay, and something, somewhere, took nine seconds.

This page wires one request through a realistic stack: a browser, a Python checkout API, a Kafka topic, a fulfilment worker and Postgres, with Collectors in between. Each tier gets the code it needs and the trap it usually falls into. The deep dives on single pieces live elsewhere: trace context propagation, Collector pipeline design and semantic conventions.

Advertisement

The goal: one trace id, three signals

Full-stack instrumentation succeeds when three things hold. Every tier joins the same trace, so a request is one tree of spans, not five disconnected fragments. Every signal from a tier carries the same resource identity, so a trace, a metric and a log can be matched to the same service and version. And logs and metrics point back at traces: log records carry the trace and span id, and latency histograms carry exemplars. Everything below serves one of those three.

Browserweb SDK, fetchCheckout APIPython, auto + manualKafkaorders topicFulfilmentconsumer workerPostgresclient spanstraceparentheadersextractPostgresclient spansCollector agentper node: limits, k8s attrsCollector agentper nodeOTLP/HTTP via edgeOTLPOTLPCollector gatewaytail sampling, redactionby trace idby trace idTraces backendMetrics and logsbackendsOne trace id flows left to right through every tier; telemetry flows down through the collectors.
The reference stack. Context flows left to right with the request; telemetry flows down to the agents, then to a gateway that routes by trace id.

Start with resource identity

A resource describes the thing producing telemetry. Decide its attributes before you instrument anything, because changing them later breaks every dashboard and query. The minimum: service.name (stable, lowercase, one per deployable), service.version (the build, so you can compare releases) and deployment.environment.name (older semantic-convention versions call it deployment.environment; pick the one your backend expects and use it everywhere).

Set these through OTEL_SERVICE_NAME and OTEL_RESOURCE_ATTRIBUTES in the deployment manifest rather than in code, so the same image gets the right identity in every environment. Let the Collector add infrastructure attributes such as pod, node and namespace with the Kubernetes attributes processor, rather than each SDK detecting them.

Advertisement

Tier 1: the browser

The trace should start where the user waits. The web SDK creates a span for each fetch and injects a traceparent header, so the API's spans become children of the browser's.

import { WebTracerProvider } from "@opentelemetry/sdk-trace-web";
import { BatchSpanProcessor } from "@opentelemetry/sdk-trace-base";
import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-http";
import { resourceFromAttributes } from "@opentelemetry/resources";
import { registerInstrumentations } from "@opentelemetry/instrumentation";
import { FetchInstrumentation } from "@opentelemetry/instrumentation-fetch";

const provider = new WebTracerProvider({
  resource: resourceFromAttributes({
    "service.name": "shop-web",
    "service.version": BUILD_SHA,
  }),
  spanProcessors: [
    new BatchSpanProcessor(new OTLPTraceExporter({ url: "https://telemetry.example.com/v1/traces" })),
  ],
});
provider.register();

registerInstrumentations({
  instrumentations: [
    new FetchInstrumentation({
      // Only first-party APIs: never send traceparent to third parties.
      propagateTraceHeaderCorsUrls: [/^https:\/\/api\.example\.com\//],
    }),
  ],
});

Two traps live here. First, browsers do not send custom headers on cross-origin requests unless the server allows them: the API's CORS preflight response must list traceparent and tracestate in Access-Control-Allow-Headers. If it does not, the preflight fails and the request breaks, which is worse than missing a trace. Second, the export endpoint is public. Put it behind an edge that rate-limits, drops oversized payloads and strips anything user-identifying, and never ship a backend API key in browser code.

Tier 2: the services

For most services, start with zero-code instrumentation: it creates server spans for incoming HTTP, client spans for outgoing HTTP and database calls, and propagates context, without touching application code. Then add manual spans only for business steps the libraries cannot see.

# Install once per image:
#   pip install opentelemetry-distro opentelemetry-exporter-otlp
#   opentelemetry-bootstrap -a install
# Run with:
#   OTEL_SERVICE_NAME=checkout-api \
#   OTEL_RESOURCE_ATTRIBUTES=service.version=4.12.0,deployment.environment.name=prod \
#   OTEL_EXPORTER_OTLP_ENDPOINT=http://$NODE_IP:4318 \
#   OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf \
#   OTEL_TRACES_SAMPLER=parentbased_traceidratio OTEL_TRACES_SAMPLER_ARG=1.0 \
#   OTEL_PYTHON_LOG_CORRELATION=true \
#   opentelemetry-instrument gunicorn app:app

from opentelemetry import trace, metrics, propagate

tracer = trace.get_tracer("checkout")
orders = metrics.get_meter("checkout").create_counter("checkout.orders", unit="{order}")

def place_order(cart, producer):
    with tracer.start_as_current_span("checkout.place_order") as span:
        span.set_attribute("checkout.items", len(cart.items))
        order = reserve_and_charge(cart)            # HTTP and DB spans come from auto-instrumentation
        headers = {}
        propagate.inject(headers)                   # writes traceparent (and tracestate)
        producer.produce("orders", key=order.id, value=order.to_json(),
                         headers=list(headers.items()))
        orders.add(1, {"payment.method": cart.payment_method})   # bounded values only
        return order

The sampler parentbased_traceidratio follows the caller's sampling decision when there is one and samples new traces at the given ratio. Keep it at 1.0 in services and do the cost-saving in the gateway's tail sampler, so the decision is made with the whole trace in view. OTEL_PYTHON_LOG_CORRELATION adds the trace and span ids to standard-library log records, which is the cheapest correlation you will ever get.

Keep metric attribute values bounded. A payment method has five values; a user id has millions, and one such attribute can multiply a metric's series count past what the backend will store.

Tier 3: asynchronous hops

Context does not cross a message queue by itself. The producer must inject it into message headers, as the checkout code does, and the consumer must extract it. Then decide how the consumer span relates to the producer. When one message causes one unit of work, make the producer context the parent, so the trace stays one tree. When a worker processes a batch of messages together, a single parent is impossible; start a new trace and attach a link to each producer's span context.

from opentelemetry import trace, propagate
from opentelemetry.trace import SpanKind, Link

tracer = trace.get_tracer("fulfilment")

def handle(msg):
    carrier = {k: v.decode() for k, v in (msg.headers() or [])}
    producer_ctx = propagate.extract(carrier)
    producer_span = trace.get_current_span(producer_ctx).get_span_context()
    # One message, one unit of work: continue the producer's trace as parent.
    with tracer.start_as_current_span(
            "process orders", context=producer_ctx, kind=SpanKind.CONSUMER,
            attributes={"messaging.system": "kafka",
                        "messaging.destination.name": "orders"}):
        fulfil(msg.value())

def handle_batch(msgs):
    # Many messages, one unit of work: a new trace that links to every producer.
    links = []
    for m in msgs:
        ctx = propagate.extract({k: v.decode() for k, v in (m.headers() or [])})
        links.append(Link(trace.get_current_span(ctx).get_span_context()))
    with tracer.start_as_current_span("process orders", kind=SpanKind.CONSUMER, links=links):
        fulfil_many([m.value() for m in msgs])

Links preserve causality without pretending a batch has one cause; the trade-offs are covered in span links. Whichever you choose, use it consistently per consumer, because trace-search tools treat the two very differently.

Tier 4: the database

Database client libraries instrumented by OpenTelemetry produce a client span per query with the system, operation and, optionally, the query text. Database semantic conventions have been renamed across versions, for example from db.system to db.system.name and from db.statement to db.query.text, so check which names your instrumentation emits before you build queries on them. Make sure query text is sanitised, with literal values replaced by placeholders, because a logged WHERE email = ... is personal data sitting in your trace store.

Spans from the database server itself are rarely available. The client span's duration is what the application waited, and that includes connection-pool wait. Instrument the pool too, or a saturated pool will look like a slow database.

Correlating logs and metrics with traces

With trace ids on log records, a trace view can show the logs from every span, and a log line can jump to its trace. Export logs through OTLP, or keep your existing log shipper and make sure it parses the trace id field. For metrics, exemplars attach a sampled trace id to a histogram bucket, so a latency spike on a dashboard links to a real slow request; see metrics exemplars. Exemplars only help if the exemplar's trace was kept, which is one more reason to keep slow and failed traces in the tail sampler.

The Collector: agents and a gateway

Run two tiers. A per-node agent receives OTLP from local pods on the standard ports (4317 for gRPC, 4318 for HTTP), protects itself with a memory limiter, enriches with Kubernetes attributes and forwards. A gateway tier makes decisions that need a whole trace, such as tail sampling and redaction, and exports to the backends.

# Agent: one per node (DaemonSet). Receives from local pods, enriches, forwards.
receivers:
  otlp:
    protocols:
      grpc: {endpoint: 0.0.0.0:4317}
      http: {endpoint: 0.0.0.0:4318}
processors:
  memory_limiter: {check_interval: 1s, limit_percentage: 80, spike_limit_percentage: 20}
  k8sattributes: {}
  batch: {}
exporters:
  loadbalancing:                     # all spans of a trace go to the same gateway pod
    routing_key: traceID
    protocol:
      otlp: {tls: {insecure: true}}
    resolver:
      dns: {hostname: otel-gateway-headless.observability.svc.cluster.local}
  otlp/gateway:
    endpoint: otel-gateway.observability.svc.cluster.local:4317
    tls: {insecure: true}
service:
  pipelines:
    traces:  {receivers: [otlp], processors: [memory_limiter, k8sattributes, batch], exporters: [loadbalancing]}
    metrics: {receivers: [otlp], processors: [memory_limiter, k8sattributes, batch], exporters: [otlp/gateway]}
    logs:    {receivers: [otlp], processors: [memory_limiter, k8sattributes, batch], exporters: [otlp/gateway]}
# Gateway: a horizontally scaled Deployment. Decides which traces to keep.
# Partial: the otlp receiver and the otlp/traces and debug exporters are defined as usual.
# loadbalancing and tail_sampling ship in the contrib distribution, not core.
processors:
  memory_limiter: {check_interval: 1s, limit_percentage: 80, spike_limit_percentage: 20}
  tail_sampling:
    decision_wait: 10s
    policies:
      - {name: errors,   type: status_code,   status_code: {status_codes: [ERROR]}}
      - {name: slow,     type: latency,       latency: {threshold_ms: 1500}}
      - {name: baseline, type: probabilistic, probabilistic: {sampling_percentage: 5}}
service:
  pipelines:
    traces: {receivers: [otlp], processors: [memory_limiter, tail_sampling], exporters: [otlp/traces, debug]}

Tail sampling only works if every span of a trace reaches the same gateway instance, which is why the agent's trace pipeline uses the load-balancing exporter keyed on trace id. The debug exporter prints what passes through and belongs on a canary instance, not on every replica. The memory limiter goes first in every pipeline. The batch processor is shown because it is widely deployed; the Collector project now favours batching inside the exporter's sending queue, which, backed by a persistent queue, survives restarts better, so check your Collector version's exporter documentation before choosing.

Worked example: the nine-second checkout

An alert fires: checkout p99 latency has tripled. The histogram's exemplar opens a kept trace. The browser span is 9.2 seconds; the API's server span is 9.1, so the network is not the problem. Inside it, the database span for the stock reservation is 40 ms, but a gap of 8.8 seconds precedes it. The pool-wait span fills the gap: the connection pool is exhausted. Following the span link from the fulfilment worker's batch span shows the cause: a new release of the worker, identified by service.version, holds connections across a slow external call. Logs filtered by that trace id show the pool-timeout warnings. Rolling the worker back clears it. Each step relied on one tier's instrumentation; with any tier missing, the chain breaks at that point.

Failure modes

  • Broken traces at a hop. A proxy, queue or thread pool that does not carry context splits one request into several traces. Test propagation per hop, not per service.
  • Inconsistent sampling. Services sampling independently at their own ratios produce traces with missing middles. Use parent-based samplers and sample in one place.
  • Identity drift. Two names for one service, or a version attribute set in code and forgotten, splits dashboards and hides regressions.
  • Cardinality blow-ups. Unbounded attribute values on metrics, or span names containing ids, overwhelm backends and bills.
  • Collector as a single point of failure. An under-sized gateway without a memory limiter crashes under a traffic spike and loses exactly the data from the incident you need.
  • Sensitive data in telemetry. Query text, headers and URLs carry personal data unless redacted in the SDK or the gateway.

Rolling it out

Instrument along one critical user journey end to end before widening coverage; one complete trace teaches more than twenty services with server spans only. A practical order: agree resource identity; deploy agents and a gateway with the debug exporter on a canary; enable zero-code instrumentation on the services in that journey; add propagation across the queue; add the browser; then logs correlation and exemplars; then tail-sampling policies tuned on real traffic. The trade-off throughout is completeness against cost: keep everything at first on one journey, measure the volume, and only then set sampling percentages from data rather than guesses.

What to do next

  1. Write down the resource attributes and service names for every deployable, and set them in manifests through environment variables.
  2. Pick one user journey and draw its hops, marking each one as HTTP, queue, thread pool or database.
  3. Deploy a Collector agent per node and a gateway with a memory limiter, the load-balancing exporter by trace id, and a debug exporter on one canary.
  4. Enable zero-code instrumentation on the journey's services with a parent-based sampler at 1.0, and add manual spans for the business steps.
  5. Inject and extract context across every queue, choosing parent or links per consumer, and add traceparent and tracestate to the API's CORS allow list before enabling browser tracing.
  6. Turn on log correlation and exemplars, then add tail-sampling policies that keep errors and slow traces, and confirm one alert-to-trace-to-log path works before the next incident.
Key takeaway: Full-stack OpenTelemetry means one trace id from the browser to the database and back, shared resource identity across traces, metrics and logs, and Collectors that make sampling decisions with the whole trace in view. Agree on identity first, propagate context explicitly across every queue and CORS boundary, sample in one place, and grow coverage one complete user journey at a time.