Jaeger is an open-source distributed tracing backend: it receives spans from instrumented services, stores them, and gives you a UI and API to search for traces and see where a request spent its time. It began at Uber, became a CNCF graduated project, and has since changed shape in a way many older tutorials miss. Jaeger v2 is built on the OpenTelemetry Collector framework, accepts OTLP natively, and expects services to be instrumented with OpenTelemetry SDKs rather than the retired Jaeger client libraries.

This article is about Jaeger as a system you operate. The tracing data model, span kinds and propagation are covered in Distributed tracing architecture; here the subject is the pieces of a Jaeger deployment, how data moves through them, which storage to pick, how sampling is controlled centrally, and what breaks. Facts follow the Jaeger documentation and repository configuration files as checked on 2026-10-01; configuration keys are quoted from those files.

Advertisement

One binary, several roles

Jaeger v1 shipped separate executables: an agent on every host, a collector, a query service and an ingester. Jaeger v2 replaces them with a single jaeger binary that is a custom distribution of the OpenTelemetry Collector. What it does is decided entirely by its YAML configuration, which uses the Collector's vocabulary of receivers, processors, exporters, extensions and pipelines. The documentation describes four roles you assemble from that vocabulary.

  • Collector: receives spans (OTLP, plus legacy Jaeger and Zipkin protocols) and writes them to storage or to Kafka.
  • Query: serves the HTTP and gRPC query APIs and the web UI, by default on port 16686.
  • Ingester: reads spans from Kafka and writes them to storage, used only in the Kafka-buffered topology.
  • All-in-one: collector and query in one process, usually with in-memory or Badger storage, for development and small installs.

The host agent is gone. If you want a local hop for batching, enrichment or tail sampling, the documentation recommends running a stock OpenTelemetry Collector as a sidecar or gateway in front of Jaeger. Because Jaeger is itself a Collector build, any skills and processors you know from the OpenTelemetry Collector transfer directly: memory limiter, batch, attribute processors and the tail sampling processor are all available inside the Jaeger binary.

Jaeger v2: one binary built on the OpenTelemetry Collector, deployed in different rolesService + OTel SDKspans, traceparentOptional OTel Collectorenrich, tail sampleOTLPjaeger (collector role)receivers: otlp, jaeger, zipkinprocessors: batch, adaptive_samplingexporter: jaeger_storage_exporter or kafkaextension: remote_sampling (5778)OTLP 4317/4318Kafka topicjaeger-spansjaeger (ingester role)kafka receiveroptional bufferTrace storage (jaeger_storage extension)Elasticsearch / OpenSearch / Cassandra / ClickHouse / Badger / memory / remote gRPCdirect writewritejaeger (query role)jaeger_query, UI 16686readpoll strategyEvery box labelled jaeger is the same binary with a different YAML pipeline. All-in-one runs collector and query in one process.Adaptive sampling closes a loop: the collector observes throughput, recalculates probabilities and serves them back to SDKs.
Jaeger v2 deployment roles. The same binary acts as collector, ingester or query depending on its pipeline configuration; Kafka is optional.

The extensions that make it Jaeger

A plain Collector has no notion of a trace store or a UI. Jaeger adds them as Collector extensions, and reading the stock config.yaml in the repository is the fastest way to understand the moving parts:

  • jaeger_storage declares named storage backends, for example some_store backed by memory, Elasticsearch or Cassandra. It is a registry; nothing writes to it until an exporter or the query extension refers to a name.
  • jaeger_storage_exporter is the pipeline exporter that writes spans into one named backend through trace_storage.
  • jaeger_query serves the UI and APIs and names which backend to read for traces and, optionally, traces_archive and metrics (the last powers the Monitor tab).
  • remote_sampling serves sampling strategies to SDKs, either from a JSON file or computed adaptively.

A minimal production-style collector-plus-query configuration with Elasticsearch looks like this. It is trimmed from the repository's config-elasticsearch.yaml; the index settings show that spans, services and dependencies get separate daily indices.

service:
  extensions: [jaeger_storage, jaeger_query, healthcheckv2]
  pipelines:
    traces:
      receivers: [otlp]
      processors: [batch]
      exporters: [jaeger_storage_exporter]

extensions:
  healthcheckv2:
    use_v2: true
    http:
  jaeger_query:
    storage:
      traces: main_store
  jaeger_storage:
    backends:
      main_store:
        elasticsearch:
          server_urls: [http://es-0:9200, http://es-1:9200]
          indices:
            index_prefix: "jaeger-main"
            spans:
              date_layout: "2006-01-02"
              rollover_frequency: "day"
              shards: 5
              replicas: 1

receivers:
  otlp:
    protocols:
      grpc:
      http:
        endpoint: "0.0.0.0:4318"

processors:
  batch:

exporters:
  jaeger_storage_exporter:
    trace_storage: main_store

Jaeger also exposes its own metrics in Prometheus format (port 8888 in the stock files). Scrape them: accepted, refused and dropped span counts and exporter queue size are the first numbers you will need when traces go missing.

Advertisement

Choosing trace storage

The documentation lists Badger, Cassandra, ClickHouse, Elasticsearch, OpenSearch, Kafka (as a buffer, not a query store) and memory, and a remote gRPC storage API lets you plug in anything else. The decision is mostly about the queries you need and what your team already operates.

BackendGood forWatch out for
MemoryLocal development, demos, CILost on restart; capped by max_traces
BadgerSingle-node all-in-one with persistenceOne node only; no horizontal scale
Elasticsearch / OpenSearchFlexible tag search, most common choiceIndex count and shard sizing; retention via rollover or index cleanup
CassandraVery high write rates, TTL-based retentionTag search needs indexed tables; less ad hoc querying
ClickHouseColumnar storage, analytical queries over spansNewer backend; check maturity of the features you need

Two sizing rules hold for all of them. First, trace volume is span count times average span size, and span size is dominated by attributes and events, not by timing fields; a service that attaches request bodies to spans can multiply cost on its own. Second, retention is the main cost lever. Seven days covers most incident investigations; keep longer history by archiving chosen traces (the traces_archive backend exists for exactly this) rather than retaining everything.

Direct writes or a Kafka buffer

In the direct topology collectors write straight to storage. It is simple and has the fewest moving parts, but a slow or unavailable store pushes back into the collectors, whose queues fill and then drop spans. In the Kafka topology, collectors export to a topic and ingesters consume it and write to storage. Kafka absorbs bursts and storage maintenance windows, and lets you replay. The repository's collector side is a normal pipeline with a kafka exporter:

exporters:
  kafka:
    brokers: [kafka-0:9092, kafka-1:9092]
    traces:
      topic: jaeger-spans
      encoding: otlp_proto

The ingester is another jaeger instance with a kafka receiver on the same topic and encoding and a jaeger_storage_exporter. Choose Kafka when your span rate is bursty, when storage is shared with other workloads, or when losing spans during a storage upgrade is unacceptable. Do not add it by default: it is another cluster to size, and consumer lag becomes a new reason traces appear late. If you already run Kafka well, the trade is usually worth it above a few tens of thousands of spans per second; below that, a bigger collector queue is cheaper.

Sampling controlled from the backend

Most production systems cannot store every trace, so something decides which ones to keep. Head sampling decides at the root span and propagates the decision in the traceparent flags; tail sampling decides after the whole trace has been seen. The trade-offs are covered in tail sampling architecture. Jaeger's distinctive feature is remote sampling: SDKs poll the backend for their head-sampling strategy, so you change rates without redeploying services.

File-based strategies are a JSON document referenced from the remote_sampling extension. Strategies are probabilistic (a probability between 0 and 1) or ratelimiting (traces per second), set per service, with per-operation overrides that must be probabilistic, and a default_strategy for everything else:

{
  "service_strategies": [
    {
      "service": "checkout",
      "type": "probabilistic",
      "param": 0.2,
      "operation_strategies": [
        { "operation": "GET /health", "type": "probabilistic", "param": 0.0 },
        { "operation": "POST /orders", "type": "probabilistic", "param": 1.0 }
      ]
    },
    { "service": "search", "type": "ratelimiting", "param": 50 }
  ],
  "default_strategy": { "type": "probabilistic", "param": 0.01 }
}

Adaptive sampling goes further. Instead of fixed probabilities, Jaeger recalculates a probability for each service and endpoint so that the volume of collected traces approaches target_samples_per_second. Endpoints with no history start at initial_sampling_probability. It needs two things in the configuration: the adaptive_sampling processor in the traces pipeline, which observes throughput, and the remote_sampling extension in adaptive mode with a sampling_store (memory, Cassandra, Badger, Elasticsearch or OpenSearch) where throughput and probabilities are kept. With several collectors, one wins a lease-based leader election and does the calculation. The concept is covered in adaptive sampling.

extensions:
  remote_sampling:
    adaptive:
      sampling_store: main_store
      initial_sampling_probability: 0.1
    http:
    grpc:

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [batch, adaptive_sampling]
      exporters: [jaeger_storage_exporter]

Instrumenting a service to use it

Services send spans over OTLP to port 4317 (gRPC) or 4318 (HTTP). To use remote sampling, the SDK needs a Jaeger-remote sampler; in Go it lives in go.opentelemetry.io/contrib/samplers/jaegerremote, whose default strategy URL is http://127.0.0.1:5778/sampling. Wrap it in a parent-based sampler so downstream services follow the root's decision instead of sampling again:

exp, _ := otlptracegrpc.New(ctx,
    otlptracegrpc.WithEndpoint("jaeger-collector:4317"),
    otlptracegrpc.WithInsecure())

remote := jaegerremote.New("checkout",
    jaegerremote.WithSamplingServerURL("http://jaeger-collector:5778/sampling"),
    jaegerremote.WithSamplingRefreshInterval(30*time.Second),
    jaegerremote.WithInitialSampler(sdktrace.TraceIDRatioBased(0.05)),
)

tp := sdktrace.NewTracerProvider(
    sdktrace.WithBatcher(exp),
    sdktrace.WithSampler(sdktrace.ParentBased(remote)),
    sdktrace.WithResource(resource.NewWithAttributes(semconv.SchemaURL,
        semconv.ServiceName("checkout"))),
)
otel.SetTracerProvider(tp)
otel.SetTextMapPropagator(propagation.TraceContext{})

Three details matter more than the code. The service.name resource attribute is what Jaeger groups by, so set it explicitly. The initial sampler is what runs before the first successful poll, so a misconfigured URL silently leaves you at that rate forever; alert on it. And the propagator must match across every hop, or traces break into fragments; see trace context propagation.

Worked example: a slow checkout

A team sees checkout p99 rise from 400 ms to 1.3 s after a deploy. Metrics show the regression but not its cause. In the Jaeger UI they search service checkout, operation POST /orders, with a minimum duration of 1 s over the last hour; because the per-operation strategy above samples that route at 1.0, slow requests are present. The scatter plot shows two clusters, near 350 ms and near 1.2 s.

Opening a slow trace, the timeline shows inventory.reserve making eleven sequential calls to SELECT stock, each about 80 ms, where a fast trace shows one batched query. Using trace comparison on one fast and one slow trace highlights the extra spans. The deploy had replaced a batch lookup with a per-item loop. The fix is obvious once seen; the point is that only a trace shows the shape of the request. Afterwards the team adds a span attribute cart.items so the next search can filter on cart size, and links metrics to traces with exemplars so the jump from a dashboard spike to a trace is one click.

Failure modes

  • Broken traces: a proxy or queue that drops traceparent splits one request into several root spans. Check that every hop propagates context and that all services use the same propagator.
  • Clock skew: spans from hosts with drifting clocks appear to start before their parents. The query extension has a max_clock_skew_adjust setting, but fix NTP rather than relying on display adjustment.
  • Silent drops at the collector: when storage slows, exporter queues fill and spans are refused. Alert on refused and dropped span metrics, not only on collector liveness.
  • Sampler stuck on its initial rate: SDKs that cannot reach port 5778 never receive strategies. Test the endpoint from inside the service network.
  • Index explosion: high-cardinality span attributes, such as user ids as tag keys, bloat Elasticsearch mappings. Put variable data in values, never in keys.
  • Retention never runs: daily indices or tables accumulate until disks fill. Configure rollover or a cleanup job on day one.
  • Adaptive sampling on memory storage with several collectors: each instance keeps its own state, so calculations diverge. Use a shared sampling store.

Trade-offs and sizing

Jaeger's strengths are a focused trace UI, central sampling control and mature storage integrations. Its limits are equally clear: it is a trace store, not a full observability platform, so logs and metrics live elsewhere and correlation depends on shared trace ids. If you want derived request rate, error and duration metrics, generate them from spans with the span metrics connector in a Collector and point the Monitor tab at that metrics backend, as described in span metrics.

Size from the span rate. Estimate spans per request, times requests per second, times the sampling rate, times average encoded span size, and add replication. For example, 2,000 requests per second with 25 spans each at a 10% sample rate and about 1 KB per span is roughly 5 MB per second, about 430 GB per day before replication. That arithmetic usually decides two things at once: the sampling rate you can afford and whether you need Kafka in front of storage.

What to do next

  1. Run the all-in-one container locally, point one service at port 4317 and confirm its traces appear in the UI on port 16686.
  2. Instrument with OpenTelemetry SDKs, set service.name explicitly and use the W3C trace context propagator on every hop.
  3. Pick a storage backend from the table, set retention and index rollover before production traffic arrives, and configure an archive store for traces you want to keep.
  4. Write a sampling strategy file with full sampling for critical routes and zero for health checks, serve it through remote_sampling, and alert when SDKs fail to poll.
  5. Scrape Jaeger's own metrics and alert on refused or dropped spans and exporter queue size.
  6. Add Kafka between collectors and storage only when bursts or storage maintenance are actually costing you spans.
Key takeaway: Jaeger v2 is one OpenTelemetry Collector-based binary that you configure into collector, query, ingester or all-in-one roles. Its own extensions add named trace storage, the query UI and remote sampling. Choose storage by query needs and operations skill, control cost with retention and centrally managed sampling, buffer through Kafka only when bursts demand it, and monitor the pipeline itself, because the most common tracing failure is spans dropped silently before anyone looks.