OpenTelemetry is the vendor-neutral standard for traces, metrics and logs, and Java has its most complete implementation: a Java agent that instruments hundreds of libraries without code changes, a stable API and SDK, and a Spring Boot starter. The pieces are easy to start and easy to misconfigure: spans that never arrive, traces that break at every thread pool, and metric bills that explode because someone added a user id as an attribute.

This article explains how the parts fit, then builds up a service's telemetry from the agent, to manual spans and metrics, to a hand-configured SDK, with the defaults that matter checked against the OpenTelemetry documentation. Library versions move monthly, so pin them with the BOMs rather than copying numbers from any article, including this one.

Advertisement

The moving parts

The API (opentelemetry-api) is what code calls: Tracer, Span, Meter and Context. On its own it does nothing; every call is a cheap no-op, which is why libraries can depend on it safely. The SDK implements the API: it samples, batches and exports, and is configured once per application. Instrumentation creates spans for libraries such as servlet containers, JDBC, HTTP clients and Kafka, and the Java agent packages the SDK plus that instrumentation, attached with -javaagent. The Collector is a separate process that receives OTLP, processes it and forwards it to backends.

OpenTelemetry in a Java service: API, SDK, agent and CollectorJVM: orders-apiYour codeTracer, Meter, @WithSpanJava agentauto-instruments librariesOpenTelemetry APIContext, Span, Meter: no-op until an SDK is setSDKResource, Sampler, BatchSpanProcessor, MetricReaderOTLPCollectorbatch, tail-sample, redactTrace storeJaeger, TempoMetricsPrometheusLogsLoki, Elastictraceparent header to downstream services00-trace id-parent span id-flagsLibraries depend only on the API; the application chooses one SDK configuration, ideally exporting to a Collector rather than to vendors directly.
Code and the agent both talk to the API; one SDK configuration exports over OTLP to a Collector, which routes to trace, metric and log backends. Context crosses process boundaries in the W3C traceparent header.

Three ways in

ApproachCode changesCoverageChoose it when
Java agentNone; one JVM flagHundreds of libraries automaticallyMost services; fastest path to useful traces
Spring Boot starterA dependency and propertiesSpring and common libraries, fewer than the agentYou cannot add JVM flags, or native images
Manual SDK plus library instrumentationExplicit setup codeOnly what you wireTight startup or size budgets, or full control

All three accept the same OTEL_* environment variables for the common settings, and your own manual spans work identically on top of any of them, because they only touch the API. Start with the agent unless a constraint rules it out.

Advertisement

Configuring the agent, and the protocol trap

The agent reads its configuration from environment variables or the equivalent system properties. A production baseline:

export OTEL_SERVICE_NAME=orders-api
export OTEL_RESOURCE_ATTRIBUTES=deployment.environment.name=prod,service.version=1.42.0
export OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4318   # agent 2.x default protocol: http/protobuf
export OTEL_TRACES_SAMPLER=parentbased_traceidratio
export OTEL_TRACES_SAMPLER_ARG=0.10
export OTEL_INSTRUMENTATION_JDBC_ENABLED=true                   # per-library switches follow this pattern
java -javaagent:/opt/otel/opentelemetry-javaagent.jar -jar orders-api.jar

One default catches people constantly. The SDK's autoconfigure module defaults the OTLP protocol to grpc on port 4317, but Java agent 2.x and the Spring Boot starter default to http/protobuf on port 4318. A team that points an agent at a Collector's gRPC port, or a hand-configured SDK at the HTTP port, gets no data and often no obvious error. Set OTEL_EXPORTER_OTLP_PROTOCOL explicitly and match the port. Traces, metrics and logs all export over OTLP by default; metrics are exported every 60 seconds unless OTEL_METRIC_EXPORT_INTERVAL says otherwise.

Always set OTEL_SERVICE_NAME; without it, every service shows up under an unknown name. Set OTEL_JAVAAGENT_DEBUG=true in a test environment to log spans as they are created, which settles most did-anything-happen questions in a minute.

Manual spans for your own logic

The agent sees HTTP requests and database calls, but not that a request spent 300 ms applying pricing rules. Add a span for any unit of business work you will want to find in a trace.

import io.opentelemetry.api.GlobalOpenTelemetry;
import io.opentelemetry.api.trace.Span;
import io.opentelemetry.api.trace.StatusCode;
import io.opentelemetry.api.trace.Tracer;
import io.opentelemetry.context.Scope;

public final class PricingService {
    private static final Tracer TRACER = GlobalOpenTelemetry.getTracer("com.example.pricing");

    public Quote quote(String sku, int quantity) {
        Span span = TRACER.spanBuilder("pricing.quote")
                .setAttribute("app.sku", sku)
                .setAttribute("app.quantity", quantity)
                .startSpan();
        try (Scope ignored = span.makeCurrent()) {      // children and logs see this span
            Quote q = rules.apply(catalog.load(sku), quantity);
            span.setAttribute("app.discount_applied", q.discounted());
            return q;
        } catch (RuntimeException e) {
            span.recordException(e);
            span.setStatus(StatusCode.ERROR, e.getClass().getSimpleName());
            throw e;
        } finally {
            span.end();                                 // always end, or the span is never exported
        }
    }
}

Three rules are visible here. Make the span current with makeCurrent() inside try-with-resources, so child spans and log lines attach to it and the previous context is restored. Record exceptions and set an error status explicitly; a thrown exception alone does not mark a manual span as failed. End the span in finally, because a span that never ends is never exported. Name spans after the operation, not the input: pricing.quote, never pricing.quote.SKU-123. Attribute names should follow the semantic conventions where one exists, with your own namespace prefix otherwise.

Context across threads

Context lives in a thread-local. When work hops to another thread, through an executor, a CompletableFuture or a reactive pipeline, the context must be carried, or the new work starts a fresh, disconnected trace. The agent instruments the common executors and async libraries and does this for you. Without the agent, wrap executors yourself:

import io.opentelemetry.context.Context;
import io.opentelemetry.instrumentation.annotations.SpanAttribute;
import io.opentelemetry.instrumentation.annotations.WithSpan;

// Without the agent, wrap executors so tasks inherit the submitting thread's context.
ExecutorService pool = Context.taskWrapping(Executors.newFixedThreadPool(8));

@WithSpan("inventory.reserve")                          // needs the agent or Spring starter
public Reservation reserve(@SpanAttribute("app.sku") String sku, int qty) { ... }

@WithSpan and @SpanAttribute come from the opentelemetry-instrumentation-annotations artifact and only produce spans when the agent or Spring starter is present. Virtual threads behave like any other thread here: the context is carried at submission, so the same wrapping applies; see the virtual threads article for their scheduling model.

Metrics without a cardinality explosion

Metrics are aggregated in the SDK, so every unique combination of attribute values becomes its own time series held in memory and stored in the backend.

Meter meter = GlobalOpenTelemetry.getMeter("com.example.orders");

LongCounter placed = meter.counterBuilder("app.orders.placed")
        .setDescription("Orders accepted").setUnit("{order}").build();
DoubleHistogram quoteTime = meter.histogramBuilder("app.pricing.quote.duration")
        .setUnit("s").build();

AttributeKey<String> METHOD = AttributeKey.stringKey("app.payment_method");
placed.add(1, Attributes.of(METHOD, "card"));           // bounded values only, never user ids
quoteTime.record(elapsedSeconds);

Use counters for things that only go up, histograms for durations and sizes, and gauges or up-down counters for current levels. Keep attribute values to small known sets: payment method, region, outcome. A user id, an order id or a raw URL path as an attribute turns one series into millions. Durations are in seconds, following current semantic conventions, which matters when you compare your histograms with the agent's HTTP server metrics.

Correlating logs with traces

With the agent, the active span's trace_id, span_id and trace_flags are injected into the logging MDC by default, so adding %X{trace_id} to a Logback pattern links each log line to its trace. The agent also bridges log records into OTLP, so logs can travel through the Collector with trace ids attached. Decide on one path to avoid storing every line twice: either ship logs over OTLP, or keep your existing log shipper and rely on the MDC fields.

Sampling: head in the SDK, tail in the Collector

The default sampler is parentbased_always_on: record everything, and honour the caller's decision. That is right for development and expensive at scale. parentbased_traceidratio with an argument of 0.10 keeps about one root trace in ten, and child services follow the root's decision, so traces stay complete. Head sampling decides before anything interesting has happened, so it discards most errors along with the noise. To keep every error and every slow request, export everything to a Collector and use tail sampling there, which needs all spans of a trace to reach the same Collector instance.

Configuring the SDK by hand

Without the agent, you build the SDK yourself once at startup and register it globally. The autoconfigure module, AutoConfiguredOpenTelemetrySdk, can build the same thing from the OTEL_* variables; this is the explicit form.

Resource resource = Resource.getDefault().merge(Resource.create(
        Attributes.of(AttributeKey.stringKey("service.name"), "orders-api")));

SdkTracerProvider tracerProvider = SdkTracerProvider.builder()
        .setResource(resource)
        .setSampler(Sampler.parentBased(Sampler.traceIdRatioBased(0.10)))
        .addSpanProcessor(BatchSpanProcessor.builder(
                OtlpHttpSpanExporter.builder()
                        .setEndpoint("http://otel-collector:4318/v1/traces").build()).build())
        .build();

OpenTelemetrySdk sdk = OpenTelemetrySdk.builder()
        .setTracerProvider(tracerProvider)
        .setPropagators(ContextPropagators.create(W3CTraceContextPropagator.getInstance()))
        .buildAndRegisterGlobal();

Runtime.getRuntime().addShutdownHook(new Thread(sdk::close));   // flush queued spans on exit

The batch span processor queues up to 2,048 spans and exports batches of up to 512 every 5 seconds by default. If spans arrive faster than export drains them, the queue fills and new spans are dropped, which shows up as traces with holes. Raise the queue or reduce volume, and close the SDK on shutdown so the final batch is flushed.

Running the agent in production

Measure the agent's cost rather than guessing it. Run the same load test with and without the agent and compare latency percentiles, CPU, heap and startup time; the difference depends on how many libraries are instrumented and how many spans each request creates. If startup or overhead matters, set OTEL_INSTRUMENTATION_COMMON_DEFAULT_ENABLED=false and enable only the modules you use, one OTEL_INSTRUMENTATION_<NAME>_ENABLED=true at a time.

Upgrades need the same care as any dependency. The agent releases frequently, and changes to semantic conventions can rename attributes: agent 2.0 switched HTTP spans and metrics to the stable conventions, so http.method became http.request.method and dashboards built on the old names went blank. Read the release notes for convention changes, upgrade staging first, and compare dashboards before and after. Add Kubernetes pod, namespace and node attributes in the Collector rather than in every JVM, so the service configuration stays small and identical across environments. For CPU and allocation problems that spans cannot explain, pair traces with Java Flight Recorder recordings.

Worked example: one slow checkout

A user reports a slow checkout. The trace for that request shows the agent's server span for POST /checkout at 2.4 s. Under it, the agent's JDBC spans total 80 ms, and the manual pricing.quote span takes 2.2 s. Inside that, an HTTP client span to the promotions service takes 2.1 s, and because the agent injected traceparent into the outgoing request, the trace continues into the promotions service, where a server span shows the same 2.1 s and a single SELECT inside it with no index. Search logs for the trace id and you find the promotion-lookup warning that came with it. Without the manual span, the trace would show 2.2 s of unexplained gap; without propagation, it would stop at the client call.

Failure modes

SymptomCauseFix
No telemetry at allProtocol and port mismatch, or wrong endpointSet protocol explicitly; enable agent debug; check Collector receiver logs
Traces break into fragmentsContext lost across a custom thread pool or queueContext.taskWrapping, or propagate context in message headers
Duplicate spans for one callAgent plus manually added library instrumentationUse one; disable the agent's module for that library
Traces with missing childrenBatch queue full, spans droppedLarger queue, lower sampling ratio, Collector close by
Metrics backend bill spikesHigh-cardinality attribute valuesBounded attribute values; views to drop attributes
Startup slower by secondsAgent instrumenting at class loadMeasure it; disable unused instrumentation modules

What to do next

  1. Attach the agent to one service in staging with OTEL_SERVICE_NAME, an explicit protocol and endpoint, and debug logging, and confirm spans arrive.
  2. Run a Collector between your services and backends, and point every service at it rather than at a vendor.
  3. Add manual spans around the two or three business operations you most often need to explain, with error status recording.
  4. Pin the agent, API and instrumentation versions with the BOMs and upgrade them on a schedule.
  5. Add trace ids to your log pattern and pick one log shipping path.
  6. Choose a sampling ratio from your traffic and budget, and add tail sampling for errors and slow requests. Read the Jaeger tracing article to explore traces once they land.
Key takeaway: OpenTelemetry Java splits into an API your code and libraries call, an SDK that samples and exports, and a Java agent that instruments libraries automatically. Start with the agent, set the service name, protocol and endpoint explicitly, and add manual spans for business work. Carry context across threads, keep metric attributes bounded, correlate logs through the MDC, and sample deliberately, keeping errors through tail sampling in a Collector.