OpenTelemetry is the vendor-neutral standard for traces, metrics and logs, and Java has its most complete implementation: a Java agent that instruments hundreds of libraries without code changes, a stable API and SDK, and a Spring Boot starter. The pieces are easy to start and easy to misconfigure: spans that never arrive, traces that break at every thread pool, and metric bills that explode because someone added a user id as an attribute.
This article explains how the parts fit, then builds up a service's telemetry from the agent, to manual spans and metrics, to a hand-configured SDK, with the defaults that matter checked against the OpenTelemetry documentation. Library versions move monthly, so pin them with the BOMs rather than copying numbers from any article, including this one.
The moving parts
The API (opentelemetry-api) is what code calls: Tracer, Span, Meter and Context. On its own it does nothing; every call is a cheap no-op, which is why libraries can depend on it safely. The SDK implements the API: it samples, batches and exports, and is configured once per application. Instrumentation creates spans for libraries such as servlet containers, JDBC, HTTP clients and Kafka, and the Java agent packages the SDK plus that instrumentation, attached with -javaagent. The Collector is a separate process that receives OTLP, processes it and forwards it to backends.
Three ways in
| Approach | Code changes | Coverage | Choose it when |
|---|---|---|---|
| Java agent | None; one JVM flag | Hundreds of libraries automatically | Most services; fastest path to useful traces |
| Spring Boot starter | A dependency and properties | Spring and common libraries, fewer than the agent | You cannot add JVM flags, or native images |
| Manual SDK plus library instrumentation | Explicit setup code | Only what you wire | Tight startup or size budgets, or full control |
All three accept the same OTEL_* environment variables for the common settings, and your own manual spans work identically on top of any of them, because they only touch the API. Start with the agent unless a constraint rules it out.
Configuring the agent, and the protocol trap
The agent reads its configuration from environment variables or the equivalent system properties. A production baseline:
export OTEL_SERVICE_NAME=orders-api
export OTEL_RESOURCE_ATTRIBUTES=deployment.environment.name=prod,service.version=1.42.0
export OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4318 # agent 2.x default protocol: http/protobuf
export OTEL_TRACES_SAMPLER=parentbased_traceidratio
export OTEL_TRACES_SAMPLER_ARG=0.10
export OTEL_INSTRUMENTATION_JDBC_ENABLED=true # per-library switches follow this pattern
java -javaagent:/opt/otel/opentelemetry-javaagent.jar -jar orders-api.jarOne default catches people constantly. The SDK's autoconfigure module defaults the OTLP protocol to grpc on port 4317, but Java agent 2.x and the Spring Boot starter default to http/protobuf on port 4318. A team that points an agent at a Collector's gRPC port, or a hand-configured SDK at the HTTP port, gets no data and often no obvious error. Set OTEL_EXPORTER_OTLP_PROTOCOL explicitly and match the port. Traces, metrics and logs all export over OTLP by default; metrics are exported every 60 seconds unless OTEL_METRIC_EXPORT_INTERVAL says otherwise.
Always set OTEL_SERVICE_NAME; without it, every service shows up under an unknown name. Set OTEL_JAVAAGENT_DEBUG=true in a test environment to log spans as they are created, which settles most did-anything-happen questions in a minute.
Manual spans for your own logic
The agent sees HTTP requests and database calls, but not that a request spent 300 ms applying pricing rules. Add a span for any unit of business work you will want to find in a trace.
import io.opentelemetry.api.GlobalOpenTelemetry;
import io.opentelemetry.api.trace.Span;
import io.opentelemetry.api.trace.StatusCode;
import io.opentelemetry.api.trace.Tracer;
import io.opentelemetry.context.Scope;
public final class PricingService {
private static final Tracer TRACER = GlobalOpenTelemetry.getTracer("com.example.pricing");
public Quote quote(String sku, int quantity) {
Span span = TRACER.spanBuilder("pricing.quote")
.setAttribute("app.sku", sku)
.setAttribute("app.quantity", quantity)
.startSpan();
try (Scope ignored = span.makeCurrent()) { // children and logs see this span
Quote q = rules.apply(catalog.load(sku), quantity);
span.setAttribute("app.discount_applied", q.discounted());
return q;
} catch (RuntimeException e) {
span.recordException(e);
span.setStatus(StatusCode.ERROR, e.getClass().getSimpleName());
throw e;
} finally {
span.end(); // always end, or the span is never exported
}
}
}Three rules are visible here. Make the span current with makeCurrent() inside try-with-resources, so child spans and log lines attach to it and the previous context is restored. Record exceptions and set an error status explicitly; a thrown exception alone does not mark a manual span as failed. End the span in finally, because a span that never ends is never exported. Name spans after the operation, not the input: pricing.quote, never pricing.quote.SKU-123. Attribute names should follow the semantic conventions where one exists, with your own namespace prefix otherwise.
Context across threads
Context lives in a thread-local. When work hops to another thread, through an executor, a CompletableFuture or a reactive pipeline, the context must be carried, or the new work starts a fresh, disconnected trace. The agent instruments the common executors and async libraries and does this for you. Without the agent, wrap executors yourself:
import io.opentelemetry.context.Context;
import io.opentelemetry.instrumentation.annotations.SpanAttribute;
import io.opentelemetry.instrumentation.annotations.WithSpan;
// Without the agent, wrap executors so tasks inherit the submitting thread's context.
ExecutorService pool = Context.taskWrapping(Executors.newFixedThreadPool(8));
@WithSpan("inventory.reserve") // needs the agent or Spring starter
public Reservation reserve(@SpanAttribute("app.sku") String sku, int qty) { ... }@WithSpan and @SpanAttribute come from the opentelemetry-instrumentation-annotations artifact and only produce spans when the agent or Spring starter is present. Virtual threads behave like any other thread here: the context is carried at submission, so the same wrapping applies; see the virtual threads article for their scheduling model.
Metrics without a cardinality explosion
Metrics are aggregated in the SDK, so every unique combination of attribute values becomes its own time series held in memory and stored in the backend.
Meter meter = GlobalOpenTelemetry.getMeter("com.example.orders");
LongCounter placed = meter.counterBuilder("app.orders.placed")
.setDescription("Orders accepted").setUnit("{order}").build();
DoubleHistogram quoteTime = meter.histogramBuilder("app.pricing.quote.duration")
.setUnit("s").build();
AttributeKey<String> METHOD = AttributeKey.stringKey("app.payment_method");
placed.add(1, Attributes.of(METHOD, "card")); // bounded values only, never user ids
quoteTime.record(elapsedSeconds);Use counters for things that only go up, histograms for durations and sizes, and gauges or up-down counters for current levels. Keep attribute values to small known sets: payment method, region, outcome. A user id, an order id or a raw URL path as an attribute turns one series into millions. Durations are in seconds, following current semantic conventions, which matters when you compare your histograms with the agent's HTTP server metrics.
Correlating logs with traces
With the agent, the active span's trace_id, span_id and trace_flags are injected into the logging MDC by default, so adding %X{trace_id} to a Logback pattern links each log line to its trace. The agent also bridges log records into OTLP, so logs can travel through the Collector with trace ids attached. Decide on one path to avoid storing every line twice: either ship logs over OTLP, or keep your existing log shipper and rely on the MDC fields.
Sampling: head in the SDK, tail in the Collector
The default sampler is parentbased_always_on: record everything, and honour the caller's decision. That is right for development and expensive at scale. parentbased_traceidratio with an argument of 0.10 keeps about one root trace in ten, and child services follow the root's decision, so traces stay complete. Head sampling decides before anything interesting has happened, so it discards most errors along with the noise. To keep every error and every slow request, export everything to a Collector and use tail sampling there, which needs all spans of a trace to reach the same Collector instance.
Configuring the SDK by hand
Without the agent, you build the SDK yourself once at startup and register it globally. The autoconfigure module, AutoConfiguredOpenTelemetrySdk, can build the same thing from the OTEL_* variables; this is the explicit form.
Resource resource = Resource.getDefault().merge(Resource.create(
Attributes.of(AttributeKey.stringKey("service.name"), "orders-api")));
SdkTracerProvider tracerProvider = SdkTracerProvider.builder()
.setResource(resource)
.setSampler(Sampler.parentBased(Sampler.traceIdRatioBased(0.10)))
.addSpanProcessor(BatchSpanProcessor.builder(
OtlpHttpSpanExporter.builder()
.setEndpoint("http://otel-collector:4318/v1/traces").build()).build())
.build();
OpenTelemetrySdk sdk = OpenTelemetrySdk.builder()
.setTracerProvider(tracerProvider)
.setPropagators(ContextPropagators.create(W3CTraceContextPropagator.getInstance()))
.buildAndRegisterGlobal();
Runtime.getRuntime().addShutdownHook(new Thread(sdk::close)); // flush queued spans on exitThe batch span processor queues up to 2,048 spans and exports batches of up to 512 every 5 seconds by default. If spans arrive faster than export drains them, the queue fills and new spans are dropped, which shows up as traces with holes. Raise the queue or reduce volume, and close the SDK on shutdown so the final batch is flushed.
Running the agent in production
Measure the agent's cost rather than guessing it. Run the same load test with and without the agent and compare latency percentiles, CPU, heap and startup time; the difference depends on how many libraries are instrumented and how many spans each request creates. If startup or overhead matters, set OTEL_INSTRUMENTATION_COMMON_DEFAULT_ENABLED=false and enable only the modules you use, one OTEL_INSTRUMENTATION_<NAME>_ENABLED=true at a time.
Upgrades need the same care as any dependency. The agent releases frequently, and changes to semantic conventions can rename attributes: agent 2.0 switched HTTP spans and metrics to the stable conventions, so http.method became http.request.method and dashboards built on the old names went blank. Read the release notes for convention changes, upgrade staging first, and compare dashboards before and after. Add Kubernetes pod, namespace and node attributes in the Collector rather than in every JVM, so the service configuration stays small and identical across environments. For CPU and allocation problems that spans cannot explain, pair traces with Java Flight Recorder recordings.
Worked example: one slow checkout
A user reports a slow checkout. The trace for that request shows the agent's server span for POST /checkout at 2.4 s. Under it, the agent's JDBC spans total 80 ms, and the manual pricing.quote span takes 2.2 s. Inside that, an HTTP client span to the promotions service takes 2.1 s, and because the agent injected traceparent into the outgoing request, the trace continues into the promotions service, where a server span shows the same 2.1 s and a single SELECT inside it with no index. Search logs for the trace id and you find the promotion-lookup warning that came with it. Without the manual span, the trace would show 2.2 s of unexplained gap; without propagation, it would stop at the client call.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| No telemetry at all | Protocol and port mismatch, or wrong endpoint | Set protocol explicitly; enable agent debug; check Collector receiver logs |
| Traces break into fragments | Context lost across a custom thread pool or queue | Context.taskWrapping, or propagate context in message headers |
| Duplicate spans for one call | Agent plus manually added library instrumentation | Use one; disable the agent's module for that library |
| Traces with missing children | Batch queue full, spans dropped | Larger queue, lower sampling ratio, Collector close by |
| Metrics backend bill spikes | High-cardinality attribute values | Bounded attribute values; views to drop attributes |
| Startup slower by seconds | Agent instrumenting at class load | Measure it; disable unused instrumentation modules |
What to do next
- Attach the agent to one service in staging with
OTEL_SERVICE_NAME, an explicit protocol and endpoint, and debug logging, and confirm spans arrive. - Run a Collector between your services and backends, and point every service at it rather than at a vendor.
- Add manual spans around the two or three business operations you most often need to explain, with error status recording.
- Pin the agent, API and instrumentation versions with the BOMs and upgrade them on a schedule.
- Add trace ids to your log pattern and pick one log shipping path.
- Choose a sampling ratio from your traffic and budget, and add tail sampling for errors and slow requests. Read the Jaeger tracing article to explore traces once they land.