ADK Java already speaks OpenTelemetry. It creates spans for every invocation, agent run, model call and tool call, and records latency and size histograms, all through the OpenTelemetry API. What it does not do is decide where that telemetry goes, how it joins the rest of your system, or which of it you keep. Those are integration decisions, and they are where most teams lose time: traces that start at the runner instead of at the HTTP request, tool calls whose downstream HTTP spans land in a separate trace, a Collector that samples away the one slow conversation you needed, and tests that never notice any of it.
This article is about those seams. It walks through the span tree ADK builds and exactly how it picks a parent, how context crosses RxJava and your own thread pools, the three ways to bootstrap the SDK, a Collector pipeline with redaction and tail sampling, and a unit test that pins the trace shape. The attribute catalogue and the content-capture switch are covered in the observability architecture article and tracing prompts and completions; this page assumes them and does not repeat them. Everything named here was read from the adk-java source on main in October 2026; check your own jar if you run an older release.
The span tree ADK builds
ADK obtains a tracer and a meter named gcp.vertex.agent from GlobalOpenTelemetry. The spans it creates form a fixed tree:
| Span name | Created by | Parent |
|---|---|---|
invocation | the Runner, once per runAsync or live run | whatever context is current when the stream is subscribed |
invoke_agent <name> | each agent run, including sub-agents | the invocation, or the calling agent's span |
call_llm | the LLM flow, once per model request | the agent span |
execute_tool <name> | each tool call, wrapping before/after callbacks and the body | the agent span |
execute_tool (merged) | when one model response asks for several tools | the agent span |
send_data | data sent on a live (bidirectional) connection | the agent span |
On the metrics side the same instrumentation name carries histograms such as gen_ai.agent.invocation.duration, gen_ai.tool.execution.duration and gen_ai.agent.workflow.steps. Two details matter when you integrate with an existing estate. First, call_llm spans set gen_ai.system to gcp.vertex.agent rather than to the model provider, so a dashboard that groups GenAI spans by provider will lump every ADK model call together whether it went to Gemini or to another model behind a custom BaseLlm. Group by gen_ai.request.model instead. Second, the GenAI semantic conventions are still marked as in development and newer versions rename keys (gen_ai.system has been superseded by gen_ai.provider.name). Build queries on the attribute names your exporter actually receives, and re-check them after every ADK upgrade.
Architecture: one trace per request
The goal is one trace per user request, rooted at the inbound HTTP or gRPC span, with ADK's spans in the middle and the tool's outbound calls (to your order service, a database, a search API) at the leaves. When that holds, a slow turn reads top to bottom: the request took 9 seconds, 7 of them in two call_llm spans, 1.5 in one execute_tool get_order whose child HTTP span shows the order service took 1.4 seconds. When it does not hold you get three disconnected traces and a guessing game.
How ADK chooses a parent
ADK's span helper, Tracing.TracerProvider, is an RxJava transformer. It wraps the stream in defer, and when the stream is subscribed it reads Context.current() (unless an explicit parent was set), starts the span with that parent, and then makes the new span's context current while it subscribes upstream and while it delivers each signal downstream. The consequence is simple and easy to miss: the parent of the invocation span is whatever context is current on the thread that subscribes, not the thread that built the Flowable.
In a blocking servlet controller those are the same thread and it just works: the server span is current, you call runAsync(...).blockingForEach(...), and the invocation becomes its child. In a reactive stack the framework often subscribes later, on another thread, where the server span is no longer current. ADK exposes the fix itself: Tracing.withContext(Context) is a public transformer that makes a captured context current during subscription. Capture at assembly time, apply it to the stream:
import com.google.adk.telemetry.Tracing;
import io.opentelemetry.context.Context;
@PostMapping(value = "/chat", produces = MediaType.TEXT_EVENT_STREAM_VALUE)
public Flux<String> chat(@RequestBody ChatRequest req) {
Context requestContext = Context.current(); // the server span, captured now
Flowable<Event> events = runner
.runAsync(req.userId(), req.sessionId(), Content.fromParts(Part.fromText(req.text())))
.compose(Tracing.withContext(requestContext)); // current again at subscribe time
return Flux.from(events)
.filter(e -> e.content().isPresent())
.map(e -> e.stringifyContent());
}If you run a message consumer instead of an HTTP server, the same rule applies with a different source: extract the W3C traceparent from the message headers with the configured propagator, build a consumer span with that parent, and subscribe the runner while it is current.
Context inside tools and thread pools
Inside a tool, ADK has made the execute_tool span current, so a synchronous tool body sees it as Span.current() and any instrumented HTTP or JDBC client it calls nests correctly. Problems start when the tool hands work to its own executor. CompletableFuture.supplyAsync on the common pool, a Guava ListeningExecutorService or a hand-rolled thread pool all run the task on a thread that has no OpenTelemetry context, and the downstream span becomes the root of a new trace. Wrap the executor once with Context.taskWrapping, which captures the caller's context for every submitted task:
public final class InventoryTool {
// Every task submitted here runs with the submitter's context current.
private static final ExecutorService POOL =
Context.taskWrapping(Executors.newFixedThreadPool(8));
@Schema(description = "Check stock for a SKU across warehouses")
public static Single<Map<String, Object>> checkStock(@Schema(name = "sku") String sku) {
return Single.fromCompletionStage(CompletableFuture.supplyAsync(
() -> Map.<String, Object>of("sku", sku, "levels", warehouses.query(sku)), POOL));
}
}The other direction matters too: outbound propagation. An instrumented HTTP client injects traceparent into the request, so the service you call joins the same trace. If a tool builds requests by hand with a client that is not instrumented, inject the header yourself with GlobalOpenTelemetry.getPropagators().getTextMapPropagator().inject(...). A remote agent called over HTTP then joins the same trace.
Three ways to bootstrap the SDK
There are three ways to get a real SDK behind GlobalOpenTelemetry. They differ in how much instrumentation you get for free and how much control you keep.
| Option | What you add | You get | Watch out for |
|---|---|---|---|
| Programmatic SDK | SDK + OTLP exporter dependencies, a builder in main | Full control, no magic | Must register before any ADK telemetry class loads; no HTTP or JDBC spans unless you add library instrumentation |
| SDK autoconfigure | opentelemetry-sdk-extension-autoconfigure, configured by OTEL_* variables | Standard env-var configuration, same as other services | Still needs instrumentation for server and client spans |
| Java agent | -javaagent:opentelemetry-javaagent.jar | Global SDK registered before your code runs, plus server, HTTP client, JDBC and many other spans | Startup cost, a large dependency you upgrade separately, instrumentation you did not choose |
For most services the Java agent is the pragmatic choice: it removes the registration-ordering trap (explained in the tool metrics article), it produces the server and client spans that give ADK's spans a parent and children, and it is configured with the same variables as everything else you run:
OTEL_SERVICE_NAME=support-agent
OTEL_RESOURCE_ATTRIBUTES=service.version=2026.10.08-1,deployment.environment=prod
OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317
OTEL_EXPORTER_OTLP_PROTOCOL=grpc
OTEL_TRACES_SAMPLER=parentbased_always_on # keep everything here; sample in the Collector
OTEL_METRIC_EXPORT_INTERVAL=30000
ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS=false
java -javaagent:/otel/opentelemetry-javaagent.jar -jar support-agent.jarPut service.version on the resource, not on spans. It is the single most useful dimension during a rollout, because every ADK histogram then splits by release without any code.
A Collector pipeline for agent traffic
Export to a Collector next to the service (a sidecar or node agent) rather than to a vendor. The Collector is where telemetry policy lives, so changing it does not need an agent redeploy. A minimal contrib-distribution pipeline for an agent service:
receivers:
otlp:
protocols: { grpc: { endpoint: 0.0.0.0:4317 } }
processors:
memory_limiter: { check_interval: 1s, limit_percentage: 80, spike_limit_percentage: 20 }
attributes/strip-content: # defence in depth if someone flips the ADK switch back on
actions:
- { key: gcp.vertex.agent.llm_request, action: delete }
- { key: gcp.vertex.agent.llm_response, action: delete }
- { key: gcp.vertex.agent.tool_call_args, action: delete }
- { key: gcp.vertex.agent.tool_response, action: delete }
tail_sampling:
decision_wait: 30s # agent turns are slow; wait for the whole trace
policies:
- { name: errors, type: status_code, status_code: { status_codes: [ERROR] } }
- { name: slow, type: latency, latency: { threshold_ms: 15000 } }
- { name: rest, type: probabilistic, probabilistic: { sampling_percentage: 10 } }
batch: {}
exporters:
otlp/traces: { endpoint: traces-backend:4317 }
service:
pipelines:
traces: { receivers: [otlp], processors: [memory_limiter, attributes/strip-content, tail_sampling, batch], exporters: [otlp/traces] }Two agent-specific settings deserve attention. decision_wait must exceed your longest normal turn: the tail sampler decides when the wait expires, and spans that arrive later are judged without the rest of their trace. And tail sampling only works when every span of a trace reaches the same Collector instance. With several replicas, put a first tier in front that uses the load-balancing exporter keyed on trace id, and run the sampler in the second tier. Metrics need no sampling; send them through a separate pipeline untouched, so error rates stay exact even when only 10 percent of healthy traces are kept.
Testing the trace shape
Trace shape is a contract, and contracts deserve tests. ADK ships hooks for exactly this: Tracing.setTracerForTesting(Tracer) and Metrics.setMeterForTesting(Meter). Combine them with the SDK's in-memory exporter and assert parentage, not just existence:
class TraceShapeTest {
private final InMemorySpanExporter spans = InMemorySpanExporter.create();
private final SdkTracerProvider provider = SdkTracerProvider.builder()
.addSpanProcessor(SimpleSpanProcessor.create(spans)).build();
@BeforeEach void wire() { Tracing.setTracerForTesting(provider.get("test")); }
@Test void toolSpanNestsUnderRequest() {
Span request = provider.get("test").spanBuilder("POST /chat").startSpan();
try (Scope s = request.makeCurrent()) {
runner.runAsync("u1", sessionId, Content.fromParts(Part.fromText("where is order 42?")))
.blockingSubscribe();
} finally { request.end(); }
Map<String, SpanData> byName = spans.getFinishedSpanItems().stream()
.collect(Collectors.toMap(SpanData::getName, d -> d, (a, b) -> a));
SpanData invocation = byName.get("invocation");
assertEquals(request.getSpanContext().getSpanId(), invocation.getParentSpanId());
SpanData tool = byName.get("execute_tool get_order");
assertEquals(invocation.getTraceId(), tool.getTraceId());
}
}Drive the agent with a scripted fake model so the test is deterministic and free; testing with custom LLMs shows how. Add one more test for any tool that uses its own executor, asserting that the downstream client span shares the tool span's trace id. That is the test that catches a missing taskWrapping.
Worked example: two traces per request after a WebFlux migration
A team moves its support agent from a servlet controller to Spring WebFlux. Dashboards look fine, but the on-call engineer notices that every request now produces two traces: one with only the HTTP server span (about 40 ms), and one rooted at invocation (several seconds). Latency SLOs computed from server spans now say the service answers in 40 ms, which is the time to open the event stream, not to answer.
Diagnosis takes one query: count traces whose root span is named invocation. Before the migration it was zero; after, it equals request volume. The cause is the subscription rule above: WebFlux subscribes to the returned Flux after the controller method returns, on a Netty event-loop thread where the server span is not current, so ADK's defer reads an empty context and starts a new trace. The fix is the two-line Tracing.withContext(requestContext) shown earlier. The team adds the parentage test and an alert on root spans named invocation, and moves the latency SLO to gen_ai.agent.invocation.duration, because the server span ends when the stream opens.
Failure modes
- Orphaned invocations.
invocationspans as trace roots. The stream was subscribed without the request context current; capture it and applyTracing.withContext. - Tool calls that leave the trace. Downstream spans in separate traces. The tool uses an unwrapped executor or an uninstrumented client.
- Spans silently no-op. No SDK registered globally before ADK's static tracer was read. Use the Java agent, or register first thing in
main. - Truncated traces after tail sampling. Spans of long turns arrive after
decision_wait, or replicas see different halves of a trace. Lengthen the wait and route by trace id. - Lost spans on shutdown. The batch processor is not flushed when the pod stops. Call
shutdown()on the SDK in a shutdown hook (the Java agent does this for you).
Trade-offs
Java agent or manual SDK. The agent gives correct ordering and rich client spans for free at the price of startup time and instrumentation you did not pick; manual wiring is lean but you own every gap. Head or tail sampling. Head sampling in the SDK is cheap and consistent across services, but it decides before it knows a turn was slow or failed; tail sampling keeps the interesting traces but costs Collector memory proportional to traffic times decision_wait. Rich spans or safe spans. Content in spans is the fastest debugger you will ever have, and the widest data leak; keep it off in production and strip it again in the Collector.
What to do next
- Run one request and confirm the trace root is your server span, not
invocation. - In reactive or queue-driven entry points, capture
Context.current()at assembly and applyTracing.withContext. - Wrap every executor a tool uses with
Context.taskWrappingand check downstream spans share the trace id. - Choose a bootstrap (the Java agent unless you have a reason not to) and set
service.name,service.versionanddeployment.environmenton the resource. - Deploy a Collector with memory limiting, content stripping and tail sampling; set
decision_waitabove your p99 turn time and route by trace id if you run replicas. - Add a parentage unit test using
Tracing.setTracerForTestingand an in-memory exporter. - Alert on root spans named
invocationand on SDK dropped-span counts. - Group model spans by
gen_ai.request.model, notgen_ai.system, and re-check attribute names after each ADK upgrade.