A single user question to an orchestrator agent can turn into a model call, a delegation over A2A to a specialist agent in another service, three more model calls and a database tool on the other side. When the answer is slow or wrong, you want one trace that shows all of it. What you usually get is two unrelated traces: the orchestrator's, ending in an HTTP call that took nine seconds, and the specialist's, starting from nowhere.
ADK Java already emits invoke_agent, call_llm and execute_tool spans in each process, as the observability deep dive describes. This article is about the gap between processes: how W3C trace context crosses an A2A call, which component injects it, which one must extract it, where it gets lost on thread hops, what the spans mean for streaming calls, how sampling must be configured on both sides, and a test that proves the link instead of hoping for it. The protocol-level rules for the header itself, including validation and push notifications, are in A2A distributed tracing.
Why traces split at the agent boundary
Distributed tracing works by passing a small context with every request: a trace id shared by all spans of the request, the id of the calling span, and a sampled flag. Over HTTP that context travels in the W3C traceparent header, formatted as 00-<trace-id>-<parent-span-id>-<flags> with a 32-hex-digit trace id, a 16-hex-digit span id and flags 01 when sampled. The receiving service extracts it and starts its server span as a child of that parent.
Three things must all go right for an A2A call. The caller must have the right context current at the moment the client sends; ADK runs agents on RxJava, so the code that sends may run on a different thread from the code that started the agent span. The client must inject the header. The server runtime must extract it and keep it current through its own asynchronous hand-offs until ADK's runner starts the specialist's spans. A miss at any point yields a new root trace on the server, and nothing errors: you simply get two traces.
The path of the context
The diagram shows the path. Note which boxes belong to which library. ADK provides the agent spans on both sides. The a2a-java SDK's optional OpenTelemetry extras provide the client-side tracing and propagation wrappers and a server-side request handler decorator. The HTTP runtime on the server, the Quarkus reference server in ADK's samples, does the extraction.
Client side: tracing and propagation
The a2a-java repository ships OpenTelemetry modules under extras/opentelemetry. Per its README, a2a-extras-opentelemetry-client wraps the client transport to create a span for every A2A operation, and a2a-extras-opentelemetry-client-propagation injects W3C trace context headers into outbound requests. Both use the group id org.a2aproject.sdk, are discovered with ServiceLoader, and activate only when their keys are present in the transport configuration's parameters:
// static imports, per the module README:
// OTEL_TRACER_KEY from OpenTelemetryClientTransportWrapper
// OTEL_OPEN_TELEMETRY_KEY from OpenTelemetryClientPropagatorTransportWrapper
// The README puts the keys in "the transport configuration parameters"; check which
// config type exposes setParameters in your SDK version.
JSONRPCTransportConfig transportConfig = new JSONRPCTransportConfig();
transportConfig.setParameters(Map.of(
OTEL_TRACER_KEY, openTelemetry.getTracer("orchestrator"),
OTEL_OPEN_TELEMETRY_KEY, openTelemetry)); // enables header injection
Client client = Client.builder(card)
.withTransport(JSONRPCTransport.class, transportConfig)
.clientConfig(new ClientConfig.Builder()
.setStreaming(card.capabilities().streaming())
.build())
.build();
BaseAgent remote = RemoteA2AAgent.builder()
.name(card.name()).agentCard(card).a2aClient(client).streaming(true).build();Two details from the README matter when you read traces. The propagation wrapper has priority 500 and the tracing wrapper 600, so headers are injected before the client span is created. The parent recorded on the server is therefore whatever span was current when the call was made, typically the remote proxy's invoke_agent span, and the A2A client span appears as its sibling rather than its parent. That is fine for analysis but surprises people who expect a strict client-to-server chain. Second, the extras are versioned with the SDK; package names have moved between releases, so take imports from the version you depend on.
If you cannot use the extras, inject manually with the OpenTelemetry API in whatever hook your transport offers for outgoing headers: openTelemetry.getPropagators().getTextMapPropagator().inject(Context.current(), headers, Map::put).
Server side: extraction belongs to the runtime
On the server, a2a-extras-opentelemetry-server registers an OpenTelemetryRequestHandlerDecorator through its bundled beans.xml and creates spans for A2A methods. The README is explicit that it does not extract trace context from incoming requests; the runtime must. With Quarkus, adding the Quarkus OpenTelemetry extension makes its HTTP server layer extract traceparent and start a server span. The README also notes that context must be carried across asynchronous work, which the reference server does with a context-aware ManagedExecutor. If you host ADK's AgentExecutor in another framework, check both halves yourself: is the header extracted, and is the context current on the thread that calls execute(RequestContext, EventQueue)?
Request and response bodies are not recorded by default. The system properties org.a2aproject.sdk.server.extract.request and org.a2aproject.sdk.server.extract.response turn that on, and the same privacy argument applies as for ADK's own ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS: useful in development, a data-protection decision in production.
Thread hops inside ADK
OpenTelemetry keeps the current context in a thread-local. RxJava operators that switch schedulers do not copy it, so a span made current in one callback can be absent when RemoteA2AAgent reaches the SDK client on another thread. There are two remedies. The OpenTelemetry Java agent includes RxJava instrumentation that carries context across operators; for library setups, the opentelemetry-rxjava-3.0 instrumentation module provides the same through an assembly hook enabled once at startup. Check the artifact name for your OpenTelemetry version. The fallback is explicit capture at a boundary you own:
// Capture on the thread where the right span is current, restore where the work runs.
Context captured = Context.current();
return Flowable.defer(() -> {
try (Scope ignored = captured.makeCurrent()) {
return delegate.runAsync(ctx); // the client call now sees the caller's span
}
});A try block only covers the synchronous part of the subscription, so prefer the instrumentation when you can. Whichever you choose, do not trust it until the test in the next section passes.
A test that proves the link
The useful assertion is narrow: the traceparent that reaches the server carries the caller's trace id and a parent span id that belongs to a span the caller exported. The test below stands up a JDK HttpServer that records the header and answers with a canned JSON-RPC response, and uses OpenTelemetry's in-memory exporter. sendOne is a two-line helper you write for your SDK version that sends one message through the configured client.
@Test
void traceparentReachesRemoteAgent() throws Exception {
InMemorySpanExporter spans = InMemorySpanExporter.create();
OpenTelemetrySdk otel = OpenTelemetrySdk.builder()
.setTracerProvider(SdkTracerProvider.builder()
.addSpanProcessor(SimpleSpanProcessor.create(spans)).build())
.setPropagators(ContextPropagators.create(W3CTraceContextPropagator.getInstance()))
.build();
AtomicReference<String> seen = new AtomicReference<>();
HttpServer server = cannedA2AServer(seen); // records "traceparent", returns a task
Client client = clientWithOtel(server, otel); // the config shown above
Span caller = otel.getTracer("test").spanBuilder("caller").startSpan();
try (Scope s = caller.makeCurrent()) {
sendOne(client, "ping");
} finally {
caller.end();
}
String[] tp = seen.get().split("-"); // 00-traceid-parentid-flags
assertEquals(caller.getSpanContext().getTraceId(), tp[1]);
Set<String> exported = spans.getFinishedSpanItems().stream()
.map(d -> d.getSpanId()).collect(toSet());
assertTrue(exported.contains(tp[2]), "server parent must be a caller-side span");
assertEquals("01", tp[3]); // sampled decision travelled
}Then write a second test that drives the same call through an ADK Runner with a scripted model that transfers to the remote agent. That version catches the thread-hop loss, which the direct test cannot, because it is the RxJava pipeline that drops context.
Streaming and long-running tasks
Streaming changes what the spans measure. For message/stream, the server decorator's span covers only the creation of the Flow.Publisher; the README says its duration does not reflect the streaming operation and sets its response attribute to Stream publisher created. A dashboard of server A2A span latency will therefore show streaming calls as nearly instant. On the client, the extras create child spans per streaming event, linked to the parent. Measure end-to-end latency from the client side, or from the ADK invoke_agent span on the server, which covers the specialist's actual run. For time to first token across the hop, see the A2A client streaming recipe.
Long-running tasks break the single-trace model altogether. A task that pauses for input and resumes an hour later arrives as a new request, often in a new trace. Record the A2A taskId and contextId as span attributes and in log context on both sides, and add a span link to the original trace when you have it, so one task can be followed across its traces.
Sampling and trust boundaries
If the specialist ignores the incoming sampled flag, its decision is independent of the caller's. A 10 percent caller and a server that samples 10 percent at random keep complete cross-service traces for about 1 percent of requests, and store many half-traces besides. Use a parent-based sampler everywhere except at the edge: OTEL_TRACES_SAMPLER=parentbased_traceidratio with OTEL_TRACES_SAMPLER_ARG=0.1 samples 10 percent of new root traces and honours the caller's flag for everything downstream. The OpenTelemetry SDK's default sampler is already parent-based, so this mostly means not overriding it with a plain ratio sampler. Both services must also export to the same backend, or to backends you can query together, before a joined trace can be seen.
At an organisational trust boundary, think before propagating. A third-party agent learns your trace ids and can set the sampled flag on your behalf. Do not send baggage across it, and consider starting a new root on the far side with a link back rather than a parent.
Worked example: the cost on the wrong side
An illustrative case, with made-up but typical timings. A support orchestrator's p95 latency rises from about 6 to 11 seconds. Before propagation was fixed, its trace ended in a 9-second A2A client span with no children, and the billing specialist's traces looked healthy. After fixing the thread hop and confirming the test above, one joined trace showed the specialist's invoke_agent span at 8.6 seconds containing four call_llm spans instead of the usual two: a new tool returned a paginated result and the model asked for every page. The server's A2A span read 3 milliseconds because the call was streamed, which is why the server dashboard never alarmed. The fix was a page-size argument on the tool; the lesson was that only the joined trace placed the cost on the right side.
Failure modes
- Two root traces per request. Header not injected, not extracted, or context lost on a thread hop. The header test and the runner test isolate which.
- Right trace, wrong parent. Spans attach to the orchestrator's root instead of the proxy span. Usually a captured context taken too early; capture where the proxy's span is current.
- Broken traces at a fraction of requests. Independent ratio samplers on each side. Switch to parent-based sampling.
- Fast-looking streaming calls. The server span measures publisher creation. Alert on client-side or
invoke_agentdurations. - Prompts in a shared trace backend. Content capture left on in one service. Set
ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS=falseand leave A2A body extraction off in production on both sides. - Errors that do not show on the caller. A2A failures can arrive in-band as a failed task state rather than an exception; record task state on the client span, as A2A error propagation explains.
What to do next
- Add the client tracing and propagation extras to every service that calls a remote agent, and the server extra plus your runtime's OpenTelemetry extension to every service that hosts one.
- Write the header test, then the runner-level test with a scripted model.
- Enable RxJava context propagation through the Java agent or the library module.
- Set parent-based sampling everywhere and ratio sampling only at the edge.
- Put
taskId,contextId, trace id and invocation id on spans and log lines on both sides. - Move streaming latency alerts off the server A2A span.
- Review content capture in both services, using prompt and completion tracing for what each attribute holds, and the A2A integration guide for the wiring.