An A2A agent that delegates work to another agent is making a network call to a service it probably does not own. For the two halves of that call to appear in one trace, the caller must send its current trace context and the receiver must read it and continue from it. The A2A project's enterprise guidance says clients and servers should take part in distributed tracing, for example by using OpenTelemetry to propagate trace context through standard HTTP headers such as W3C Trace Context. The protocol specification itself (version 1.0 at the time of writing, checked on 2026-10-02) defines no tracing field, so everything about how context crosses an agent boundary is a convention the two sides must share.
Why tracing matters for multi-agent systems, and what a good agent trace looks like, is covered in A2A distributed tracing architecture. This page is about the wire: what is in the traceparent header, how to validate it, where to inject and extract it for each A2A interaction style, and how to keep a trace meaningful when a task finishes a day after the request that started it.
What traceparent actually carries
W3C Trace Context defines two HTTP headers. traceparent is fixed-format and carries the identity of the caller's span; tracestate carries optional vendor-specific data. A traceparent value has four dash-separated fields, all lowercase hexadecimal: a two-character version (currently 00), a 32-character trace id (16 bytes), a 16-character parent id (8 bytes, the id of the span that made the call) and a two-character flags byte. An all-zero trace id or parent id is invalid, and version ff is forbidden.
The flags byte is a bit field, so test bits with a mask, never with equality. Bit 0 (0x01) is the sampled flag: the caller may have recorded its span. Trace Context Level 2, still a W3C Candidate Recommendation Draft, adds bit 1 (0x02) to say that at least the right-most 7 bytes of the trace id are random, which lets downstream samplers make consistent decisions from the id alone. The remaining bits are reserved and must be zero when you set them.
tracestate is a list of up to 32 key=value members, each key owned by a tracing vendor or system, for example a sampling probability or a vendor's own span id. Separately, the W3C Baggage header carries application key-value pairs that every downstream hop can read. Baggage is the most likely to leak data to a third-party agent.
Validating what you receive
The rule for a receiver is simple and strict: if traceparent is invalid, ignore it, start a new trace and drop tracestate. For a version higher than 00, parse the first three fields and the sampled bit and ignore anything after them; this is how the format stays forward compatible. OpenTelemetry's propagators already implement these rules, so in practice you call extract and get either a valid remote context or an empty one. The parser below shows what the propagator does.
import re
_TP = re.compile(r"^([0-9a-f]{2})-([0-9a-f]{32})-([0-9a-f]{16})-([0-9a-f]{2})(-.*)?$")
def parse_traceparent(value):
"""Return (trace_id, parent_id, flags) or None if the header must be ignored."""
m = _TP.match(value.strip())
if not m:
return None
version, trace_id, parent_id, flags, rest = m.groups()
if version == "ff":
return None
if version == "00" and rest:
return None # version 00 has exactly four fields
if trace_id == "0" * 32 or parent_id == "0" * 16:
return None
return trace_id, parent_id, int(flags, 16)
CASES = {
"sampled": "00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01",
"sampled+random": "00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-03",
"uppercase hex": "00-4BF92F3577B34DA6A3CE929D0E0E4736-00f067aa0ba902b7-01",
"zero trace-id": "00-00000000000000000000000000000000-00f067aa0ba902b7-01",
"future version": "01-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01-xyz",
}
for label, v in CASES.items():
r = parse_traceparent(v)
print(f"{label:15} -> " + ("ignore, start a new trace" if r is None
else f"sampled={r[2] & 1} random={r[2] >> 1 & 1}"))Running it prints:
sampled -> sampled=1 random=0
sampled+random -> sampled=1 random=1
uppercase hex -> ignore, start a new trace
zero trace-id -> ignore, start a new trace
future version -> sampled=1 random=0Upper-case hex, usually from a hand-written client, is a common invalid value; the visible symptom is a trace that breaks at exactly that hop.
Injecting and extracting on each A2A binding
A2A 1.0 defines three protocol bindings: JSON-RPC over HTTP, gRPC, and HTTP+JSON. For both HTTP bindings, trace context goes in HTTP request headers, exactly as for any instrumented HTTP call. For gRPC, it goes in request metadata under the lowercase key traceparent, which the OpenTelemetry gRPC instrumentation injects and extracts for you. Keep it out of the JSON-RPC body: proxies, meshes and auto-instrumentation read headers, not params.
On the client side, open a CLIENT span for the call, then inject the current context into the outgoing headers. The OpenTelemetry GenAI semantic conventions name such a span invoke_agent {gen_ai.agent.name} with gen_ai.operation.name set to invoke_agent. Those agent conventions are at Development stability, so attribute names may still change; pin the semantic conventions version your collector and dashboards expect.
import httpx
from opentelemetry import trace
from opentelemetry.propagate import inject
from opentelemetry.trace import SpanKind
tracer = trace.get_tracer("orchestrator")
PROVIDER = "my-agent-platform" # set per the GenAI conventions version you pin
def call_agent(agent_url: str, agent_name: str, message: dict) -> dict:
# message is built for the A2A version your peers speak
with tracer.start_as_current_span(
f"invoke_agent {agent_name}", kind=SpanKind.CLIENT,
attributes={"gen_ai.operation.name": "invoke_agent",
"gen_ai.provider.name": PROVIDER, "gen_ai.agent.name": agent_name},
) as span:
headers = {"content-type": "application/json"}
inject(headers) # writes traceparent (and tracestate, baggage if set)
body = {"jsonrpc": "2.0", "id": 1, "method": "SendMessage",
"params": {"message": message}}
resp = httpx.post(agent_url, json=body, headers=headers, timeout=30)
span.set_attribute("http.response.status_code", resp.status_code)
return resp.json()On the server side, extract the context from the request headers and start a SERVER span with that context as its parent. If the request carried no valid context, extract returns an empty context and the span becomes a new root, which is the correct behaviour for a first hop.
from opentelemetry import trace
from opentelemetry.propagate import extract, inject
from opentelemetry.trace import SpanKind
tracer = trace.get_tracer("visa-agent")
async def handle_jsonrpc(request):
ctx = extract(dict(request.headers)) # empty context if header absent or invalid
with tracer.start_as_current_span(
"a2a.server SendMessage", context=ctx, kind=SpanKind.SERVER
) as span:
payload = await request.json()
task = await create_task(payload["params"]["message"])
span.set_attribute("a2a.task.id", task.id)
# Save the *current* context with the task so later work can find it.
carrier = {}
inject(carrier)
await task_store.save_trace_context(task.id, carrier)
return task.to_response()Keep A2A's own identifiers separate from tracing: contextId groups one conversation and taskId names one unit of work, both living for minutes or days, while a trace id names one causal tree. Record contextId and the task id as span attributes so you can search for every trace of a conversation, but never derive one from the other. If an intermediary strips headers and you must carry context in the payload, the message's metadata object is the natural place, under a key both sides agree on. That is a convention of your own, not part of the specification, so document it next to your agent card.
Streaming responses
With SendStreamingMessage or SubscribeToTask, the server answers one HTTP request with a stream of events, which on the HTTP bindings means Server-Sent Events. Trace context travels once, on the request. The server's SERVER span covers the request and stays open as long as the stream does, and the client's CLIENT span does the same.
For an hour-long stream that is a problem: SDKs export a span when it ends, so the work is invisible until then and a crash loses it. Bound the span to the work instead of the connection. Close the request span once the task is accepted, and represent each meaningful status or artifact update as its own short span, or as a span event on the span doing the work. A SubscribeToTask call that reconnects to a running task is a new request, so it gets a new span in a trace of its own, and should link to the task's saved context rather than pretend to be its child. Streaming mechanics themselves are in A2A streaming.
Long-running tasks: persist the context with the task
An A2A task can outlive the request that created it by hours. The handling agent replies quickly with a task in a working state, and the actual work happens later, perhaps on a queue worker in another process. The fix is to serialize the context when the task is created, store it with the task, and restore it wherever the task is processed. inject into a plain dictionary produces exactly the strings you need, and extract on that dictionary brings the context back.
A worker that picks the task up within seconds can continue the same trace. Work that resumes hours later, or is triggered by something else such as a human approval, should start a new trace with a span link to the saved context. Otherwise the original trace's duration stretches to a day, tail-based samplers wait for a trace that never seems to finish, and latency percentiles computed from root spans become meaningless.
Push notifications are a request in the other direction
With push notifications, the client registers a webhook url and optional token or authentication details through the push notification config methods, and the server later POSTs task updates to that URL. That POST is a new request, started by the server, often long after the original call. The specification says nothing about trace context on it, so decide your convention explicitly.
The convention that works: the server starts a span for the notification, linked to the task's saved context, and injects its own current context into the webhook request headers. The client's webhook handler extracts that context like any other inbound request, so the notification delivery is one connected trace, and both sides can follow the link back to the trace that created the task. The webhook endpoint must validate the token or authentication before trusting anything in the request, trace headers included. Delivery semantics and retries are in A2A push notifications.
from opentelemetry import context, trace
from opentelemetry.propagate import extract, inject
from opentelemetry.trace import Link, SpanKind
async def finish_and_notify(task_id: str, result: dict):
saved = extract(await task_store.load_trace_context(task_id))
origin = trace.get_current_span(saved).get_span_context()
links = [Link(origin, {"a2a.link.reason": "task resumed"})] if origin.is_valid else []
# New root: the original request finished hours ago.
with tracer.start_as_current_span(
"a2a.task complete", context=context.Context(), links=links, kind=SpanKind.INTERNAL,
attributes={"a2a.task.id": task_id},
):
with tracer.start_as_current_span("a2a.push POST", kind=SpanKind.CLIENT):
headers = {"authorization": f"Bearer {await webhook_token(task_id)}"}
inject(headers) # webhook receiver continues *this* trace
await http.post(await webhook_url(task_id), json=result, headers=headers)Sampling across agents you do not own
With parent-based sampling, the usual default, each service records a span only if the incoming sampled flag says the caller did. Across organizations, that lets a partner agent that sets the flag on every request force your agent to record and export everything.
At a trust boundary, treat inbound context as information, not instruction. Start a new root under your own sampler and link to the caller's context, so the relationship is preserved for anyone who can see both sides. Do the same in reverse: strip baggage and any internal tracestate members before calling an external agent, because both are forwarded verbatim. Tail sampling, where the keep-or-drop decision is made after the whole trace is seen, needs every span of a trace to reach the same collector; that pipeline design is covered in OpenTelemetry pipeline design.
Worked example: a trip planner with three agents
An orchestrator plans a trip by calling three remote agents. The flight agent answers SendMessage synchronously in two seconds. The hotel agent streams options through SendStreamingMessage for twenty seconds. The visa agent accepts a task that needs a human officer's review and finishes the next morning, notifying the orchestrator through a registered webhook.
The trace for the user's request has the orchestrator's root span, two CLIENT spans for the flight and hotel calls with the agents' SERVER spans nested under them, and a third CLIENT span for the visa call that ends as soon as the task is accepted. The visa agent's server span saves the context with the task. The next morning, the visa worker starts a new trace whose first span links to that saved context, and its push POST carries the new trace's traceparent to the orchestrator's webhook, whose handler span joins the new trace.
Root-span latency for trip planning stays an honest 22 seconds instead of 14 hours, and the link leads from the request to the completion trace.
Failure modes
- Broken at one hop: an agent framework or gateway that does not extract headers starts a fresh trace. Check every hop with a test request carrying a known trace id.
- Context in the body only: trace context put in JSON-RPC params is invisible to proxies and auto-instrumentation, so half the spans miss it.
- Day-long traces: resuming a long task as a child of the original request inflates latency and stalls tail samplers.
- Lost on the queue: a task handed to a worker without its saved context produces orphan spans with no way back.
- Forced sampling and leaked baggage: trusting inbound flags and forwarding baggage across an organizational boundary.
What to do next
- Send a request with a fixed, known
traceparentthrough your agent graph and confirm that every hop's spans share that trace id. - Make sure every A2A client and server, on every binding, uses an OpenTelemetry propagator rather than hand-built headers.
- Store serialized trace context with each task, restore it in workers, and use span links for work that resumes after more than a few seconds.
- Decide and document your push notification convention: a linked span on the server and an injected context on the webhook POST.
- At each trust boundary, start new roots with links, and strip baggage and internal tracestate on outbound calls to external agents.
- Record contextId and task id as span attributes, and pin the GenAI semantic conventions version your dashboards depend on.