A2A is specified over ordinary HTTP: a JSON agent card fetched with GET, JSON-RPC 2.0 requests sent with POST, and Server-Sent Events for streams. In ADK Java you rarely touch any of it: RemoteA2AAgent calls the a2a-java client, which calls a transport, which calls an HTTP client. But timeouts, authentication, tracing, proxies and most production incidents live at exactly that bottom layer, and the defaults there are not production defaults.
This page walks the stack from the agent down to the bytes and back up the server: what crosses the wire, how errors are encoded, what the SDK's HTTP client does and does not configure, where to plug in timeouts, auth and trace propagation, and how to keep streams alive through load balancers. Names are from a2a-java 0.3.2.Final, which google-adk-a2a pins on adk-java main as of 2026-10-08; differences in the 1.0 specification are called out where they matter.
The stack, layer by layer
Five client layers sit between your orchestrator and the socket. RemoteA2AAgent turns an ADK invocation into an A2A Message and turns responses back into ADK events. The SDK Client chooses a transport that both it and the agent card support and delivers events to consumers. JSONRPCTransport builds the JSON-RPC envelope and parses responses and SSE frames. Each ClientCallInterceptor registered on the transport config can rewrite the payload and headers. Finally an A2AHttpClient performs the call; the default is JdkA2AHttpClient, a thin wrapper over java.net.http.HttpClient.
On the server, the reference JSON-RPC app (Quarkus) exposes a single POST / route annotated @Authenticated and a GET route for /.well-known/agent-card.json. The handler reads the method field, sends message/stream and tasks/resubscribe down the SSE path and everything else down the plain JSON path, and the ADK AgentExecutor runs the agent.
The seam that matters is A2AHttpClient. It is a small interface: createGet(), createPost() and createDelete() return builders with url, addHeader and addHeaders; GET and POST builders can either return an A2AHttpResponse (status(), success(), body()) or run an SSE call with getAsyncSSE / postAsyncSSE, which take a per-message consumer, an error consumer and a completion callback. Everything below the transport is replaceable through it.
What crosses the wire
A non-streaming turn is one POST with a JSON-RPC body. In the 0.3 binding:
POST / HTTP/2
content-type: application/json
authorization: Bearer eyJ...
{"jsonrpc":"2.0","id":"7f1c","method":"message/send",
"params":{"message":{"role":"user","messageId":"m-41",
"contextId":"ctx-9","parts":[{"kind":"text","text":"Is 1987 prime?"}]}}}
HTTP/2 200
content-type: application/json
{"jsonrpc":"2.0","id":"7f1c","result":{"kind":"task","id":"t-3","contextId":"ctx-9",
"status":{"state":"completed"},"artifacts":[{"artifactId":"a-1",
"parts":[{"kind":"text","text":"Yes, 1987 is prime."}]}]}}A streaming turn uses message/stream with the same params. The response is text/event-stream: each event is a data: line holding a complete JSON-RPC response whose result is a task, a message, a status update or an artifact update, with a blank line between events. The stream ends when a status update arrives with final set, or when the server closes it.
The 1.0 specification keeps this shape but renames the methods to PascalCase (SendMessage, SendStreamingMessage, GetTask), adds an A2A-Version request header (an empty value is read as 0.3) and a VersionNotSupportedError (-32009). Its HTTP+JSON binding uses paths such as POST /message:send and GET /tasks/{id}. A caller pinned to 0.3.x talking to a 1.0-only server fails at the method name, so check versions before blaming the network.
Status codes versus protocol errors
HTTP status and protocol errors are separate channels, and mixing them up is the most common bug in hand-written A2A clients and dashboards. In the 0.3.2 reference server, a JSON-RPC failure, including a parse error or an unknown task, is written as HTTP 200 with an error object in the body. Only the layers around the handler produce non-200 statuses: the @Authenticated route yields 401 or 403, and proxies and platforms yield 404, 413, 502, 503 and 504.
| Signal | Where it comes from | Typical cause |
|---|---|---|
200 + -32001 TaskNotFoundError | handler | polling a task on a replica that never had it, or after it was evicted |
200 + -32002 TaskNotCancelableError | handler | cancelling a task already in a terminal state |
200 + -32004 UnsupportedOperationError | handler | calling a method the agent does not implement |
200 + -32005 ContentTypeNotSupportedError | handler | a part MIME type outside the card's input modes |
200 + -32600 to -32603 | handler | malformed request, bad params, internal error |
| 401 / 403 | route auth | missing or rejected credential |
| 502 / 503 / 504 | proxy or platform | replica down, overloaded, or a request timeout |
So an HTTP-level dashboard on the callee shows 100% success while every call fails. Count JSON-RPC error codes from the response bodies, or from server callbacks, as a separate metric. On the client side, JdkA2AHttpClient's SSE path checks for 401 and 403 before reading the body and signals an IOException with an authentication or authorization message to the error consumer; other non-success statuses on a stream surface as an IOException carrying the status and body. Both reach you as a failed call, not as an A2A task state.
What the default HTTP client leaves out
Read what JdkA2AHttpClient's no-argument constructor builds: an HttpClient with Version.HTTP_2 and Redirect.NORMAL, and nothing else. Three consequences follow.
- No connect timeout. A connect to a black-holed address waits for the operating system's TCP timeout, which can be minutes.
- No request timeout. The request builders set none either, so a non-streaming
message/sendto a hung agent waits until some proxy gives up. Your orchestrator's own turn budget is not enforced at this layer. - Redirects are followed silently (except HTTPS to HTTP). A card
urlthat points at a host which redirects costs an extra round trip on every call and can drop headers you expected to arrive; make the card URL final.
HTTP/2 preference also changes balancing: one multiplexed connection per caller means an L4 balancer pins that caller to one backend.
A client with timeouts
Because the transport accepts any A2AHttpClient, the fix is your own implementation with explicit timeouts. The interface is small enough to implement directly. This sketch shows the non-streaming POST path, the one that needs a timeout most:
public final class TimedA2AHttpClient implements A2AHttpClient {
private final HttpClient http = HttpClient.newBuilder()
.version(HttpClient.Version.HTTP_2)
.connectTimeout(Duration.ofSeconds(3))
.followRedirects(HttpClient.Redirect.NEVER) // a redirect here is a config bug
.build();
private final Duration callTimeout; // whole non-streaming turn
public TimedA2AHttpClient(Duration callTimeout) { this.callTimeout = callTimeout; }
@Override public PostBuilder createPost() { return new Post(); }
// createGet()/createDelete() follow the same pattern; the SSE methods use
// http.sendAsync with a line-by-line body subscriber and no request timeout.
private final class Post implements PostBuilder {
private String url, body = "";
private final Map<String, String> headers = new LinkedHashMap<>();
public PostBuilder url(String u) { url = u; return this; }
public PostBuilder addHeader(String n, String v) { headers.put(n, v); return this; }
public PostBuilder addHeaders(Map<String, String> h) { headers.putAll(h); return this; }
public PostBuilder body(String b) { body = b; return this; }
public A2AHttpResponse post() throws IOException, InterruptedException {
HttpRequest.Builder b = HttpRequest.newBuilder(URI.create(url))
.timeout(callTimeout)
.header("Content-Type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString(body));
headers.forEach(b::header);
HttpResponse<String> r = http.send(b.build(), HttpResponse.BodyHandlers.ofString());
return new A2AHttpResponse() {
public int status() { return r.statusCode(); }
public boolean success() { return r.statusCode() >= 200 && r.statusCode() < 300; }
public String body() { return r.body(); }
};
}
public CompletableFuture<Void> postAsyncSSE(Consumer<String> onMessage,
Consumer<Throwable> onError, Runnable onDone) { /* see comment above */ throw new UnsupportedOperationException(); }
}
}Do not put a request timeout on SSE calls: HttpRequest.timeout bounds the time to response headers in the JDK, but proxies and the stream's own lifetime are better bounded by an idle watchdog in your consumer that cancels the stream when no event has arrived for, say, twice the server's longest quiet period. Reuse one client instance per process; each HttpClient owns a connection pool and selector thread.
Interceptors for auth and tracing
Interceptors are where cross-cutting headers belong: credentials, W3C trace context and routing keys. They run for every method, receive the method name, and must return a new PayloadAndHeaders:
public final class AuthAndTraceInterceptor extends ClientCallInterceptor {
private final Supplier<String> tokens; // cached, refreshed before expiry
public AuthAndTraceInterceptor(Supplier<String> tokens) { this.tokens = tokens; }
@Override
public PayloadAndHeaders intercept(String method, Object payload, Map<String, String> headers,
AgentCard card, ClientCallContext ctx) {
Map<String, String> h = new HashMap<>(headers);
h.put("Authorization", "Bearer " + tokens.get());
SpanContext sc = Span.current().getSpanContext(); // OpenTelemetry API
if (sc.isValid()) {
h.put("traceparent", "00-" + sc.getTraceId() + "-" + sc.getSpanId() + "-"
+ sc.getTraceFlags().asHex());
}
return new PayloadAndHeaders(payload, h);
}
}
JSONRPCTransportConfig cfg = new JSONRPCTransportConfigBuilder()
.httpClient(new TimedA2AHttpClient(Duration.ofSeconds(45)))
.addInterceptor(new AuthAndTraceInterceptor(tokenCache::current))
.build();Fetch tokens outside the interceptor and cache them; an interceptor that calls an identity provider on every request adds its latency and failure rate to every agent call. Read the card's securitySchemes to decide which credential to send, and never send a token to a card URL whose host you did not expect: card resolution is an outbound fetch to whatever the configuration names.
Streams through proxies and platforms
Streams fail in the infrastructure, not the code. Three settings decide whether an SSE response survives the path:
- Response buffering. A proxy that buffers responses holds every event until the stream ends, so the client sees nothing and then everything. In NGINX set
proxy_buffering offfor the agent location, or have the server sendX-Accel-Buffering: no. - Idle and request timeouts. Load balancers close connections that carry no bytes for their idle timeout, and serverless platforms cap total request time: Cloud Run's request timeout defaults to 5 minutes and can be raised to 60. A tool call that keeps the agent silent for longer than the idle timeout drops the stream. Raise the limit, or confirm your server version emits keep-alive comments.
- Compression. Gzip on
text/event-streamcan delay flushing; exclude that content type.
When a stream drops mid-task, the task usually keeps running on the server. Recover with tasks/resubscribe on the same replica, or poll tasks/get, rather than resending the message, which starts a second task.
Worked example: a stalled summary
An orchestrator on Cloud Run gives each user request 60 seconds. It delegates a summarisation to a remote agent over HTTPS through an internal load balancer.
Before tuning: the remote agent's model call stalled. JdkA2AHttpClient had no request timeout, so the orchestrator's thread waited until the orchestrator's own Cloud Run request hit its limit and the user got a 504 with no partial answer. The callee's HTTP dashboard showed 0% errors throughout, because its JSON-RPC failures had all been HTTP 200.
After tuning: connect timeout 3 s, call timeout 45 s, leaving 15 s for the orchestrator to answer from what it has. The stalled call fails at 45 s with an HttpTimeoutException, which reaches the orchestrator as a failed delegation (check how your ADK version surfaces it in events); the root agent's instruction tells it to say the summary is unavailable and offer the raw sources. A metric counting JSON-RPC error codes by method now shows the callee's real failure rate, and the traceparent header joins both services' spans into one trace.
Failure modes
- Unbounded waits. Default client, no timeouts. Fix: your own
A2AHttpClientwith connect and call timeouts. - Green dashboards, failing calls. JSON-RPC errors are HTTP 200. Fix: count error codes from bodies.
- Buffered streams. All events arrive at once. Fix: disable proxy buffering and compression for SSE.
- Card URL drift. The card advertises
localhostor an internal name, or a redirecting host. Fix: generate the card URL from deployment config and test it from a caller's network. - Version mismatch. A 0.3 client against a 1.0-only server fails on method names. Fix: pin versions together and test across the boundary.
- Duplicate tasks on retry. A retried send after a dropped stream. Fix: resubscribe or poll instead.
Choosing a transport
JSON-RPC over HTTP is the transport ADK's samples use and the simplest to debug with curl. The HTTP+JSON binding maps more naturally onto gateways and per-path policies; google-adk-a2a already depends on the SDK's REST transport module, and the card's preferredTransport tells clients which to use. gRPC gives typed contracts and efficient streaming but needs HTTP/2 end to end, which some proxies and serverless front ends do not provide. Choose one per agent, advertise it in the card, and test the whole path, not just the code.
What to do next
- Replace
JdkA2AHttpClientwith anA2AHttpClientthat sets connect and per-call timeouts inside your turn budget. - Add an interceptor for credentials and
traceparent; cache tokens outside it. - Count JSON-RPC error codes per method on both sides, separate from HTTP status.
- Audit every proxy between caller and agent for SSE buffering, compression and idle timeouts.
- Verify the card
urlfrom a caller's network and make sure it does not redirect. - Pin the A2A SDK and protocol version on both ends and test a 0.3 to 1.0 mismatch before you upgrade one side.
- Continue with ADK Java and A2A, the client streaming recipe, A2A error propagation, A2A agent identity verification and Cloud Run deployment.