An orchestrator that delegates to a downstream A2A agent sees one URL: the url field in that agent's card. Behind it there are usually several replicas, and how calls are spread across them decides whether multi-turn conversations keep their context, whether a client that polls a task finds it, and whether one replica melts while three idle. Ordinary HTTP load balancing assumes every request is independent. A2A requests are not, and the defaults in the Java stack make the mismatch worse rather than better.
This page explains where an A2A server keeps state, why round-robin and plain Kubernetes Services go wrong, and three designs that work: affinity routing on a key the caller sets, interchangeable replicas backed by shared stores, and least-loaded picking for the stateless part of the traffic. The Java names are from the a2a-java SDK at 0.3.2.Final, the version google-adk-a2a pins on adk-java main as of 2026-10-08, so method names are the 0.3 slash style (message/send, tasks/get). Spec 1.0 renamed them; the routing logic is unchanged.
Where an A2A server keeps state
Start from the server. In the a2a-java reference server, a request reaches a JSON-RPC handler that keeps three kinds of state, and ADK adds a fourth:
| State | Default in 0.3.x / ADK | Who needs it |
|---|---|---|
| Task records (status, history, artifacts) | InMemoryTaskStore | tasks/get, tasks/cancel, follow-up message/send with a taskId |
| Live event queues for running tasks | InMemoryQueueManager | message/stream and tasks/resubscribe |
| Push-notification configs | InMemoryPushNotificationConfigStore | the replica that later sends the webhook |
| ADK session (events, state) | whatever sessionService the executor was built with; samples use InMemorySessionService | every turn in the same contextId |
With in-memory defaults, all four live in one JVM. A tasks/get that lands on a different replica gets TaskNotFoundError (JSON-RPC code -32001) for a task that exists. A second turn in the same conversation that lands elsewhere starts with an empty ADK session, so the remote agent forgets what it was told one turn ago, and nothing errors: the answer is just worse. That silent failure is the expensive one.
The requests a client sends fall into two groups. Calls that start something new, a first message/send or message/stream with no contextId, can go anywhere. Everything else is addressed to existing state, by contextId or taskId, and must reach the replica that holds it, or a replica that can read it from a shared store.
Why the defaults balance badly
Two defaults defeat naive balancing. The first is connection reuse. The SDK's JdkA2AHttpClient builds one java.net.http.HttpClient with HttpClient.Version.HTTP_2. When HTTP/2 is actually negotiated (ALPN over TLS, or an accepted h2c upgrade on plain HTTP), every request from that caller is multiplexed over a single TCP connection, and a Kubernetes ClusterIP Service balances connections, not requests. All of one orchestrator's calls then go to one pod. Even on HTTP/1.1 the client pools keep-alive connections, so an L4 balancer sees a few long-lived flows rather than a stream of requests.
The second is variance. An A2A call is an agent turn: one may finish in 800 ms, the next may run tools for 40 seconds, and a streaming call holds a connection open for its whole life. Round-robin assumes requests cost about the same; here they differ by 50x, so equal request counts produce very unequal load. The signal that matters is outstanding work per replica, not requests per second.
Choosing and publishing the affinity key
Affinity needs a key. contextId is the right one for conversations: every turn of one conversation carries it, and the tasks inside it are created by the replica that served those turns. taskId is the fallback for calls that carry only a task, such as tasks/get and tasks/cancel.
There is a trap at the start of every conversation. If the caller sends its first message without a contextId, the server mints one, and the first request was routed on an empty key: effectively at random. The second turn hashes the new contextId and may land on a different replica from the one that holds the session. The fix is to mint the contextId on the caller before the first send; the message schema lets a client supply one, but confirm your server accepts it. If your caller cannot do that, you need shared session storage, covered below.
The interceptor hook in the SDK is the clean place to publish the key. ClientCallInterceptor receives the method name and the JSON-RPC request object before it is serialised, and returns a PayloadAndHeaders:
import io.a2a.client.transport.spi.interceptors.ClientCallContext;
import io.a2a.client.transport.spi.interceptors.ClientCallInterceptor;
import io.a2a.client.transport.spi.interceptors.PayloadAndHeaders;
import io.a2a.spec.*;
import java.util.HashMap;
import java.util.Map;
/** Publishes the routing key as a header so an L7 proxy can hash on it. */
public final class RouteKeyInterceptor extends ClientCallInterceptor {
static final String HEADER = "x-a2a-route-key";
@Override
public PayloadAndHeaders intercept(String method, Object payload, Map<String, String> headers,
AgentCard card, ClientCallContext ctx) {
String key = null;
if (payload instanceof JSONRPCRequest<?> req) {
Object params = req.getParams();
if (params instanceof MessageSendParams msp && msp.message() != null) {
Message m = msp.message();
key = m.getContextId() != null ? m.getContextId() : m.getTaskId();
} else if (params instanceof TaskQueryParams q) {
key = q.id(); // tasks/get
} else if (params instanceof TaskIdParams t) {
key = t.id(); // tasks/cancel, tasks/resubscribe
}
}
Map<String, String> out = new HashMap<>(headers);
if (key != null) out.put(HEADER, key);
return new PayloadAndHeaders(payload, out);
}
}Hashing a taskId only lands on the right replica if the task was created under the same hash, so for task-only calls prefer the contextId you already know: keep a small map from task to context in the caller, filled from the response events, and look it up here. Register the interceptor on the transport, which is where 0.3.x keeps the interceptor list:
Client client = Client.builder(card)
.withTransport(JSONRPCTransport.class,
new JSONRPCTransportConfigBuilder()
.httpClient(new JdkA2AHttpClient())
.addInterceptor(new RouteKeyInterceptor()))
.build();
Balancing in the proxy
With the header in place, the balancing itself belongs in an L7 proxy that terminates each request, so HTTP/2 multiplexing no longer pins a caller to a pod. In Istio that is a DestinationRule with a consistent hash on the header:
apiVersion: networking.istio.io/v1
kind: DestinationRule
metadata:
name: research-agent
spec:
host: research-agent.agents.svc.cluster.local
trafficPolicy:
loadBalancer:
consistentHash:
httpHeaderName: x-a2a-route-key
connectionPool:
http:
idleTimeout: 600s # longer than your longest silent SSE gapRequests without the header (first sends without a client-minted context, card fetches) hash to an arbitrary host, which is fine because they address no state. Envoy's ring hash and Maglev balancers, an NGINX hash $http_x_a2a_route_key consistent upstream, and most API gateways can do the same; the requirement is only request-level, header-keyed hashing.
If you run no proxy, the same logic can live in the caller as a custom A2AHttpClient that rewrites the host of each request to a replica chosen from a discovered endpoint list. That puts discovery, health and ring maintenance into every calling service, which is why the proxy is usually the better home. Endpoint discovery through DNS (headless Service A records, or SRV records) is generic infrastructure; A2A itself only defines the card url.
Interchangeable replicas and least-request
Affinity is the cheap answer. The robust one is to make replicas interchangeable for everything that is stored: implement the SDK's TaskStore and PushNotificationConfigStore interfaces over your database, and build the ADK executor with a durable sessionService instead of the in-memory one. Then tasks/get, follow-up turns and push delivery work from any replica, and the balancer can pick purely on load.
One thing stays local: the live event queue of a running task. In 0.3.x the queue manager is in memory, so a tasks/resubscribe for a task still running has to reach the replica that is executing it. Keep affinity for streaming and resubscribe even after you add shared stores, and treat a resubscribe that lands elsewhere as a signal to fall back to polling tasks/get.
For the load-picking part, power-of-two-choices on outstanding requests is the standard algorithm: sample two healthy replicas, send to the one with fewer calls in flight. It avoids the herding of always-pick-the-least and needs only local counts:
final class P2CPicker {
private final List<Replica> replicas; // refreshed from discovery
Replica pick() {
var r = ThreadLocalRandom.current();
Replica a = replicas.get(r.nextInt(replicas.size()));
Replica b = replicas.get(r.nextInt(replicas.size()));
return a.inFlight.get() <= b.inFlight.get() ? a : b;
}
}
// around each call: replica.inFlight.incrementAndGet(); try { send } finally { decrementAndGet(); }In a mesh, the equivalent is simple: LEAST_REQUEST in the DestinationRule, which Envoy implements with the same two-choice sampling.
Worked example: four replicas, 120 conversations
A research agent runs on four replicas, each sized for about 30 concurrent turns. Three orchestrator pods call it, holding 120 concurrent conversations between them. Turns take 2 seconds when the agent answers from memory and up to 40 seconds when it searches and summarises.
Behind a ClusterIP Service with HTTP/2: each orchestrator pod holds one connection, so three replicas get roughly 40 conversations each and the fourth gets none. The three busy replicas run 33% above their sizing, p95 turn latency climbs, and the autoscaler adds a fifth replica that also receives no traffic, because no new connection is opened.
Round-robin per request through a proxy, in-memory state: load is even, about 30 per replica, but turn two of each conversation has a 1-in-4 chance of reaching the replica that holds its session. Three quarters of follow-up turns run without context, and about 75% of tasks/get polls return -32001.
Consistent hash on contextId: every turn finds its session. Load is uneven by luck of the hash: with 120 keys on 4 replicas, a spread of roughly 24 to 36 is normal. Adding a fifth replica remaps about one fifth of the keys, so about 24 conversations lose their in-memory session at the scale-up; that is the price of affinity with local state.
Shared stores plus least-request, affinity only for streams: load evens to 30 each, scale-up moves nothing that matters, and only running streams stay pinned. This is the target for an agent that many teams call.
Operating it
Scale and drain deliberately. When a replica leaves the ring, the conversations hashed to it move. Give pods a pre-stop delay and a termination grace period longer than your longest turn, stop advertising readiness first so the proxy removes them, and let running tasks reach a terminal state before the JVM exits. Readiness should reflect whether the replica can take a new turn (model quota, thread pool headroom), not only whether the agent card endpoint answers.
Measure what the balancer cannot see. Export in-flight tasks per replica, turn duration by replica, and the rate of -32001 responses on the callee. A nonzero TaskNotFoundError rate on a healthy service almost always means routing, not lost data. On the caller, log the route key with the trace ID so a bad answer can be traced to the replica that produced it.
Retries need care. A retried message/send after a timeout may start a second task on another replica while the first is still running; if the agent has side effects, reuse the same messageId and make the server-side handling idempotent before you enable retries at the proxy.
Failure modes
- All traffic on one pod. HTTP/2 or keep-alive behind an L4 Service. Fix: request-level L7 balancing, or cap connection lifetime so connections rebalance.
- Context amnesia without errors. Follow-up turns land on a replica with an empty in-memory session. Fix: affinity on
contextIdor a durable session service; detect it by logging session event counts per turn. - First-turn split. The server minted the
contextId, so turn one and turn two hashed differently. Fix: mint it on the caller. - Ring churn at scale events. Autoscaling every few minutes reshuffles keys continuously. Fix: slower scale-down, larger bounded-load hash rings, or shared state.
- Resubscribe to the wrong replica. Shared stores fixed polling but not live queues. Fix: keep affinity for
tasks/resubscribe. - Hot keys. One context that fans out many long tasks overloads its replica. Fix: bounded-load consistent hashing, or move that caller to shared state and least-request.
Trade-offs
| Design | Good at | Costs |
|---|---|---|
| Consistent hash on route key, local state | simple, no new storage, keeps sessions warm | uneven load, lost sessions on scale events |
| Shared stores + least-request | even load, painless scaling, any replica serves reads | a database on the hot path, store implementations to own |
| Client-side picker | no proxy, per-caller policy | discovery and health logic in every caller |
| Plain Service, no affinity | nothing | only correct for single-turn, non-streaming agents |
What to do next
- List where your A2A server keeps tasks, queues, push configs and ADK sessions; anything in memory needs affinity.
- Check whether your callers reach the agent through an L4 Service with HTTP/2; if so, move to request-level L7 balancing.
- Mint
contextIdon the caller and add a route-key interceptor like the one above. - Configure consistent hashing on that header, with an idle timeout above your longest SSE gap.
- Plan the move to a durable session service and
TaskStore; then switch non-streaming calls to least-request. - Alert on the callee's
-32001rate and on in-flight spread across replicas. - Read ADK Java and A2A for the client and executor wiring, the client streaming recipe for resubscribe handling, A2A multi-region topology for routing across regions, health, liveness and readiness for drain signals and Kubernetes deployment for the pod settings.