Putting an ADK Java agent behind the Agent2Agent (A2A) protocol takes about forty lines of wiring: an executor bean, an agent card and a server transport. That wiring is covered in ADK Java + A2A. This page is about what happens after the wiring works. A single execute call decides which ADK user and session a remote caller lands in, whether the second message in a conversation sees the first, what happens when two messages for the same task arrive at once, and whether a cancel request actually stops a model call on another replica.
Those questions are answered by one class, com.google.adk.a2a.executor.AgentExecutor, and the answers are not all obvious. Everything below was read from the google/adk-java main branch on 2026-10-08. The module is young and changes quickly, so each behaviour comes with a test you can run against the version you pin, and you should trust that test over this page.
The execute lifecycle, step by step
The a2a-java SDK owns HTTP, JSON-RPC, the task store and the event queue. When a message arrives it builds a RequestContext (task ID, context ID, the message, the existing task if there is one, and the send parameters) and calls the executor. ADK's executor then does the following, in order:
- Wraps the context and queue in the SDK's
TaskUpdater. Ifctx.getTask()is null, this is a new task, and it callssubmit(). - Registers the task in a
ConcurrentHashMapcalledactiveTaskswithputIfAbsent. If the task ID is already present, it throwsIllegalStateException("Task ... already running"). - Converts the A2A message parts to ADK
Contentand calls yourbeforeExecuteCallback, if any. If that returnstrue, the executor cancels the task instead of running it. - Looks up or creates the ADK session (the next section explains how), marks the task
working, and callsrunner.runAsync(userId, sessionId, content, runConfig). - Converts each ADK event to a
TaskArtifactUpdateEventaccording to the output mode, and passes it throughafterEventCallbackbefore enqueueing it. - When the event stream ends, builds one final
TaskStatusUpdateEvent:COMPLETED, orFAILEDwith an opaqueerror_id. It passes that event throughafterExecuteCallback, enqueues the result, and removes the task fromactiveTasks.
Two defaults come from AgentExecutorConfig: the run config disables model streaming and caps a run at 20 LLM calls, and the output mode is ARTIFACT_PER_RUN. Everything else on this page follows from the steps above.
The context ID becomes the user
Look at how step 4 names the user. The executor builds the ADK user ID as "A2A_USER_" + ctx.getContextId(). It does not use the caller's identity, because the executor never sees one; authentication happens in the transport, outside this class. There are three consequences.
user:-prefixed state is per conversation, not per caller. If your agent stores a preference withuser:state, it is scoped to the A2A context. A new context from the same calling agent starts from nothing.- Memory search is per conversation too. A memory service keyed by user will not connect two contexts from the same tenant. If you want cross-conversation recall, key it on a tenant identifier you pass explicitly (see below).
app:state is shared by every caller. This is the one prefix that crosses contexts, so never write anything caller-specific to it from an A2A-exposed agent.
None of this is wrong; a context is the protocol's unit of conversation. It is just different from a local deployment, where you choose the user ID from your own login. Agents written for local use often depend on user: state meaning a real person. Before you expose them, audit every user: key they read.
Session continuity across turns
Now the session. The executor calls getSession(appName, userId, contextId, Optional.empty()). If that finds nothing, it calls createSession(appName, userId). The ADK interface documents that two-argument overload as letting the service generate a unique session ID. So, as of main on 2026-10-08, the first message in a context creates a session whose ID is not the context ID. The second message in the same context looks up the context ID, misses, and creates another fresh session. Unless your session service maps the IDs some other way, the agent sees each turn of a multi-turn A2A conversation as a new conversation.
You may not notice in a demo, because single-shot tasks work perfectly. It shows up when a remote agent asks a clarifying question, enters input-required, and the caller's answer arrives with no history. You can close the gap without forking ADK. The before-execute hook runs before the lookup, so it can create the session under the context ID:
static final String USER_PREFIX = "A2A_USER_"; // mirrors the executor's private constant
static AgentExecutorConfig continuityConfig(BaseSessionService sessions, String appName) {
return AgentExecutorConfig.builder()
.beforeExecuteCallback(ctx -> {
String userId = USER_PREFIX + ctx.getContextId();
return sessions
.getSession(appName, userId, ctx.getContextId(), Optional.empty())
.switchIfEmpty(Maybe.defer(() -> sessions
.createSession(appName, userId, Map.of(), ctx.getContextId())
.toMaybe()))
.map(s -> false) // false = do not skip execution
.defaultIfEmpty(false);
})
.build();
}This depends on two internals: the prefix and the lookup order. Pin it with a test that fails loudly if either changes. Send two messages with the same context ID through the executor using an in-memory session service. Then assert that the second run's session contains the first run's user event, and that exactly one session exists for that user. If a future ADK release fixes the lookup itself, the test still passes and you can delete the callback.
Carrying the caller's identity inward
Because the executor knows nothing about the caller, carrying identity is your job. There are two steps. First, authenticate at the HTTP layer, before the SDK handler runs. Use the security scheme your card advertises, typically OAuth bearer tokens or mTLS between services, as described in Verifying Agent Identity in A2A. Do not use beforeExecuteCallback as the authentication gate: its skip path cancels the task rather than rejecting the request, so an unauthenticated caller sees a cancelled task instead of an authorization error.
Second, pass the verified facts inward. The executor copies the send request's metadata map into RunConfig.customMetadata() under the key a2a_metadata. That map comes from the caller, so treat it as untrusted. The safe pattern is for your HTTP filter to verify the token and then overwrite a server-owned key in that metadata, such as the tenant, before the handler sees it. A tool then reads it through its context:
@SuppressWarnings("unchecked")
static String tenantOf(ToolContext ctx) {
Map<String, Object> custom = ctx.invocationContext().runConfig().customMetadata();
Object a2a = custom.get("a2a_metadata");
if (!(a2a instanceof Map<?, ?> m) || !(m.get("x-verified-tenant") instanceof String t)) {
throw new IllegalStateException("no verified tenant on this invocation");
}
return t;
}Fail closed. A tool that falls back to a default tenant when the key is missing is simply waiting for a misconfigured filter to leak data across tenants.
Replicas, concurrent messages and cancel
activeTasks is an ordinary field on the executor, so it describes one JVM. This matters as soon as you run two replicas behind a load balancer.
| Situation | Same replica | Different replica |
|---|---|---|
| Second message for a task that is still running | Rejected: IllegalStateException | Runs concurrently against the same session |
| Cancel while the model is mid-call | updater.cancel(), then the subscription is disposed | Task marked cancelled (shared task store); the run on the other replica keeps going |
| Follow-up after the task finished | Normal | Normal, if the session and task stores are shared |
The disposal in the cancel path is what actually stops work: it unsubscribes from runAsync, which stops further model and tool calls. A cancel that lands on a different replica only changes the task's recorded state, while the original replica keeps spending tokens and may append events to a session that the caller believes is finished. Two designs fix this:
- Affinity by context. Route on a hash of the context ID (or the task ID) at the load balancer or a thin router, so every request for a conversation reaches the replica that runs it. This is simple and is the same idea as task ownership in A2A multi-region topology.
- A cancel channel. Publish cancels on a shared bus (Pub/Sub, Redis); every replica subscribes and calls
executor.cancelfor task IDs it holds. This survives rebalancing, but adds a component to operate.
Either way, use a durable, shared session service. With InMemorySessionService behind two replicas, each replica has half the conversations, and the continuity fix above makes things worse because it creates the session wherever the message happens to land.
The server's callback hooks
The three callbacks in AgentExecutorConfig are the server's policy layer. Their signatures, from Callbacks.java:
Single<Boolean> BeforeExecuteCallback.call(RequestContext ctx); // true = skip
Maybe<TaskArtifactUpdateEvent> AfterEventCallback.call(
RequestContext ctx, TaskArtifactUpdateEvent processed, Event adkEvent);
Maybe<TaskStatusUpdateEvent> AfterExecuteCallback.call(
RequestContext ctx, TaskStatusUpdateEvent finalEvent);Use afterEventCallback for redaction and audit: it sees both the converted artifact and the raw ADK event, so it can log tool calls that never leave the server while stripping internal fields from what does. Choose the output mode deliberately. ARTIFACT_PER_EVENT shows the caller intermediate model text and is useful for progress, but every intermediate thought becomes something the peer can store. ARTIFACT_PER_RUN releases only the final result.
Be careful with afterExecuteCallback. The executor enqueues whatever it returns. If it returns an empty Maybe, for example because a metrics call inside it failed and you swallowed the error, the final status event is never enqueued. The caller then waits on a task that never reaches a terminal state. Return the input event unchanged unless you mean to rewrite it, and test the empty case.
Worked example: a billing agent on two replicas
Here is a worked example: a billing agent exposed to an orchestrator owned by another team, running on two replicas.
- The orchestrator sends "Why was invoice 8812 higher than last month?" with a new context
ctx-41. The HTTP filter verifies the token, maps it to tenantacme, and writesx-verified-tenant=acmeinto the request metadata. - The router hashes
ctx-41to replica B. Execute submits the task, and the continuity callback creates sessionctx-41for userA2A_USER_ctx-41. - The agent calls
getInvoice. The tool reads the tenant froma2a_metadataand queries only acme's rows. The model asks which line item the user means, and the task ends ininput-required. - The orchestrator replies "the compute line" on the same context. It hashes to replica B again, the lookup finds session
ctx-41with the previous turn, and the agent answers with a comparison artifact. - Without the callback, step 4 starts from an empty session and asks the clarifying question again. Without affinity, step 4 can land on replica A, and if the session store is in memory, that is an empty session too.
Failure modes
| Symptom | Likely cause | Fix |
|---|---|---|
| Remote agent forgets the previous turn | Session created under a generated ID, not the context ID | Continuity callback plus two-turn test |
| "Task ... already running" in logs | Caller retried or sent a follow-up while the task was working | Caller waits for a terminal or input-required state; idempotent retries |
| Cancelled tasks keep billing tokens | Cancel served by a replica that is not running the task | Context affinity or a cancel bus |
| Tasks stuck in working forever | afterExecuteCallback returned empty, or the process died mid-run | Return the event; a sweeper that fails stale tasks |
| Preferences leak between tenants | Caller-specific data in app: state | Keep tenant data in tenant-keyed stores |
| Peer sees only an error_id | Working as designed: exception text is withheld | Look up the id in server logs; ADK_DEBUG_ERRORS is local-only |
The last row deserves emphasis. The executor deliberately sends the peer "Agent execution failed. (error_id: ...)" and logs the throwable under that ID. Setting ADK_DEBUG_ERRORS=1 puts exception text back into responses, so never set it on a reachable deployment. A2A error propagation covers how the calling side should act on these failures.
Trade-offs
Everything here trades simplicity against correctness under scale. One replica with an in-memory store has none of the replica problems, and is fine for an internal tool with a restart budget. Context affinity keeps the executor's in-process assumptions true, but makes a hot conversation a hot replica and complicates draining during deploys. A cancel bus lets requests balance freely, but cancellation becomes eventually consistent and you operate another moving part. Similarly, ARTIFACT_PER_RUN minimises what leaves the service, at the cost of progress visibility; streaming every event gives a better calling experience and a larger disclosure surface. Pick per agent, based on what its outputs contain.
What to do next
- Write the two-turn continuity test against your pinned ADK version before anything else, and add the before-execute callback if it fails.
- Grep your agent for
user:andapp:state keys and decide what each one should mean when the user is a context. - Put authentication in an HTTP filter, write a server-owned tenant key into the request metadata, and make tools fail closed when it is missing.
- Replace in-memory session and artifact services with shared durable ones before you add a second replica.
- Choose context affinity or a cancel bus, then test cancel from the replica that is not running the task.
- Audit your
afterExecuteCallbackfor paths that return empty, and add a sweeper that fails tasks stuck in working past your run timeout. - Set the output mode per agent based on what intermediate events would disclose.