When one ADK Java agent delegates to another over A2A, a failure no longer happens inside your process. The remote agent runs in a different service with its own session, its own logs and its own deployment schedule, and all you see is what the protocol carries back. A model call that throws in the remote service, a network timeout and a remote agent that simply refuses the task all look different on the calling side, and some do not look like errors at all.

This page traces every path a failure can take in each direction. It assumes the setup from ADK Java and A2A, which covers exposing an agent with AgentExecutor, consuming it with RemoteA2AAgent, the agent card, streaming and timeouts. In-process tool and model failures are covered in ADK Java error recovery. Behaviour described here was read from the current google/adk-java source and the javadoc on adk.dev; the A2A module is young and changes between releases, so verify each point against the version you depend on.

Advertisement

Why a boundary changes error handling

Inside one JVM, an exception carries its type, message and stack trace to whoever catches it. Across A2A, the remote side must turn the failure into protocol data: a task state, a status message, or a JSON-RPC error object. The caller then turns that data back into something Java can handle. Each conversion loses information, sometimes on purpose. A remote service should not send its stack traces, SQL fragments or internal hostnames to every client, and the A2A specification treats tasks as having a lifecycle rather than a return value.

So the questions change: did a task get created, what terminal state did it reach, what text came with it, and can the remote operator find the real cause from what you have?

The failure channels, at a glance

Calling service (orchestrator)Remote service (specialist)A2A boundary: HTTP + JSON-RPC, task states, messages1. AgentCardResolutionErrorthrown while you build RemoteA2AAgent2. A2AClientErrortransport, protocol or conversion failure; Flowable onError3. Event with errorMessageremote task FAILED; turnComplete = true; no exception4. Ordinary eventsCANCELED, REJECTED, INPUT_REQUIRED, AUTH_REQUIREDYour handlerfallback, retry, tell the user, record error_idAgent card endpoint/.well-known/agent-card.jsonAgentExecutor.executeruns the ADK agent through a Runnerrunner errorlogged with a 12-character error_idfinal TaskStatusUpdateEventFAILED + text: Agent execution failed. (error_id: ...)afterExecuteCallbackmay rewrite the final statuscard fetch failsFAILED statusHTTP / JSON-RPC error
Where failures appear on the calling side. Only channels 1 and 2 throw; a remote FAILED task arrives as an ordinary event carrying errorMessage, and other terminal or interrupted states carry no error marker at all.

There are four outcomes to design for on the calling side, and only two of them are exceptions.

  1. Card resolution failure. If the agent card cannot be fetched or parsed, RemoteA2AAgent.AgentCardResolutionError is thrown when the agent is constructed. This happens at startup, not on the first request.
  2. Client error. Connection refused, HTTP failures, JSON-RPC error responses and failures converting the remote response into ADK events all surface as A2AClientError (in com.google.adk.a2a.common) through the error signal of the agent's Flowable<Event>.
  3. Remote task FAILED. The remote agent ran, failed and reported a FAILED task. The ADK response converter turns this into a normal Event whose errorMessage() holds the status text and whose turnComplete() is true. Nothing is thrown.
  4. Other states. CANCELED, REJECTED, INPUT_REQUIRED and AUTH_REQUIRED end the remote turn too, but in the source examined they are not mapped to an error message. Your code must look at them explicitly if they matter.
Advertisement

What the server sends when the agent fails

On the remote side, ADK's AgentExecutor implements the a2a-java AgentExecutor interface: execute(RequestContext, EventQueue) runs the ADK agent through a Runner and streams results into the event queue. When the run ends with an error, the executor marks the task FAILED in a final status update. The design choice that matters most is what that status contains.

// What ADK's AgentExecutor does when the runner fails (current adk-java source, paraphrased):
//   state   = error != null ? FAILED : COMPLETED
//   errorId = 12 hex characters from a random UUID
//   log.error("Runner failed to execute [error_id={}]", errorId, error)   <- full stack trace, server only
//   message = AGENT message, text "Agent execution failed. (error_id: <id>)"
//             (+ ": <exception class>: <message>" only when ADK_DEBUG_ERRORS is 1 or true)
//   final TaskStatusUpdateEvent(state, message, isFinal = true)
//     -> passed through afterExecuteCallback if configured -> enqueued for the client

// An afterExecuteCallback that adds a stable, non-sensitive hint for callers.
Callbacks.AfterExecuteCallback addRetryHint = (ctx, statusEvent) -> {
  if (statusEvent.getStatus().state() != TaskState.FAILED) {
    return Maybe.just(statusEvent);
  }
  Message original = statusEvent.getStatus().message();
  String text = firstText(original) + " [retryable=false; contact=payments-oncall]";
  Message rewritten = new Message.Builder(original)
      .parts(List.of(new TextPart(text)))
      .build();
  return Maybe.just(new TaskStatusUpdateEvent.Builder(statusEvent)
      .status(new TaskStatus(TaskState.FAILED, rewritten, null))
      .build());
};
// Register it on AgentExecutorConfig. Check the builder method name and the copy-constructor
// builders above against the a2a-java and adk-java versions you actually depend on.

By default the peer receives only Agent execution failed. (error_id: 3f9c0a7b12de). The exception class, message and stack trace stay in the server log, keyed by the same 12-character id. Setting the environment variable ADK_DEBUG_ERRORS to 1 or true appends the exception class and message to the text. That helps on a laptop and leaks data in production, where exception messages contain SQL, file paths and user input. Keep it off outside development.

Two configuration hooks shape this path. A beforeExecuteCallback returns Single<Boolean>; if it emits true, the executor cancels the task instead of running it, so a caller sees CANCELED, not REJECTED or FAILED. Use it for cheap admission checks, and remember that callers will not see an error marker. An afterExecuteCallback receives the final TaskStatusUpdateEvent and may return a replacement, which lets you add a stable hint such as whether a retry makes sense, without exposing internals.

Protocol errors versus task failures

A FAILED task means the server accepted the request, created a task and the work failed. A protocol error means the request never became a successful task operation: the task id was not found, the operation or content type is not supported, the protocol version is wrong, or the parameters were invalid. The A2A specification names these errors, among them TaskNotFoundError, TaskNotCancelableError, UnsupportedOperationError, ContentTypeNotSupportedError and VersionNotSupportedError, and its JSON-RPC binding carries them as JSON-RPC error objects alongside the standard codes such as -32602 for invalid params and -32603 for internal errors. In the source examined, failures reported to the client's error handler surface as an A2AClientError.

The distinction drives retry decisions. The specification states that a task in a terminal state, including FAILED, cannot accept further messages; to try again, the client starts a new task. A transport error before any task id came back may mean the request was never received, or that it was received and the response was lost. The client cannot tell these apart, so retries of side-effecting work need an idempotency key.

Handling failures in the calling service

Start at construction. Because card resolution happens when you build the agent tree, a remote service that is down during your deployment can stop your whole application from starting. Decide whether that is what you want. For a non-critical specialist, catch the error and substitute a stub agent that explains the outage:

// Channel 1: the card is resolved while the agent tree is built, not on first use.
BaseAgent bookingAgent;
try {
  bookingAgent = remoteBookingAgent(bookingUrl);   // card resolution + RemoteA2AAgent.builder()
} catch (RemoteA2AAgent.AgentCardResolutionError e) {
  log.error("booking agent card unavailable at {}; starting without delegation", bookingUrl, e);
  bookingAgent = unavailableStub("booking_agent",
      "Booking is temporarily unavailable. Tell the user and offer to retry later.");
}

At run time, handle channels 2 and 3 in one place. Errors from the remote agent propagate through the parent agent's flowable to Runner.runAsync, possibly wrapped, so walk the cause chain rather than checking the top-level type. FAILED tasks must be caught in doOnNext, because they never reach the error path.

// Caller side: one handler for all three channels.
Flowable<Event> events = runner.runAsync(userId, sessionId, userMessage);

events
    .doOnNext(ev -> ev.errorMessage().ifPresent(msg -> {
        // Channel 3: remote task FAILED. Arrives as a normal event; nothing is thrown.
        String errorId = extractErrorId(msg);                   // "error_id: 3f9c0a7b12de"
        log.warn("remote agent {} failed, error_id={}", ev.author(), errorId);
        metrics.counter("a2a.remote_failed", "agent", ev.author()).increment();
    }))
    .onErrorResumeNext(t -> {
        // Channel 2: transport, protocol or conversion failure surfaces as onError.
        if (hasCause(t, A2AClientError.class)) {
            metrics.counter("a2a.client_error").increment();
            return Flowable.just(fallbackEvent(
                "The booking service is unavailable right now. Your request was not submitted."));
        }
        return Flowable.error(t);                               // a local bug: let it fail loudly
    })
    .blockingSubscribe(this::render);

static boolean hasCause(Throwable t, Class<? extends Throwable> type) {
  for (Throwable c = t; c != null; c = c.getCause()) {
    if (type.isInstance(c)) return true;
  }
  return false;
}

static String extractErrorId(String text) {
  Matcher m = Pattern.compile("error_id: ([0-9a-f]{12})").matcher(text);
  return m.find() ? m.group(1) : "unknown";
}

Note the fallback text. When a transport error interrupts a delegation, the honest message is that the request's outcome is unknown or was not confirmed, not that it failed. Say "was not submitted" only when you know the request never left the process.

Letting the orchestrating model see the failure

Handling an error in Java code is only half the job when the parent is an LlmAgent that decides what to say next. Whether the parent model sees a remote failure depends on how events are converted into the next model request in your ADK version. An event whose only failure signal is the errorMessage field may not appear in the model's context as text. Do not rely on it. A small wrapper agent turns every failure into explicit content with a fixed prefix the instruction can refer to:

// Sketch: a wrapper agent that turns every failure channel into explicit content the parent
// model can read. BaseAgent's constructor arguments differ between ADK releases; adapt them.
final class GuardedRemoteAgent extends BaseAgent {
  private final BaseAgent remote;

  @Override
  protected Flowable<Event> runAsyncImpl(InvocationContext ctx) {
    return remote.runAsync(ctx)
        .map(ev -> ev.errorMessage()
            .map(msg -> textEvent(ctx, "REMOTE_FAILED " + msg))   // channel 3 -> content
            .orElse(ev))
        .onErrorResumeNext(t -> hasCause(t, A2AClientError.class)
            ? Flowable.just(textEvent(ctx, "REMOTE_UNAVAILABLE: no task was confirmed"))
            : Flowable.error(t));
  }

  private Event textEvent(InvocationContext ctx, String text) {
    return Event.builder()
        .id(Event.generateEventId())
        .invocationId(ctx.invocationId())
        .author(name())
        .content(Content.fromParts(Part.fromText(text)))
        .build();
  }
}

Then tell the parent: "If a sub-agent replies with REMOTE_FAILED or REMOTE_UNAVAILABLE, do not claim the action succeeded; tell the user it could not be completed and offer to retry." The model now has a deterministic signal. The same pattern is a natural place to add a circuit breaker, as in the ADK Java circuit breaker guide, so a dead specialist fails fast instead of timing out on every turn. Agent-level ADK Java callbacks are an alternative hook for the same mapping.

Worked example: a payment failure, traced end to end

A travel orchestrator delegates "book flight 4821 for Ana" to a remote booking agent. In the booking service, the agent's payment tool throws a SQLTransientConnectionException because the connection pool is exhausted. The Runner's flowable errors, and the executor logs Runner failed to execute [error_id=3f9c0a7b12de] with the stack trace, then enqueues a final FAILED status whose text is Agent execution failed. (error_id: 3f9c0a7b12de).

In the orchestrator, RemoteA2AAgent receives the status update and the converter emits an event from booking_agent with errorMessage set to that text and turnComplete true. The handler logs the id and increments a2a.remote_failed; the wrapper turns it into a REMOTE_FAILED message; the parent model tells the user the booking could not be completed. The user asks again. Because FAILED is terminal, the retry is a new task, and because the booking tool used an idempotency key derived from the conversation and flight, the remote side can detect whether the first attempt reserved a seat before the payment step failed.

The on-call engineer copies error id 3f9c0a7b12de from the orchestrator log, searches the booking service logs and finds the pool exhaustion; the exception text never needed to cross the network.

Observability across the boundary

Log the same identifiers on both sides: the A2A task id, the context id, the error id for failures and your own trace id. Count each channel separately: client errors point to networking, certificates or version drift; FAILED tasks point to the remote agent; CANCELED may be your own admission callback. Propagate trace context over HTTP so a delegation shows up as one trace, as described in ADK Java observability.

Failure modes

SymptomCauseFix
Orchestrator reports success after a remote failureFAILED arrives as an event; only onError was handledCheck errorMessage in doOnNext; wrap the remote agent
Application fails to start when a specialist is downCard resolved at build timeCatch AgentCardResolutionError; substitute a stub agent
Caller sees CANCELED, nobody knows whybeforeExecuteCallback returned trueLog the reason server-side; count cancels separately
Stack traces or SQL in client logsADK_DEBUG_ERRORS enabled in productionRemove it from deployment config; rely on error_id
Retry rejected or ignoredNew message sent to a terminal taskStart a new task; use idempotency keys for side effects
Caller hangs until timeout after a remote failureafterExecuteCallback returned an empty Maybe, so no final status was sentAlways return a status event
Failures nobody can correlateerror_id not logged on the calling sideExtract and log it with task and context ids

Trade-offs

  • Detail versus leakage: the error_id design gives callers almost nothing to act on automatically. Adding structured, non-sensitive hints in an afterExecuteCallback restores some machine-readable information without exposing internals.
  • Wrapper agents versus callbacks: a wrapper is explicit and testable but depends on BaseAgent details that change between releases.
  • Automatic retries versus correctness: retrying transport errors is safe only for idempotent work; for everything else, surface the uncertainty to the user.

What to do next

  1. List every RemoteA2AAgent in your system and decide, for each, whether a card resolution failure should stop startup or degrade.
  2. Add a doOnNext check for errorMessage and an onErrorResumeNext for A2AClientError around your runner calls.
  3. Wrap remote agents so failures become explicit content, and update the parent instruction to refer to it.
  4. Confirm ADK_DEBUG_ERRORS is unset in every non-development environment.
  5. Log task id, context id and error_id on both sides, and build a runbook step that searches remote logs by error_id.
  6. Add idempotency keys to side-effecting remote tools before enabling any automatic retry.
Key takeaway: In ADK Java, a remote A2A failure reaches the caller through one of three channels: AgentCardResolutionError while the agent is built, A2AClientError on the Flowable for transport, protocol and conversion failures, or a normal Event with errorMessage set when the remote task ends FAILED, which throws nothing. The remote executor sends only a short error_id unless ADK_DEBUG_ERRORS is on, so log and correlate that id. Handle all three channels explicitly, turn failures into content the parent model can read, treat FAILED as terminal, and retry side-effecting work only with idempotency keys.