An ADK Java agent is a loop: the runner appends the user message to a session, the flow asks the model for the next step, the model either answers or requests tool calls, the tools run, and their results go back to the model. Every one of those hops can fail. A model endpoint returns a 429 or a 503, a tool throws because a downstream API timed out, a callback has a bug, or the loop runs past its call budget. What the user experiences depends entirely on where the failure happens and which recovery hook, if any, catches it.

This article is a map of that territory. It explains what the framework already does for you, which is more in some places and less in others than most people assume, then shows how to use the two error callbacks, how to add model retries without breaking the call budget, and how a caller should react when an error escapes the agent. Specific mechanisms each have their own page here: circuit breakers, tool timeouts and idempotency. Behaviour described below was checked against the google/adk-java main branch on 2026-10-01; the framework moves quickly, so re-read the source for your version before relying on a detail.

Advertisement

The two currencies of recovery

Every recovery decision in ADK Java comes down to one question: should this failure become data that the model reads, or an exception that the caller handles? These are different currencies and mixing them up causes most error-handling bugs.

When a tool failure becomes data, it is returned as the function response, a map such as {"status": "error", "reason": "rate_limited"}. The model sees it in the next turn and can apologise, try a different tool, ask the user for missing input, or retry with corrected arguments. The invocation continues. This is the right currency when the model can do something useful with the information.

When a failure becomes an exception, the reactive stream returned by Runner.runAsync, a Flowable<Event>, terminates with onError. The model never sees it. Events already emitted are already in the session. The caller must decide whether to retry, resume, degrade or report. This is the right currency when the model cannot help: the model provider is down, the budget is exhausted, or the agent itself is misconfigured.

One invocation, four places to fail, two ways to recoverCallerrunAsync(...)Runnersession + eventsLLM flow stepcount, then callmaxLlmCallslimit exceptionBaseLlmgenerateContentmodel callonModelErrorCallbackresponse or rethrowerrorfallbackTool dispatchfunction callsfunction callonToolErrorCallbackmap or rethrowerrorFunction responsemodel sees it, can adaptresult or error mapFlowable onErrorcaller decides: retry, resume, failRecovered errors become data the model reads.Unrecovered errors end the stream for the caller.Names follow google/adk-java main as checked on 2026-10-01; there is no built-in model retry
Model and tool errors pass through their error callbacks. A callback can turn the error into a response that the loop continues with; if it declines, the error terminates the stream returned to the caller.

Tool failures: what FunctionTool already does

Most tools are plain Java methods wrapped in FunctionTool. In the source checked, FunctionTool.runAsync wraps the method invocation in a try/catch. If your method throws synchronously, the exception is logged and the tool returns the map {"status": "error", "message": "An internal error occurred."}. The model gets a function response, the invocation continues, and no error callback fires.

That default is safe, because the agent does not crash, but it is uninformative. The model cannot tell a missing record from an expired credential from a downstream outage, so it either retries blindly or tells the user something vague. The catch also covers only the synchronous part. If your method returns a Single or Maybe that later fails, that error is not caught there; it travels to the tool error callback described next.

So return structured errors yourself for every failure the model can act on.

public static Map<String, Object> getOrder(
        @Schema(name = "orderId", description = "Order id, e.g. A-1042") String orderId) {
    try {
        Order o = orders.find(orderId);
        if (o == null) {
            return Map.of("status", "not_found",
                          "hint", "Ask the user to check the order id.");
        }
        return Map.of("status", "ok", "order", o.toMap());
    } catch (DownstreamTimeoutException e) {
        return Map.of("status", "error", "reason", "orders_service_timeout",
                      "retryable", true);
    } catch (AuthException e) {
        // Not something the model can fix; say so plainly.
        return Map.of("status", "error", "reason", "not_authorised",
                      "retryable", false);
    }
}

Keep error maps small and stable: a status, a machine-readable reason, whether a retry could help, and at most a short hint. Never put stack traces, SQL or internal host names in them; everything in a function response is model input and may be echoed to the user.

Advertisement

onToolErrorCallback: the last chance before an exception

Errors that do escape a tool, such as async failures or exceptions from tools that are not FunctionTool, go to the tool error callbacks. LlmAgent.Builder accepts onToolErrorCallback in an async form returning Maybe<Map<String, Object>> and a synchronous form, onToolErrorCallbackSync, returning Optional<Map<String, Object>>. Both receive the invocation context, the tool, its input arguments, the tool context and the exception.

In the flow code, plugin callbacks are tried first and then the agent's callbacks in order. The first one that returns a map wins and that map becomes the function response, exactly as if the tool had returned it. If every callback returns empty, the original exception is rethrown and the invocation stream fails. Returning empty is therefore a deliberate choice to escalate, not a neutral no-op.

LlmAgent agent = LlmAgent.builder()
    .name("support_agent")
    .model("gemini-2.5-flash")
    .tools(FunctionTool.create(OrderTools.class, "getOrder"))
    .onToolErrorCallbackSync((ctx, tool, args, toolCtx, error) -> {
        Throwable root = rootCause(error);
        log.warn("tool {} failed: {}", tool.name(), root.toString());
        if (root instanceof java.util.concurrent.TimeoutException
                || root instanceof java.net.SocketTimeoutException) {
            return Optional.of(Map.of("status", "error",
                "reason", "timeout", "retryable", true));
        }
        if (root instanceof IllegalArgumentException) {
            return Optional.of(Map.of("status", "error",
                "reason", "bad_arguments", "detail", root.getMessage()));
        }
        return Optional.empty();   // unknown: escalate to the caller
    })
    .build();

Unwrap the cause chain, because reactive operators wrap exceptions in RuntimeException. And convert only errors the model can act on; turning an authorisation failure into data invites the model to retry in a loop.

Also note: in the source checked, a function call naming an unregistered tool is logged as a warning and skipped, not raised. Alert on that warning; it usually means a renamed tool.

Model failures: there is no built-in retry

Model calls fail more often than tools: quota errors, overloaded endpoints, network resets. As of this check, ADK Java has no built-in retry for model calls. Issue 1397 on google/adk-java proposes an opt-in retry policy at the model-call boundary and is still open. Retrying the whole runAsync call is the wrong fix, because the session already holds the events from earlier steps, and a rerun can repeat tool calls that are not idempotent.

The right place to retry is around the single model call. LlmAgent.Builder.model accepts either a model name or a BaseLlm instance, so you can wrap the real model in a delegating BaseLlm that retries transient failures.

public final class RetryingLlm extends BaseLlm {
    private final BaseLlm delegate;
    private final int maxAttempts;

    public RetryingLlm(BaseLlm delegate, int maxAttempts) {
        super(delegate.model());
        this.delegate = delegate;
        this.maxAttempts = maxAttempts;
    }

    @Override
    public Flowable<LlmResponse> generateContent(LlmRequest request, boolean stream) {
        AtomicBoolean emitted = new AtomicBoolean(false);
        return Flowable.defer(() -> delegate.generateContent(request, stream))
            .doOnNext(r -> emitted.set(true))
            .retryWhen(errors -> errors.zipWith(
                Flowable.range(1, maxAttempts), (err, attempt) -> {
                    // Never retry after output has reached the caller,
                    // and never retry non-transient errors.
                    if (emitted.get() || !isTransient(err) || attempt == maxAttempts) {
                        throw Exceptions.propagate(err);
                    }
                    return attempt;
                })
                .flatMap(attempt -> Flowable.timer(backoffMillis(attempt),
                                                   TimeUnit.MILLISECONDS)));
    }

    @Override
    public BaseLlmConnection connect(LlmRequest request) {
        return delegate.connect(request);   // live sessions are not retried here
    }

    private static long backoffMillis(int attempt) {
        long base = 500L << (attempt - 1);              // 500, 1000, 2000 ...
        return base / 2 + ThreadLocalRandom.current().nextLong(base / 2 + 1);
    }
}

// Wiring: wrap the registered model rather than constructing it yourself.
BaseLlm real = LlmRegistry.getLlm("gemini-2.5-flash");
LlmAgent agent = LlmAgent.builder().name("support_agent")
    .model(new RetryingLlm(real, 3)).build();

The isTransient classifier is provider-specific: rate limiting, overload and connection resets are transient; invalid requests, safety blocks and authentication failures are not, and retrying them only burns quota. The emitted flag matters when streaming: once partial responses have reached the caller, retrying would duplicate text, so the wrapper gives up and lets the error callback or the caller handle it.

maxLlmCalls and what your retries do to it

RunConfig.builder().maxLlmCalls(n) sets a per-invocation ceiling on model calls. In the flow source, the invocation context increments its counter once per step, before calling the model, and throws LlmCallsLimitExceededException (package com.google.adk.models) when the limit is passed. That exception is your protection against a model that keeps calling tools forever.

Notice what that means for the wrapper above: three attempts inside one generateContent call count as one step. The budget then limits reasoning steps, not provider requests. That is often what you want, but if you pay per request or the provider has a strict quota, add your own attempt counter to the wrapper and export it as a metric. The open retry proposal takes the opposite view and counts every provider attempt against the budget, so check which semantics your version implements before upgrading.

A healthy agent rarely hits the limit; when it does, look in the session for a tool that keeps returning an error map the model believes it can fix.

onModelErrorCallback: a fallback answer

onModelErrorCallback receives the callback context, the request and the exception; the sync form returns Optional<LlmResponse>. If it returns a response, the flow uses it in place of the failed call. If it returns empty, the exception is rethrown. Use it for graceful degradation: a short, honest message to the user, or a response from a smaller fallback model.

.onModelErrorCallbackSync((cbCtx, request, error) -> {
    metrics.counter("adk.model.error", "type", rootCause(error).getClass().getSimpleName())
           .increment();
    if (!isTransient(error)) {
        return Optional.empty();                 // bugs and auth errors escalate
    }
    return Optional.of(LlmResponse.builder()
        .content(Content.fromParts(Part.fromText(
            "I could not reach the assistant service just now. "
          + "Your request was not completed; please try again in a minute.")))
        .turnComplete(true)
        .build());
})

Be careful with what the fallback claims. If tools already ran in this invocation, a message saying nothing happened may be false. And do not return a response that only sets errorCode expecting the loop to stop: in the source checked, the loop ends on a final response or an end-invocation action, not on an error code alone.

Recovering at the caller

Whatever escapes the callbacks reaches the caller as a stream error. Three facts guide the response. Events emitted before the error are already appended to the session, including any tool calls and their results. Side effects those tools caused have happened. And the failed step left no final response, so the session ends with an unanswered turn.

A safe caller therefore does not blindly re-send the user message. It classifies the error, then either reports it, or retries with a guard that makes repeated tool effects harmless: idempotency keys derived from the invocation id and tool arguments, as described in the idempotency article. Work that cannot be completed is recorded for later handling, as in the dead-letter pattern. ADK Java also has an @Experimental runAsync overload that takes an invocation id and resumes an existing invocation instead of starting a new one; it requires resumability to be configured on the app. Because it is experimental, test it against your session service before you depend on it.

runner.runAsync(userId, sessionId, userMsg, runConfig)
    .doOnNext(event -> stream.send(render(event)))
    .doOnError(err -> {
        Throwable root = rootCause(err);
        if (root instanceof LlmCallsLimitExceededException) {
            alerts.loopSuspected(sessionId);
            stream.send("I got stuck on this request. A person will follow up.");
        } else if (isTransient(root)) {
            deadLetters.record(sessionId, userMsg, root);   // replayed later, idempotently
            stream.send("Temporary problem; your request is queued.");
        } else {
            stream.send("Something went wrong and the request was not completed.");
        }
    })
    .onErrorComplete()
    .blockingSubscribe();

Worked example: an order-support agent under a provider outage

A support agent uses getOrder and refundOrder tools, a RetryingLlm with three attempts, both error callbacks, and maxLlmCalls(12). At 14:02 the model provider starts returning overload errors for about 40 percent of requests.

Most model calls succeed on the first or second attempt; latency rises slightly. A few fail all three attempts, and for those the model error callback returns the fallback message and adk.model.error spikes on the dashboard. One user was mid-refund: refundOrder had already run when the next model call failed, so the fallback wrongly said nothing was completed. The team changes the callback to check the session for a completed refund in this invocation. Meanwhile getOrder returns the structured timeout map when the orders service slows; the model retries once, then explains the delay. No invocation hits the call budget, so the error maps are not causing loops.

Failure modes

FailureWhat happensDefence
Tool throws synchronouslyGeneric error map; model is vagueReturn structured error maps from the tool
Async tool failure, no callbackInvocation stream failsonToolErrorCallback for known causes
Callback converts everythingModel loops on unfixable errorsConvert only actionable errors; escalate the rest
Rerun of runAsync after errorDuplicate tool side effectsIdempotency keys; resume instead of rerun
Retry after streamed outputDuplicated partial textStop retrying once anything is emitted
Renamed toolCalls silently skipped with a warningAlert on the tool-not-found log line

Trade-offs

DecisionOption AOption B
Error currencyData to the model: graceful, risk of loopsException to caller: explicit, abrupt for users
Retry locationInside BaseLlm wrapper: precise, invisible to budgetCaller rerun: simple, duplicates side effects
Budget sizeLow maxLlmCalls: stops loops fast, may cut real workHigh: completes long tasks, costs more when stuck

For deeper reading on failures across branches of a fan-out, see parallel agents and failures.

What to do next

  1. List every tool and the failures it can produce; return a structured error map for each one the model can act on.
  2. Add onToolErrorCallbackSync that unwraps the cause chain, converts only actionable errors and returns empty for the rest.
  3. Wrap your model in a retrying BaseLlm with jittered backoff, a transient-error classifier and a stop-after-emission rule.
  4. Set maxLlmCalls explicitly in RunConfig and alert on LlmCallsLimitExceededException.
  5. Add onModelErrorCallbackSync with an honest fallback that checks which tools already ran.
  6. Make write tools idempotent and route caller-side failures to a dead-letter store instead of rerunning.
Key takeaway: ADK Java gives you two recovery currencies: errors that become data the model reads, and errors that end the stream for the caller. FunctionTool already turns synchronous exceptions into a generic error map, so return specific maps yourself. Use onToolErrorCallback and onModelErrorCallback to convert only the failures the model or user can act on, retry model calls inside a BaseLlm wrapper rather than rerunning the invocation, remember how those retries interact with maxLlmCalls, and make the caller resume or dead-letter rather than repeat side effects.