An ADK Java agent calls tools that talk to the outside world: payment APIs, search back ends, databases, other agents. Some of those calls fail for reasons that disappear if you wait a moment, such as a 503 from a load balancer that is shedding load, a 429 from a rate limiter, or a connection reset during a deploy. Retrying those calls with backoff turns a visible failure into a slightly slower success. Retrying the wrong calls, or retrying them in the wrong place, does the opposite: it charges a card twice, or multiplies a small outage into a retry storm that keeps the dependency down.

This article builds tool-level retry for ADK Java from first principles: failure classification, backoff arithmetic, a RetryingTool decorator on the real BaseTool API, idempotency keys, budgets and a worked example. One correction to older notes first: ADK Java has no RetryPolicy argument when you register a tool and no ToolErrorKind enum. Retry is something you add by wrapping the tool, as shown below. The API surface used here was checked against the google/adk-java main branch on 2026-10-05.

Where retries happen in an invocation

A single user turn can be retried at four layers, and the attempt counts multiply. Your HTTP client may retry a request on its own (many do so by default for idempotent methods). A tool decorator can call the tool again. The model can decide to call the tool again on its next turn after it reads an error, until the RunConfig maxLlmCalls limit stops the invocation. Finally, the caller that drives Runner.runAsync may replay the whole turn. Three attempts at each of three layers is 27 requests against a dependency that was already struggling.

Four places a tool call can be retried; their attempt counts multiplycaller / runnerreplays the user turnLlmAgent loopmodel re-issues the callRetryingToolbackoff + jitterHTTP clientits own retriesdependencyrate limit, 503attemptsclassifierexception or result mapretry budgetdeadline, token bucketretry?function responseenvelope to modelgive upidempotency keyfunctionCallId, or invocationId + args hashWorst case per user turn = caller x model x tool x client attempts, e.g. 2 x 3 x 3 x 3 = 54 hits.Pick one layer to own backoff for each dependency; make every other layer retry at most once, or not at all.
Figure: retry layers around one tool call. The decorator is the right owner for backoff because it sees both exceptions and error maps, knows the deadline and can attach an idempotency key; the other layers should be capped.

The fix is ownership: for each dependency one layer applies backoff and the others make one attempt. The tool decorator is usually the best owner, because it sees the full failure and can hand the model a clear "do not retry" envelope. Leave model re-calls for errors the model can fix, such as a bad argument. For how the model reads those envelopes, see Wrapping Errors from Tools Back to the LLM; for retrying the model call itself, see ADK Java Error Recovery.

Classify before you retry

Retry is only correct when two things hold: the failure is likely to be transient, and running the operation again is safe. The first is about the error, the second about the tool. Classify before you retry, and keep the classifier in one place so every tool agrees.

FailureTransient?Retry in the decorator?
Connection refused or reset, DNS failureUsuallyYes, with backoff
HTTP 503, 502, 504UsuallyYes; honour Retry-After if sent
HTTP 429 rate limitedYes, after the windowYes, but wait at least Retry-After
Per-attempt timeoutMaybeOnly if the tool is idempotent or keyed
HTTP 400, 422, validation errorNoNo; return fix_arguments to the model
HTTP 401, 403NoNo; refreshing a token is a separate path
HTTP 404, business rule rejectionNoNo
NullPointerException, ClassCastExceptionNo, it is a bugNever

Timeouts deserve special care. A timeout tells you that you stopped waiting, not that the operation failed. The payment may have gone through and only the response was lost. That is why the next sections tie retry of timeouts to idempotency. For a broader catalogue of failure shapes, LLM Error Taxonomy in ADK Java classifies thrown versus in-band failures on the model side.

Backoff and jitter

Backoff spaces attempts out so that a recovering dependency is not hit again immediately. Exponential backoff doubles the ceiling each attempt: with a base of 200 ms the ceilings for the gaps before attempts 2, 3 and 4 are 200, 400 and 800 ms, capped at some maximum such as 2 s. Without jitter, every client that failed at the same moment retries at the same moment, so the load arrives in synchronized waves and the dependency fails again. Jitter randomizes the delay. The usual choice is full jitter: pick the delay uniformly between zero and the ceiling.

// Full jitter: uniform in [0, min(cap, base * 2^(attempt-1))]
static long backoffMillis(int attempt, long baseMs, long capMs) {
    long ceiling = Math.min(capMs, baseMs << Math.min(attempt - 1, 20));
    return ThreadLocalRandom.current().nextLong(ceiling + 1);
}

Full jitter halves the mean delay and spreads retries evenly over the window. If a server sends Retry-After, it knows more than your formula does: wait at least that long, and add a little jitter on top so that clients told the same value do not return together.

A RetryingTool decorator

In ADK Java a tool is a BaseTool. Its constructor takes a name, a description and a long-running flag, and the work happens in Single<Map<String, Object>> runAsync(Map<String, Object> args, ToolContext ctx). A decorator that extends BaseTool, forwards declaration() and processLlmRequest to the inner tool, and wraps runAsync can add retry to any tool, including function tools, MCP tools and OpenAPI tools you did not write. The Tool interface deep dive covers that contract in detail.

public final class RetryingTool extends BaseTool {
    private final BaseTool inner;
    private final RetrySpec spec;          // maxAttempts, base, cap, timeouts, classifier
    private final RetryBudget budget;      // shared per dependency, see below

    public RetryingTool(BaseTool inner, RetrySpec spec, RetryBudget budget) {
        super(inner.name(), inner.description(), inner.longRunning());
        this.inner = inner; this.spec = spec; this.budget = budget;
    }

    @Override public Optional<FunctionDeclaration> declaration() { return inner.declaration(); }

    @Override
    public Completable processLlmRequest(LlmRequest.Builder b, ToolContext ctx) {
        return inner.processLlmRequest(b, ctx);
    }

    @Override
    public Single<Map<String, Object>> runAsync(Map<String, Object> args, ToolContext ctx) {
        long deadline = System.nanoTime() + spec.totalBudget().toNanos();
        AtomicInteger attempt = new AtomicInteger();
        return Single.defer(() -> {                       // re-runs the tool on every subscribe
                attempt.incrementAndGet();
                return inner.runAsync(args, ctx)
                    .timeout(spec.perAttempt().toMillis(), TimeUnit.MILLISECONDS);
            })
            .flatMap(r -> spec.isRetryableResult(r)          // errors that arrive as maps
                ? Single.<Map<String, Object>>error(new RetryableResult(r))
                : Single.just(r))
            .doOnSuccess(r -> budget.recordCall())
            .retryWhen(errors -> errors.flatMap(err -> {
                int n = attempt.get();
                Throwable root = rootCause(err);
                long delay = spec.delayMillis(n, root);       // jitter, Retry-After aware
                boolean inTime = System.nanoTime() + delay * 1_000_000L < deadline;
                boolean retryable = err instanceof RetryableResult || spec.isRetryable(root);
                if (n >= spec.maxAttempts() || !retryable
                        || !inTime || !budget.tryWithdraw()) {
                    return Flowable.error(err);
                }
                return Flowable.timer(delay, TimeUnit.MILLISECONDS);
            }))
            .onErrorResumeNext(err -> err instanceof RetryableResult rr
                ? Single.just(rr.result)                      // hand the last map to the model
                : Single.error(err));                         // let the error callbacks see it
    }

    static final class RetryableResult extends RuntimeException {
        final Map<String, Object> result;
        RetryableResult(Map<String, Object> r) { super(null, null, false, false); result = r; }
    }
}

Three details matter. Single.defer is required: some tools do their work when runAsync is called rather than when the result is subscribed, so a plain retryWhen on an already-created Single could resubscribe without running anything again. The per-attempt timeout stops one hung call from eating the whole budget. And the decorator ends with either a map the model reads or an error for the tool error callbacks, so nothing is swallowed. Register it like any tool: LlmAgent.builder().tools(new RetryingTool(FunctionTool.create(Payments.class, "charge"), spec, budget)).

Errors that arrive as results

A plain exception-based retry misses an important case. When a method behind a FunctionTool throws synchronously, FunctionTool.runAsync catches the exception, logs it and completes successfully with {"status": "error", "message": "An internal error occurred."}. No error ever reaches retryWhen. Tools that follow an envelope convention also report outages as maps, for example {"status": "error", "error": {"code": "UNAVAILABLE", "retryable": true}}. That is why the decorator converts retryable result maps into a private exception, retries, and converts the last one back at the end.

boolean isRetryableResult(Map<String, Object> r) {
    if (!"error".equals(r.get("status"))) return false;
    Object e = r.get("error");
    if (e instanceof Map<?, ?> m) {
        Object code = m.get("code");
        return "UNAVAILABLE".equals(code) || "RATE_LIMITED".equals(code);
    }
    return false;   // the generic FunctionTool message hides the cause: do not guess
}

The generic message hides the cause, so retrying it would also retry bugs. Fix it upstream: have tool methods return a coded envelope for transient errors.

Idempotency keys for side effects

Reads can be retried freely. Writes with side effects (charge a card, send an email, create a ticket) need an idempotency key: a value the dependency stores with the first successful request so that a repeat with the same key returns the original result instead of acting again. Many payment and messaging APIs accept such a key in a request header. Your own services can implement it with a unique constraint on the key.

Where should the key come from? ToolContext.functionCallId() returns an Optional<String> with the id of the function call the model emitted. It is stable across the decorator's own retries, which is exactly the window you need. It is not stable across a model re-issue: if the model calls the tool again on its next turn, that is a new function call with a new id. To deduplicate those too, derive the key from invocationId() (inherited from ReadonlyContext), the tool name and a hash of the canonical arguments.

public static Map<String, Object> charge(
        @Schema(name = "orderId") String orderId,
        @Schema(name = "amountCents") long amountCents,
        ToolContext toolContext) {           // FunctionTool injects it by this name
    String key = toolContext.invocationId() + ":charge:" + sha256(orderId + "|" + amountCents);
    ChargeResult r = payments.charge(orderId, amountCents, key);   // key sent as a header
    return Map.of("status", "ok", "chargeId", r.id(), "replayed", r.replayed());
}

The invocation-scoped key folds two genuinely identical charges in one turn into one, which is the safe direction for payments. If a dependency has no idempotency mechanism, do not retry its writes after a timeout; return an "outcome unknown" envelope and have the model check state with a read tool.

Deadlines and retry budgets

Two limits keep retries from making an outage worse. The first is a deadline: give each tool call a total budget (say 8 s) and never schedule an attempt whose backoff ends past it. The second is a retry budget shared by every session calling the same dependency, as a token bucket: each success deposits a fraction of a token, each retry withdraws one. In a broad outage the bucket drains and retries stop, so load stays near normal.

public final class RetryBudget {
    private final double ratio, max;          // e.g. 0.1 retries per call, burst of 20
    private double tokens;

    public RetryBudget(double ratio, double max) { this.ratio = ratio; this.max = max; tokens = max; }
    public synchronized void recordCall()  { tokens = Math.min(max, tokens + ratio); }
    public synchronized boolean tryWithdraw() {
        if (tokens < 1.0) return false;
        tokens -= 1.0;
        return true;
    }
}

A retry budget limits extra load; a circuit breaker stops calls entirely once a dependency is clearly down. They work well together, with the breaker outside the retry decorator so an open breaker fails fast without consuming attempts. See ADK Java circuit breaker architecture.

Worked example: a payment tool during a deploy

Take a charge tool in front of a payment API that starts returning 503 for 30 seconds during a deploy. The spec is: 4 attempts, base 200 ms, cap 2 s, full jitter, per-attempt timeout 3 s, total budget 8 s, retry budget of 0.1 per call with a burst of 20. Two hundred agent sessions are active.

The first sessions to fail wait a random 0 to 200, 0 to 400 and 0 to 800 ms, about 700 ms on average, spread out rather than landing together. After 20 retries the shared bucket is empty and later sessions fail after one attempt, so the dependency sees near-normal load, not four times normal. Failing sessions return UNAVAILABLE with action: retry_later and the model tells the user the payment has not gone through. A request that timed out at 3 s but actually succeeded is replayed safely thanks to its key: the API returns the original charge with replayed=true.

Watch it in your metrics. Count attempts per call, record the final outcome, and alert when retry budget exhaustion happens. Tool Observability + Metrics in ADK Java shows the built-in tool span and how to add outcome counters.

Failure modes

  • Retry storms. No jitter, or no shared budget, so every session hammers the dependency in step. Symptom: request rate jumps to a multiple of normal exactly when errors start.
  • Duplicate side effects. A write retried after a timeout without an idempotency key. Symptom: double charges or duplicate emails that line up with latency spikes.
  • Stacked layers. HTTP client, decorator and model all retry. Symptom: one user turn produces dozens of upstream requests.
  • Retrying bugs. Treating every exception as transient.

Trade-offs

More attempts improve success during short blips and hurt during long outages. Longer backoff protects the dependency and costs user latency. Retrying in the decorator is fast and precise but hides failures from the model; leaving retries to the model makes them visible and explainable but costs a model call each time, which is far slower and more expensive than a local timer. Resilience4j offers retry and circuit breaker primitives with metrics built in; a small decorator is easier to read. Either way the policy decisions stay yours.

What to do next

  1. List every tool and mark it read, idempotent write or non-idempotent write.
  2. For each dependency, pick the one layer that owns backoff, and turn off retries in the HTTP client or the decorator accordingly.
  3. Wrap remote tools in a RetryingTool with full jitter, a cap, a per-attempt timeout and a total deadline.
  4. Make the classifier inspect both exceptions and result maps, and never retry the generic internal error.
  5. Add idempotency keys to every write tool, or forbid retries after timeouts for that tool.
  6. Share a retry budget per dependency and put a circuit breaker outside the decorator.
  7. Emit attempts-per-call and budget-exhausted metrics, then test with a fault injector that returns 503s for 30 seconds.
Key takeaway: Retry a tool call only when the failure is transient and repeating it is safe. Own backoff in one layer, usually a BaseTool decorator, use exponential backoff with full jitter under a total deadline, retry error maps as well as exceptions, key every write, and share a retry budget so an outage does not turn into a storm.