A model call in an ADK Java agent fails in boring ways: a rate limit, a 503 from an overloaded endpoint, a reset connection, a hung request. Retrying is the obvious response, and it is also how agents turn a short provider incident into a long one, because retries get added in several places by people who each assume they are the only ones doing it.

This article is about retries at the boundary around BaseLlm.generateContent. It builds on ADK Java error recovery, which shows a basic retrying decorator and the maxLlmCalls interaction. Here we cover the retry layer you probably already have without knowing it, how layers multiply, failure classification, deadlines, retry budgets, the streaming point of no return, and virtual-time tests.

API names were checked on 2026-10-04 against google/adk-java main, which pins google-genai 1.75.0, and the java-genai source at the v1.75.0 tag. Check your own versions; retry behaviour changes between releases.

Three layers that can each retry

Three places a model call can be retried, and how the attempts multiplyCallerreruns runAsync?ADK flowcounts maxLlmCallsRetry decoratoryour BaseLlm wrapperGemini (BaseLlm)GenAI SDK clientRetryInterceptoron by default: 5Model endpointHTTP 429 / 503layer 3: C attemptslayer 2: D attemptslayer 1: S attemptsWorst case per logical call = C x D x Sdefaults 1 x 1 x 5; a naive stack 2 x 3 x 5 = 30Retry in exactly one layer with a deadline and a budget; turn the others down.
A logical model call passes through up to three retry layers. Attempt counts multiply, so an outage is amplified by the product, not the sum.

Count the places a failed model request can be repeated. Layer 3 is the caller: code that catches the error from the stream returned by Runner.runAsync and runs the invocation again. Layer 2 is a decorator you write around the model, a BaseLlm subclass that delegates and retries. Layer 1 is the HTTP client inside the Google GenAI Java SDK, which ADK's Gemini class uses to talk to the Gemini API or Vertex AI.

The layers nest. If the caller tries C times, the decorator D times and the SDK S times, one logical call can send C x D x S requests to a provider that is already struggling, and because each layer backs off independently, the waits stack too.

The layers see different things. The SDK sees HTTP status codes but nothing about the agent. The caller sees the whole invocation, but rerunning runAsync replays the turn, and tools may run again. Only the decorator knows whether partial output has reached the user, and it can count attempts per model. Make the decorator the home of policy, keep the SDK layer small, and keep the caller out of retrying.

Layer 1: the retries already in the SDK

Layer 1 is the one people miss. In the java-genai v1.75.0 source, the client's HTTP setup always installs a RetryInterceptor, using the retry options from HttpOptions if you set any and a default HttpRetryOptions otherwise. The defaults in the interceptor are 5 attempts including the original, an initial delay of 1 second, a maximum delay of 60 seconds, an exponential base of 2 and a jitter factor of 1.0. It retries HTTP 408, 429, 500, 502, 503 and 504, and it also retries IOExceptions until the last attempt.

The delay before retry n is 1 second times 2 to the power n-1, scaled by a random factor between 0 and 2 (jitter 1.0) and capped at 60 seconds. The four waits average 1, 2, 4 and 8 seconds: about 15 seconds of expected backoff, up to about 30, before the SDK gives up. The source checked does not read a Retry-After header.

In the adk-java source checked, Gemini.Builder offers modelName, apiClient, apiKey and vertexCredentials, and nothing in Gemini changes the SDK's retry options. So an agent created with a plain model name such as gemini-2.5-flash already retries up to five times at the HTTP layer. To own the policy, pass a GenAI Client whose retry options you chose:

// Turn the SDK's transport retries down to one attempt ("0 or 1 means no retries")
// HttpOptions.timeout (ms) is a whole-call ceiling that includes a streamed body:
// size it above your longest answer and leave idle detection to the decorator.
HttpOptions http = HttpOptions.builder()
    .retryOptions(HttpRetryOptions.builder().attempts(1).build())
    .timeout(60_000)
    .build();

Client genai = Client.builder()
    .apiKey(System.getenv("GOOGLE_API_KEY"))
    .httpOptions(http)
    .build();

BaseLlm gemini = Gemini.builder()
    .modelName("gemini-2.5-flash")
    .apiClient(genai)
    .build();

BaseLlm model = new LayeredRetryLlm(gemini, RetryPolicy.interactive(), budget, Schedulers.computation());
LlmAgent agent = LlmAgent.builder().name("support_agent").model(model).build();

Keeping the SDK at two attempts, to absorb a single connection reset, is also reasonable. What matters is that you choose and write the number down.

Classifying failures

A retry policy is mostly a classifier. In java-genai, HTTP failures surface as ApiException with code(), status() and message(); 4xx responses raise the subclass ClientException and 5xx responses ServerException. Reactive operators wrap exceptions, so walk the cause chain before deciding.

enum Failure {
  RATE_LIMITED(true), OVERLOADED(true), TIMEOUT(true), NETWORK(true),
  BAD_REQUEST(false), AUTH(false), NOT_FOUND(false), UNKNOWN(false);

  final boolean retryable;
  Failure(boolean retryable) { this.retryable = retryable; }

  static Failure classify(Throwable t) {
    for (Throwable c = t; c != null; c = c.getCause()) {
      if (c instanceof ApiException) {
        int code = ((ApiException) c).code();
        if (code == 429) return RATE_LIMITED;
        if (code == 408 || code == 504) return TIMEOUT;
        if (code == 500 || code == 502 || code == 503) return OVERLOADED;
        if (code == 401 || code == 403) return AUTH;
        if (code == 404) return NOT_FOUND;
        return BAD_REQUEST;           // 400, 413, other 4xx: the request itself is wrong
      }
      if (c instanceof TimeoutException || c instanceof SocketTimeoutException) return TIMEOUT;
      if (c instanceof IOException) return NETWORK;
    }
    return UNKNOWN;                   // bugs stay loud: do not retry what you cannot name
  }
}

Never retry a request the provider rejected as invalid: a prompt over the context limit fails identically every time. Never retry authentication failures. And treat UNKNOWN as fatal, so a bug in your own request assembly fails fast instead of hiding behind three polite retries.

Deadlines and per-attempt timeouts

Users do not care how many times you tried; they care how long they waited. Give every logical model call a deadline, say 20 seconds for a chat turn or 10 minutes for a batch job, and refuse to start a backoff sleep that cannot finish inside it.

Each attempt also needs its own timeout, because a hung request holds a connection and makes no progress. In the adk-java source checked, when Gemini builds the client itself from a model name, API key or Vertex credentials, it hands the SDK an OkHttp client with zero connect, read and write timeouts, which OkHttp treats as no timeout. The SDK still adds its retry interceptor on top. So do not assume the transport will save you: set HttpOptions.timeout, which the SDK applies as an OkHttp call timeout covering the whole call, streamed body included, and also sends to the server as an X-Server-Timeout header. Size it above your longest expected response and let the decorator's per-item RxJava timeout catch stalls. For streaming, prefer an idle timeout between chunks over a total one.

A retry budget

Deadlines bound one call; a retry budget bounds the service. In an outage every in-flight call retries at once, and three retries per call quadruples load on a provider that is already failing. A token bucket fed by successes absorbs blips but caps retries in an outage: each success deposits a fraction of a token, each retry spends a whole one.

/** Retries are allowed only while recent successes have paid for them. */
final class RetryBudget {
  private final double max;          // e.g. 10 tokens: a burst allowance
  private final double perSuccess;   // e.g. 0.1: at most ~10% extra load in steady state
  private double tokens;

  RetryBudget(double max, double perSuccess) {
    this.max = max; this.perSuccess = perSuccess; this.tokens = max;
  }
  synchronized void onSuccess() { tokens = Math.min(max, tokens + perSuccess); }
  synchronized boolean tryAcquire() {
    if (tokens < 1.0) return false;
    tokens -= 1.0;
    return true;
  }
}

Share one budget per provider and model, not per agent: the quota belongs to the API key. When the budget is empty, the decorator fails fast and the error callback degrades gracefully.

The decorator, assembled

Putting the pieces together gives a decorator with four guards: nothing emitted yet, a retryable failure, attempts and deadline remaining, and budget available. The scheduler is injected so tests can control time.

public final class LayeredRetryLlm extends BaseLlm {
  private final BaseLlm delegate;
  private final RetryPolicy policy;     // maxAttempts, attemptTimeout, deadline, delayMillis(n)
  private final RetryBudget budget;
  private final Scheduler scheduler;

  public LayeredRetryLlm(BaseLlm delegate, RetryPolicy policy, RetryBudget budget, Scheduler scheduler) {
    super(delegate.model());
    this.delegate = delegate; this.policy = policy; this.budget = budget; this.scheduler = scheduler;
  }

  @Override
  public Flowable<LlmResponse> generateContent(LlmRequest request, boolean stream) {
    return Flowable.defer(() -> {
      AtomicBoolean emitted = new AtomicBoolean();
      AtomicInteger attempt = new AtomicInteger();
      long start = scheduler.now(TimeUnit.MILLISECONDS);
      return Flowable.defer(() -> {
            attempt.incrementAndGet();
            return delegate.generateContent(request, stream)
                // Idle timeout: max gap before the first item and between items.
                .timeout(policy.attemptTimeout().toMillis(), TimeUnit.MILLISECONDS, scheduler);
          })
          .doOnNext(r -> emitted.set(true))
          .doOnComplete(budget::onSuccess)
          .retryWhen(errors -> errors.flatMap(err -> {
            Failure f = Failure.classify(err);
            int n = attempt.get();
            long delay = policy.delayMillis(n);
            long elapsed = scheduler.now(TimeUnit.MILLISECONDS) - start;
            boolean retry = !emitted.get() && f.retryable
                && n < policy.maxAttempts()
                && elapsed + delay < policy.deadline().toMillis()
                && budget.tryAcquire();
            if (!retry) return Flowable.<Long>error(err);
            Metrics.retry(model(), f, n);
            return Flowable.timer(delay, TimeUnit.MILLISECONDS, scheduler);
          }));
    });
  }

  @Override
  public BaseLlmConnection connect(LlmRequest request) {
    return delegate.connect(request);   // live sessions need their own reconnect logic
  }
}

The outer defer creates fresh counters per subscription, so agents sharing one model instance cannot corrupt each other's counts. Give RATE_LIMITED a longer base delay than OVERLOADED: quota windows last seconds to minutes.

Streaming: the point of no return

With stream set to true, the Flowable emits partial LlmResponse objects that ADK forwards as partial events. Once one has gone out, a retry restarts the answer and the user sees the first sentence twice. So the decorator stops retrying once anything is emitted, and failure belongs to onModelErrorCallback, which can say plainly that the answer was cut off.

The SDK layer has the same limit for a different reason. An HTTP interceptor makes its decision when the response status arrives; a stream that breaks after the body has started arriving is, by inference from how interceptors work, beyond its reach. So the mid-stream failures that users actually notice are retried by nobody unless you design for it.

If mid-stream retries matter, either buffer the whole response before emitting it, giving up streaming latency, or restart visibly by telling the UI to discard the partial answer first. Do not ask the model to continue from the partial text.

Retries and maxLlmCalls

Retries inside a decorator are invisible to ADK: the flow counts one call per step, so three attempts count once against maxLlmCalls, which therefore does not cap provider requests. Export attempts per model and failure class, and alert on attempts per logical call; a ratio drifting from 1.02 to 1.4 warns of an incident before error rates move.

Testing with virtual time

Retry code is easy to get subtly wrong and impossible to test with real sleeps. Inject an RxJava TestScheduler and a fake model, as testing custom LLMs recommends for BaseLlm implementations generally, and advance time by hand.

final class FlakyLlm extends BaseLlm {
  private final int failures; private final AtomicInteger calls = new AtomicInteger();
  FlakyLlm(int failures) { super("fake-model"); this.failures = failures; }

  @Override public Flowable<LlmResponse> generateContent(LlmRequest req, boolean stream) {
    return Flowable.defer(() -> calls.incrementAndGet() <= failures
        ? Flowable.error(new ApiException(503, "UNAVAILABLE", "overloaded"))
        : Flowable.just(LlmResponse.builder()
            .content(Content.fromParts(Part.fromText("ok"))).build()));
  }
  @Override public BaseLlmConnection connect(LlmRequest req) { throw new UnsupportedOperationException(); }
  int calls() { return calls.get(); }
}

@Test void retriesTwiceThenSucceeds() {
  TestScheduler time = new TestScheduler();
  FlakyLlm flaky = new FlakyLlm(2);
  BaseLlm llm = new LayeredRetryLlm(flaky, RetryPolicy.interactive(), new RetryBudget(10, 0.1), time);

  TestSubscriber<LlmResponse> ts = llm.generateContent(LlmRequest.builder().build(), false).test();
  time.advanceTimeBy(30, TimeUnit.SECONDS);

  ts.assertComplete().assertValueCount(1);
  assertEquals(3, flaky.calls());
}

Add negative tests: a 400, an empty budget and a too-short deadline each mean one attempt, and a fake that emits a partial response then fails is never retried.

Worked example: a 503 storm

Consider a support agent serving 200 concurrent conversations, each turn making about four model calls. The provider has a ten-minute partial outage in which half of requests return 503. The team had added caller retries (2 attempts), a decorator copied from a blog (3 attempts) and never touched the SDK defaults (5 attempts).

Worst case, each logical call now sends 2 x 3 x 5 = 30 requests, and the extra load pushes the provider's error rate higher, feeding more retries. An unlucky user waits through about 15 seconds of SDK backoff per decorator attempt, three times, twice: the agent looks hung for over a minute, fails anyway, and the caller rerun repeats a refund tool call that had already succeeded.

The fix: SDK attempts 2, no caller retries, outages routed to onModelErrorCallback, and a decorator with 3 attempts under a 20-second deadline and a shared budget of 10 tokens at 0.1 per success. The worst case fell to 6 requests per logical call, the budget throttled retries within seconds, users got a clear message within 20 seconds, and no tool ran twice.

Failure modes

  • Retry storms. Unbudgeted retries at several layers multiply load during outages. Measure attempts per logical call.
  • Duplicated text. Retrying after partial output was streamed. Guard on an emitted flag.
  • Retrying the unfixable. 400, 401 and context-length errors retried until the attempt limit, wasting quota and time.

What to do next

  1. Check which google-genai version your ADK release pulls in and whether its HTTP client retries by default.
  2. Choose one layer to own policy, normally a BaseLlm decorator, and set the other layers to one or two attempts explicitly.
  3. Implement a classifier on ApiException.code() and the cause chain; treat unknown failures as fatal.
  4. Add a per-call deadline, a per-attempt idle timeout and a shared retry budget per provider and model.
  5. Stop retrying once partial output has been emitted; route those failures to onModelErrorCallback.
  6. Write TestScheduler tests for success-after-retries, non-retryable errors, empty budgets, deadlines and mid-stream failures.
  7. Read the BaseLlm interface overview and model call orchestration to see exactly where your decorator sits in a turn.
Key takeaway: An ADK Java model call can be retried by the GenAI SDK's HTTP interceptor, your BaseLlm decorator and the caller, and the attempts multiply. Let the decorator own policy: classify by status code, add a deadline and per-attempt timeout, share a retry budget, never retry after streamed output, and test it in virtual time.