In the Agent Development Kit for Java, an agent turn looks like one method call: you send a user message to a runner and receive a stream of events. Underneath, the runtime may call the model several times, run tools between calls, stream partial text, consult your callbacks, and stop because of a budget you may not know you have. When something goes wrong, such as a tool looping, a doubled response, a missing instruction or a quota error that surfaces as an agent failure, you need to know exactly what happens around each model call.

This article follows one model call through the runtime as it exists on the google/adk-java main branch in September 2026: how the request is assembled, where your code can intercept it, how the model is resolved and invoked, how streaming and errors are handled, and how the response turns into an event that decides whether the loop continues. The surrounding loop is covered in the execution loop anatomy and callbacks in general in ADK Java callbacks; here the focus is the call itself. The runtime's class and method names below were checked against that source; the code samples also use builder and SDK helpers, and details move between releases, so compile the samples against the version you depend on.

Advertisement

Where the model call lives

An LlmAgent does not call the model itself. It delegates to a flow, a subclass of BaseLlmFlow. When an agent disallows transfer to both its parent and its peers and has no sub-agents, it gets a SingleFlow; otherwise an AutoFlow, which adds the machinery for transferring control between agents. Both share the same core.

BaseLlmFlow.run(InvocationContext) returns a Flowable<Event> (the runtime is built on RxJava 3) and repeats runOneStep until an event is a final response, an action ends the invocation, a step limit is reached, or a step produces no events. Each step is: preprocess (request processors build an LlmRequest), callLlm (callbacks, budget, model invocation), and postprocess (response processors, conversion to events, and function-call handling). A step is one model call; an agent turn is one or more steps.

One step of BaseLlmFlow: from session events to a model call and backInvocationContextsession, agent, RunConfigRequest processorsinstructions, contents, toolsbefore-modelmay short-circuitLlmRequestLLM call budgetmaxLlmCallsResolve BaseLlminstance or LlmRegistrygenerateContentstream = (mode == SSE)on-model-errorreplacement or rethrowerrorafter-modelrewrite or passLlmResponseResponse processorse.g. code executionEventpartial / final, functionCallsFunction calls? run tools, append responses, next steploop ends on final response, endInvocation or step limitSession: events appended; next step rebuilds the requesthistory, not in-memory state, carries the conversation
One step of BaseLlmFlow. Request processors assemble the LlmRequest from the invocation context; the before-model callback may answer instead of the model; the call is counted against maxLlmCalls; the model is resolved and invoked; errors go to the on-model-error callback; responses pass the after-model callback and response processors and become events. Function calls in the event trigger tool execution and another step.

Step 1: assembling the LlmRequest

The request starts empty and a chain of RequestProcessors transforms it in order, each returning an updated request. In SingleFlow the chain sets the model name and generation config, attaches the output schema if one is configured, adds the agent's instruction (with state placeholders filled in), adds identity text, converts the session's event history into model contents, and handles code execution. The tools you registered contribute their own processors, which add function declarations to the request. The exact list changes between releases, so read SingleFlow in your version rather than relying on a list from an article.

Two consequences matter operationally. First, the request is rebuilt from the session on every step. The conversation lives in the session's events, not in any in-memory transcript, so whatever you put into state or events before a step is visible to the next call, and anything you only keep in a local variable is not. Second, the contents processor decides which events go into the prompt. The agent's includeContents setting can exclude prior history for agents that should see only their instruction and current input, and event authorship and branches determine what a sub-agent sees in a multi-agent tree. When the model seems to have forgotten something, dump the final LlmRequest first.

Advertisement

Step 2: the before-model callback

Before the model is called, the flow runs the agent's before-model callbacks. A BeforeModelCallback receives a CallbackContext and an LlmRequest.Builder and returns Maybe<LlmResponse>. Returning empty means continue, with whatever changes you made to the builder; returning a response means that response is used instead of calling the model. The synchronous variant, BeforeModelCallbackSync, returns Optional<LlmResponse>.

That makes the callback the right place for request-level guardrails, caching and cheap refusals:

LlmAgent agent = LlmAgent.builder()
    .name("order_support")
    .model(MODEL_NAME)
    .instruction("Answer questions about the customer's orders. Use lookup_order.")
    .tools(lookupOrderTool)
    .beforeModelCallbackSync((ctx, request) -> {
        Optional<String> cached = responseCache.get(cacheKey(request));
        if (cached.isPresent()) {
            return Optional.of(LlmResponse.builder()
                .content(Content.fromParts(Part.fromText(cached.get())))
                .build());                               // model is not called
        }
        return Optional.empty();                         // continue to the model
    })
    .afterModelCallbackSync((ctx, response) -> {
        auditLog.record(ctx.invocationId(), response.usageMetadata());
        return Optional.empty();                         // keep the response as is
    })
    .build();

Keep before-model callbacks fast and side-effect free, since they run on every step of every turn, and be careful with short-circuit responses that contain no function calls: they become a final response and end the turn. A cache keyed on the full request, including tool results already in the history, is safe; a cache keyed on the user's last message alone will return a stale answer in the middle of a tool-using turn.

Step 3: the call budget and model resolution

Every real call increments a per-invocation counter through incrementLlmCallsCount() on the invocation context. When the count exceeds RunConfig.maxLlmCalls(), the call fails with LlmCallsLimitExceededException and the flow emits an error. The default is 500; a negative value turns enforcement off, which the runtime logs as a warning. This budget is your last defence against a model that keeps calling tools forever, and 500 is a very generous ceiling for an interactive agent: at a few seconds and a few thousand tokens per call, a runaway turn can cost a great deal before it trips. Set it deliberately per use case, for example 10 to 30 for a support agent.

The model is then resolved. If the agent was built with a BaseLlm instance, that instance is used; if it was given a model name string, the flow asks LlmRegistry.getLlm(name), which maps name patterns to factories. Registering your own factory, or passing an instance, is how you route a name to a proxy, a different provider or a test double. BaseLlm has two abstract methods: generateContent(LlmRequest, boolean stream), returning Flowable<LlmResponse>, and connect(LlmRequest) for live bidirectional sessions. Writing one is covered in implementing a custom LLM.

Step 4: streaming, partial responses and aggregation

The flow passes stream = true to generateContent only when RunConfig.streamingMode() is SSE. With NONE the model returns one LlmResponse; with SSE it returns several, and the ones carrying incremental text are marked partial(). BIDI does not go through this path at all: it uses runLive and connect, with a live request queue for audio and text.

Each response becomes an event, so under SSE your client receives a sequence of partial events followed by a complete one. Partial events exist for display: render them, but treat the complete, non-partial event as the record of what the model said. The most common streaming bug in client code is concatenating every event's text, which prints the answer twice, once from the partials and once from the final aggregate. The second most common is acting on function calls from partial events; wait for the complete event.

Step 5: errors, the on-model-error callback and retries

If generateContent fails, the flow runs the agent's on-model-error callbacks. An OnModelErrorCallback receives the callback context, the request and the Exception, and returns Maybe<LlmResponse>. A response replaces the failed call and processing continues as if the model had produced it; empty means the original error propagates and ends the invocation's event stream with an error. This is the hook for graceful degradation, such as returning an apology on a content-filter failure, but it is a poor place for retries, because it cannot re-enter the model call with backoff cleanly.

The flow itself contains no retry or backoff logic. Whether a transient failure is retried depends on the BaseLlm implementation and the HTTP client under it, which vary by version and configuration. If you need a guaranteed policy, wrap the model:

public final class RetryingLlm extends BaseLlm {
    private static final int MAX_ATTEMPTS = 3;
    private final BaseLlm delegate;

    public RetryingLlm(BaseLlm delegate) {
        super(delegate.model());
        this.delegate = delegate;
    }

    @Override
    public Flowable<LlmResponse> generateContent(LlmRequest request, boolean stream) {
        return Flowable.defer(() -> {
            AtomicBoolean emitted = new AtomicBoolean(false);
            return Flowable.defer(() -> delegate.generateContent(request, stream))
                .doOnNext(r -> emitted.set(true))
                .retryWhen(errors -> errors
                    .zipWith(Flowable.range(1, MAX_ATTEMPTS), (err, attempt) -> {
                        // never retry once output reached the caller: partials would repeat
                        if (emitted.get() || !isTransient(err) || attempt == MAX_ATTEMPTS) {
                            throw Exceptions.propagate(err);
                        }
                        return attempt;
                    })
                    .flatMap(attempt -> Flowable.timer(
                        (250L << attempt) + ThreadLocalRandom.current().nextLong(250),
                        TimeUnit.MILLISECONDS)));
        });
    }

    @Override
    public BaseLlmConnection connect(LlmRequest request) {
        return delegate.connect(request);
    }
}

Here isTransient is yours to define: typically rate-limit and unavailable responses, and timeouts on idempotent calls. Pass new RetryingLlm(baseModel) to the agent builder's model(BaseLlm). Because retries happen below the flow, one logical call still counts once against maxLlmCalls, so budget retries separately. Retry policy at this layer is discussed further in LLM-layer retries.

Step 6: after-model callbacks, response processors and the loop decision

A successful response runs through the after-model callbacks, which receive the LlmResponse and may return a replacement: redact a secret, attach metadata, or record usage from usageMetadata(). Then the flow's ResponseProcessors run (in SingleFlow, code execution handling), and the response becomes an Event authored by the agent.

If the event contains function calls, the flow assigns client-side IDs to calls that lack them, identifies long-running tools, executes the tools (through before-tool and after-tool callbacks), and appends a function-response event. That event is not a final response, so the loop takes another step, and the next request contains the call and its result. If the event has no function calls and is not partial, it is the final response and the turn ends. An event can also end the invocation through its actions, for example after an agent transfer. The mechanics of tool execution are in tool dispatch mechanics.

Worked example: a two-call turn

A user asks the order_support agent, with streaming off, 'Where is order 812?'. The runtime appends the user message to the session and starts the flow.

StepWhat happensEvents emittedLLM calls
1Processors build the request: instruction, the user message, and the lookup_order declaration. The before-model callback misses the cache. The model returns a function call lookup_order(order_id=812).model event with functionCalls1
1 (tools)The tool runs and returns status 'shipped', carrier and ETA.function response event1
2Processors rebuild the request from the session, which now holds the user message, the call and its result. The model answers in text.final model event2

The turn made two model calls and emitted three events. Had the tool returned an error the model could not handle, the model might call it again on every step; with the default maxLlmCalls of 500 that loop could run for a long time, while a budget of 15 stops it within seconds and surfaces a clear exception you can alert on.

Failure modes and operational guidance

  • Runaway tool loops. Set maxLlmCalls per use case, alert on LlmCallsLimitExceededException, and return structured, actionable tool errors so the model can stop.
  • Duplicated text in clients. Render partial events, but store only complete ones.
  • Retrying after partial output. Retrying a stream that already emitted text duplicates it; the wrapper above refuses to.
  • Hidden context growth. The request grows with every step because history is replayed. Log input tokens per call and cap history for long sessions.
  • Blocking in callbacks. Synchronous callbacks run on the flow's thread; slow I/O there stalls the stream. Use the Maybe variants for I/O.
  • Assuming default retries. Test transient-error behaviour against your version; wrap the model if the policy matters.
  • Observability. Record per call: agent, step number, model, input and output tokens, latency, whether a callback short-circuited, and the error type.

What to do next

  1. Read BaseLlmFlow and SingleFlow in the ADK Java version you depend on and note its request processors.
  2. Add an after-model callback that logs usage per call, and a debug hook that dumps the final LlmRequest.
  3. Set RunConfig.maxLlmCalls for each agent to a deliberate number and alert when it trips.
  4. Decide your retry policy, verify the current behaviour with an injected 429, and add a BaseLlm wrapper if needed.
  5. If you stream with SSE, fix the client to render partial events and keep only complete ones.
  6. Write a test with a fake BaseLlm that returns a function call and then text, and assert the two-call trace above.
Key takeaway: Each step of an ADK Java agent is one model call: request processors rebuild the LlmRequest from the session, a before-model callback may answer instead, the call is counted against maxLlmCalls, the BaseLlm is resolved and invoked with streaming only in SSE mode, errors go to on-model-error callbacks, and responses pass after-model callbacks and processors to become events. Function calls trigger another step. Set the call budget deliberately, handle partial events correctly, and put retry policy in a BaseLlm wrapper rather than assuming it exists.