In the Agent Development Kit for Java, an agent turn looks like one method call: you send a user message to a runner and receive a stream of events. Underneath, the runtime may call the model several times, run tools between calls, stream partial text, consult your callbacks, and stop because of a budget you may not know you have. When something goes wrong, such as a tool looping, a doubled response, a missing instruction or a quota error that surfaces as an agent failure, you need to know exactly what happens around each model call.
This article follows one model call through the runtime as it exists on the google/adk-java main branch in September 2026: how the request is assembled, where your code can intercept it, how the model is resolved and invoked, how streaming and errors are handled, and how the response turns into an event that decides whether the loop continues. The surrounding loop is covered in the execution loop anatomy and callbacks in general in ADK Java callbacks; here the focus is the call itself. The runtime's class and method names below were checked against that source; the code samples also use builder and SDK helpers, and details move between releases, so compile the samples against the version you depend on.
Where the model call lives
An LlmAgent does not call the model itself. It delegates to a flow, a subclass of BaseLlmFlow. When an agent disallows transfer to both its parent and its peers and has no sub-agents, it gets a SingleFlow; otherwise an AutoFlow, which adds the machinery for transferring control between agents. Both share the same core.
BaseLlmFlow.run(InvocationContext) returns a Flowable<Event> (the runtime is built on RxJava 3) and repeats runOneStep until an event is a final response, an action ends the invocation, a step limit is reached, or a step produces no events. Each step is: preprocess (request processors build an LlmRequest), callLlm (callbacks, budget, model invocation), and postprocess (response processors, conversion to events, and function-call handling). A step is one model call; an agent turn is one or more steps.
Step 1: assembling the LlmRequest
The request starts empty and a chain of RequestProcessors transforms it in order, each returning an updated request. In SingleFlow the chain sets the model name and generation config, attaches the output schema if one is configured, adds the agent's instruction (with state placeholders filled in), adds identity text, converts the session's event history into model contents, and handles code execution. The tools you registered contribute their own processors, which add function declarations to the request. The exact list changes between releases, so read SingleFlow in your version rather than relying on a list from an article.
Two consequences matter operationally. First, the request is rebuilt from the session on every step. The conversation lives in the session's events, not in any in-memory transcript, so whatever you put into state or events before a step is visible to the next call, and anything you only keep in a local variable is not. Second, the contents processor decides which events go into the prompt. The agent's includeContents setting can exclude prior history for agents that should see only their instruction and current input, and event authorship and branches determine what a sub-agent sees in a multi-agent tree. When the model seems to have forgotten something, dump the final LlmRequest first.
Step 2: the before-model callback
Before the model is called, the flow runs the agent's before-model callbacks. A BeforeModelCallback receives a CallbackContext and an LlmRequest.Builder and returns Maybe<LlmResponse>. Returning empty means continue, with whatever changes you made to the builder; returning a response means that response is used instead of calling the model. The synchronous variant, BeforeModelCallbackSync, returns Optional<LlmResponse>.
That makes the callback the right place for request-level guardrails, caching and cheap refusals:
LlmAgent agent = LlmAgent.builder()
.name("order_support")
.model(MODEL_NAME)
.instruction("Answer questions about the customer's orders. Use lookup_order.")
.tools(lookupOrderTool)
.beforeModelCallbackSync((ctx, request) -> {
Optional<String> cached = responseCache.get(cacheKey(request));
if (cached.isPresent()) {
return Optional.of(LlmResponse.builder()
.content(Content.fromParts(Part.fromText(cached.get())))
.build()); // model is not called
}
return Optional.empty(); // continue to the model
})
.afterModelCallbackSync((ctx, response) -> {
auditLog.record(ctx.invocationId(), response.usageMetadata());
return Optional.empty(); // keep the response as is
})
.build();Keep before-model callbacks fast and side-effect free, since they run on every step of every turn, and be careful with short-circuit responses that contain no function calls: they become a final response and end the turn. A cache keyed on the full request, including tool results already in the history, is safe; a cache keyed on the user's last message alone will return a stale answer in the middle of a tool-using turn.
Step 3: the call budget and model resolution
Every real call increments a per-invocation counter through incrementLlmCallsCount() on the invocation context. When the count exceeds RunConfig.maxLlmCalls(), the call fails with LlmCallsLimitExceededException and the flow emits an error. The default is 500; a negative value turns enforcement off, which the runtime logs as a warning. This budget is your last defence against a model that keeps calling tools forever, and 500 is a very generous ceiling for an interactive agent: at a few seconds and a few thousand tokens per call, a runaway turn can cost a great deal before it trips. Set it deliberately per use case, for example 10 to 30 for a support agent.
The model is then resolved. If the agent was built with a BaseLlm instance, that instance is used; if it was given a model name string, the flow asks LlmRegistry.getLlm(name), which maps name patterns to factories. Registering your own factory, or passing an instance, is how you route a name to a proxy, a different provider or a test double. BaseLlm has two abstract methods: generateContent(LlmRequest, boolean stream), returning Flowable<LlmResponse>, and connect(LlmRequest) for live bidirectional sessions. Writing one is covered in implementing a custom LLM.
Step 4: streaming, partial responses and aggregation
The flow passes stream = true to generateContent only when RunConfig.streamingMode() is SSE. With NONE the model returns one LlmResponse; with SSE it returns several, and the ones carrying incremental text are marked partial(). BIDI does not go through this path at all: it uses runLive and connect, with a live request queue for audio and text.
Each response becomes an event, so under SSE your client receives a sequence of partial events followed by a complete one. Partial events exist for display: render them, but treat the complete, non-partial event as the record of what the model said. The most common streaming bug in client code is concatenating every event's text, which prints the answer twice, once from the partials and once from the final aggregate. The second most common is acting on function calls from partial events; wait for the complete event.
Step 5: errors, the on-model-error callback and retries
If generateContent fails, the flow runs the agent's on-model-error callbacks. An OnModelErrorCallback receives the callback context, the request and the Exception, and returns Maybe<LlmResponse>. A response replaces the failed call and processing continues as if the model had produced it; empty means the original error propagates and ends the invocation's event stream with an error. This is the hook for graceful degradation, such as returning an apology on a content-filter failure, but it is a poor place for retries, because it cannot re-enter the model call with backoff cleanly.
The flow itself contains no retry or backoff logic. Whether a transient failure is retried depends on the BaseLlm implementation and the HTTP client under it, which vary by version and configuration. If you need a guaranteed policy, wrap the model:
public final class RetryingLlm extends BaseLlm {
private static final int MAX_ATTEMPTS = 3;
private final BaseLlm delegate;
public RetryingLlm(BaseLlm delegate) {
super(delegate.model());
this.delegate = delegate;
}
@Override
public Flowable<LlmResponse> generateContent(LlmRequest request, boolean stream) {
return Flowable.defer(() -> {
AtomicBoolean emitted = new AtomicBoolean(false);
return Flowable.defer(() -> delegate.generateContent(request, stream))
.doOnNext(r -> emitted.set(true))
.retryWhen(errors -> errors
.zipWith(Flowable.range(1, MAX_ATTEMPTS), (err, attempt) -> {
// never retry once output reached the caller: partials would repeat
if (emitted.get() || !isTransient(err) || attempt == MAX_ATTEMPTS) {
throw Exceptions.propagate(err);
}
return attempt;
})
.flatMap(attempt -> Flowable.timer(
(250L << attempt) + ThreadLocalRandom.current().nextLong(250),
TimeUnit.MILLISECONDS)));
});
}
@Override
public BaseLlmConnection connect(LlmRequest request) {
return delegate.connect(request);
}
}Here isTransient is yours to define: typically rate-limit and unavailable responses, and timeouts on idempotent calls. Pass new RetryingLlm(baseModel) to the agent builder's model(BaseLlm). Because retries happen below the flow, one logical call still counts once against maxLlmCalls, so budget retries separately. Retry policy at this layer is discussed further in LLM-layer retries.
Step 6: after-model callbacks, response processors and the loop decision
A successful response runs through the after-model callbacks, which receive the LlmResponse and may return a replacement: redact a secret, attach metadata, or record usage from usageMetadata(). Then the flow's ResponseProcessors run (in SingleFlow, code execution handling), and the response becomes an Event authored by the agent.
If the event contains function calls, the flow assigns client-side IDs to calls that lack them, identifies long-running tools, executes the tools (through before-tool and after-tool callbacks), and appends a function-response event. That event is not a final response, so the loop takes another step, and the next request contains the call and its result. If the event has no function calls and is not partial, it is the final response and the turn ends. An event can also end the invocation through its actions, for example after an agent transfer. The mechanics of tool execution are in tool dispatch mechanics.
Worked example: a two-call turn
A user asks the order_support agent, with streaming off, 'Where is order 812?'. The runtime appends the user message to the session and starts the flow.
| Step | What happens | Events emitted | LLM calls |
|---|---|---|---|
| 1 | Processors build the request: instruction, the user message, and the lookup_order declaration. The before-model callback misses the cache. The model returns a function call lookup_order(order_id=812). | model event with functionCalls | 1 |
| 1 (tools) | The tool runs and returns status 'shipped', carrier and ETA. | function response event | 1 |
| 2 | Processors rebuild the request from the session, which now holds the user message, the call and its result. The model answers in text. | final model event | 2 |
The turn made two model calls and emitted three events. Had the tool returned an error the model could not handle, the model might call it again on every step; with the default maxLlmCalls of 500 that loop could run for a long time, while a budget of 15 stops it within seconds and surfaces a clear exception you can alert on.
Failure modes and operational guidance
- Runaway tool loops. Set
maxLlmCallsper use case, alert onLlmCallsLimitExceededException, and return structured, actionable tool errors so the model can stop. - Duplicated text in clients. Render partial events, but store only complete ones.
- Retrying after partial output. Retrying a stream that already emitted text duplicates it; the wrapper above refuses to.
- Hidden context growth. The request grows with every step because history is replayed. Log input tokens per call and cap history for long sessions.
- Blocking in callbacks. Synchronous callbacks run on the flow's thread; slow I/O there stalls the stream. Use the
Maybevariants for I/O. - Assuming default retries. Test transient-error behaviour against your version; wrap the model if the policy matters.
- Observability. Record per call: agent, step number, model, input and output tokens, latency, whether a callback short-circuited, and the error type.
What to do next
- Read BaseLlmFlow and SingleFlow in the ADK Java version you depend on and note its request processors.
- Add an after-model callback that logs usage per call, and a debug hook that dumps the final LlmRequest.
- Set RunConfig.maxLlmCalls for each agent to a deliberate number and alert when it trips.
- Decide your retry policy, verify the current behaviour with an injected 429, and add a BaseLlm wrapper if needed.
- If you stream with SSE, fix the client to render partial events and keep only complete ones.
- Write a test with a fake BaseLlm that returns a function call and then text, and assert the two-call trace above.