An ADK Java agent has three kinds of configuration, and most production incidents come from confusing them. Process configuration, such as API keys and model names, is read once when the service starts. Assembly configuration, the agent tree, the session and artifact services and the plugins, is fixed when the Runner is built. And runtime configuration, how one particular invocation should behave, is passed with every call as a RunConfig.
This article is about that third layer. It explains what each RunConfig field does and what it defaults to, how the LLM call budget stops runaway loops, what the streaming and tool execution modes really change, and how to build a RunConfig per request from environment defaults and tenant policy, so it is validated, testable and observable. The field list and semantics below were checked against the current adk-java source; some fields are recent, so confirm them against the version you depend on.
Three layers of configuration
| Layer | Examples | When it changes | Where it is covered |
|---|---|---|---|
| Process | API keys, model id, database URL, feature flags | Deploy or restart | environment and configuration management |
| Assembly | Agent tree, tools, session / artifact / memory services, plugins | Build of the Runner | runtime boot sequence |
| Invocation | LLM call budget, streaming, tool concurrency, session creation | Every call | This article |
The distinction matters because the layers have different lifetimes and different owners. An operator changes process configuration. A developer changes assembly. But invocation configuration can legitimately differ between two requests arriving in the same millisecond: a free-tier user gets a smaller LLM budget, a browser that accepts server-sent events gets streaming, a batch job does not. If those decisions leak into process configuration, you end up redeploying to change a per-customer limit. If they are hard-coded at the call site, nobody can see or change them.
Where RunConfig enters the runtime
The Runner exposes runAsync overloads that take a user id, a session id (or a SessionKey), the new message and a RunConfig, optionally with an initial state delta. For live, bidirectional sessions there is runLive, which takes a LiveRequestQueue instead of a single message. Overloads that omit the config use RunConfig.builder().build(), which means every default described below applies silently.
Inside, the Runner creates an InvocationContext holding the config for the duration of that invocation. The LLM flow reads the streaming mode, the context counts model calls against the budget, and the function-calling code reads the tool execution mode. The config is an immutable value object built with a builder, so it is safe to share across threads and trivial to log.
The fields and their defaults
| Field | Default | What it controls |
|---|---|---|
maxLlmCalls | 500 | Maximum model calls in one invocation; 0 or negative disables the check |
streamingMode | NONE | NONE, SSE or BIDI |
toolExecutionMode | NONE | How several function calls from one model turn are executed |
autoCreateSession | false | Whether an unknown session id is created or rejected |
saveInputBlobsAsArtifacts | false | Whether inline binary parts of the user message are stored as artifacts |
responseModalities | empty | Requested output modalities, for example audio for live agents |
speechConfig, inputAudioTranscription, outputAudioTranscription | unset | Voice and transcription settings for live sessions |
customMetadata | empty map | Caller-supplied metadata carried with the config |
The source also contains fields for avatar settings and an override for how function responses are grouped in history; treat those as advanced and check the Javadoc of your version before relying on them. The two fields that matter most for reliability are the call budget and the tool execution mode, and both have defaults you should override deliberately.
import com.google.adk.agents.RunConfig;
import com.google.adk.agents.RunConfig.StreamingMode;
import com.google.adk.agents.RunConfig.ToolExecutionMode;
RunConfig cfg = RunConfig.builder()
.maxLlmCalls(25) // default 500; 0 or negative disables the check
.streamingMode(StreamingMode.SSE) // default NONE
.toolExecutionMode(ToolExecutionMode.PARALLEL) // default NONE, which behaves like PARALLEL
.autoCreateSession(false) // default false: unknown session is an error
.saveInputBlobsAsArtifacts(true) // default false
.build();
runner.runAsync(userId, sessionId, userMessage, cfg)
.blockingForEach(event -> log.info("{} {}", event.author(), event.stringifyContent()));
maxLlmCalls: the budget that stops runaway agents
An agent loop calls the model, the model asks for a tool, the tool result goes back to the model, and so on until the model answers without a tool call. A confused model can loop: calling the same search tool with slightly different arguments, or two agents transferring control back and forth. Each iteration costs latency and money. maxLlmCalls is the circuit breaker.
The semantics are precise. Before each model call, the InvocationContext increments its counter and, only if maxLlmCalls is greater than zero, throws LlmCallsLimitExceededException when the count exceeds the limit. The flow logs the error and propagates it down the event stream, so the invocation ends with an error rather than a final answer. The builder rejects Integer.MAX_VALUE with an IllegalArgumentException and logs a warning for negative values, because a negative value means no limit at all.
The default of 500 is a safety net, not a budget. A typical tool-using conversation turn needs a handful of calls; a planner with sub-agents may need a few dozen. Measure the distribution of calls per invocation in your traces, set the budget a little above the high percentiles for each agent, and treat every budget exhaustion as a signal worth investigating. Never set it to zero to silence an error; that removes the only hard stop on cost. How calls are sequenced inside one step is covered in model call orchestration.
// Map the budget failure to a clear response instead of a 500 with a stack trace.
runner.runAsync(userId, sessionId, msg, factory.forRequest(req))
.onErrorResumeNext(err -> {
if (err instanceof LlmCallsLimitExceededException) {
metrics.counter("adk.llm_budget_exhausted", "agent", agentName).increment();
return Flowable.just(budgetExhaustedEvent(req)); // "I could not finish; here is what I found"
}
return Flowable.error(err);
})
.subscribe(sink::send, sink::fail, sink::complete);
Streaming modes
NONE asks the model for a complete response and emits one event per model turn. It is the simplest to reason about and right for batch jobs, tests and back-end callers that only want the final answer.
SSE asks the model for a streamed response and emits partial events as text arrives. Use it when a person is watching a chat interface and time to first token matters. Your transport must then forward partial events, and your consumers must distinguish partial from final events, or they will render text twice or persist fragments.
BIDI is for live sessions driven through runLive, where audio or text flows in both directions over a persistent connection, together with the response modality, speech and transcription fields. It is a different interaction model rather than a faster version of SSE, and it is covered in streaming in ADK Java.
Tool execution modes
When one model turn requests several function calls, toolExecutionMode decides how they run. In the current source the function-calling code behaves as follows:
| Mode | Behaviour |
|---|---|
SEQUENTIAL | Each tool is subscribed only after the previous one completes. Also used whenever there is only one call. |
PARALLEL and NONE | All tools are subscribed eagerly on the caller thread. Asynchronous tools run concurrently; tools that block the subscribing thread still run one after another. Result order is preserved. |
PARALLEL_SUBSCRIBE | Like PARALLEL, but each tool is subscribed on a worker scheduler, so blocking tools also run concurrently. Result order is still preserved. |
The trap is that the default, NONE, looks parallel but gives you no speed-up for ordinary blocking tools such as a JDBC query or a synchronous HTTP client. Conversely, switching to PARALLEL_SUBSCRIBE makes blocking tools concurrent, which is only safe if they are thread-safe and do not share mutable state, and if downstream systems can take the extra concurrency. Choose SEQUENTIAL when tools have side effects that must happen in order, such as reserve-then-charge.
Sessions and input blobs
With autoCreateSession(false), the default, a call with a session id that the session service does not know fails with an IllegalArgumentException whose message begins Session not found. With true, the Runner creates a session with that id. Auto-creation is convenient in prototypes and dangerous in multi-tenant services: a typo, a stale client or a guessed id silently starts a fresh, empty conversation instead of surfacing an error. Keep it off and create sessions explicitly, and read session context in depth for what a session carries.
With saveInputBlobsAsArtifacts(true), inline binary parts of the incoming message, such as an uploaded image or PDF, are saved through the artifact service under names of the form artifact_<invocationId>_<index>, and the message passed to the agent carries a reference instead of the bytes. That keeps large payloads out of session history, but it means those files now live in your artifact store, with its retention, access control and cost. See artifacts in ADK Java.
Building RunConfig from defaults and policy
The robust pattern is a small factory. Process configuration supplies validated defaults at start-up; the factory turns a request context into a RunConfig using explicit policy. Policy may lower limits but never raise them above the operator's ceiling, and every field is set explicitly so a library upgrade that changes a default does not change your behaviour.
public record RuntimeDefaults(int maxLlmCalls, StreamingMode streaming, ToolExecutionMode tools) {
static RuntimeDefaults fromEnv(Map<String, String> env) {
int calls = Integer.parseInt(env.getOrDefault("ADK_MAX_LLM_CALLS", "40"));
if (calls < 1 || calls > 200) {
throw new IllegalStateException("ADK_MAX_LLM_CALLS must be 1..200, got " + calls);
}
return new RuntimeDefaults(
calls,
StreamingMode.valueOf(env.getOrDefault("ADK_STREAMING_MODE", "NONE")),
ToolExecutionMode.valueOf(env.getOrDefault("ADK_TOOL_EXECUTION", "SEQUENTIAL")));
}
}
public final class RunConfigFactory {
private final RuntimeDefaults defaults;
public RunConfigFactory(RuntimeDefaults defaults) { this.defaults = defaults; }
public RunConfig forRequest(RequestContext req) {
int budget = switch (req.tier()) { // policy can lower the budget, never raise it
case FREE -> Math.min(defaults.maxLlmCalls(), 10);
case PRO -> defaults.maxLlmCalls();
};
StreamingMode streaming = req.acceptsEventStream() ? StreamingMode.SSE : StreamingMode.NONE;
return RunConfig.builder()
.maxLlmCalls(budget)
.streamingMode(defaults.streaming() == StreamingMode.NONE ? StreamingMode.NONE : streaming)
.toolExecutionMode(defaults.tools())
.autoCreateSession(false)
.saveInputBlobsAsArtifacts(req.hasAttachments())
.build();
}
}Validation happens once, at boot, so a bad ADK_MAX_LLM_CALLS stops the deployment instead of the first request, and valueOf fails fast on a misspelled enum name. Because the output is an immutable value, you can log it with the invocation id and replay any incident with exactly the same configuration.
Worked example: one agent, three callers
Take a support agent with a search tool and a ticket tool, served to a web chat, a mobile app and a nightly batch that summarises open tickets. Traces show a median of 4 model calls per invocation and a 99th percentile of 14. The operator sets ADK_MAX_LLM_CALLS=40, ADK_STREAMING_MODE=SSE and ADK_TOOL_EXECUTION=SEQUENTIAL, because the ticket tool must not run concurrently with a search that feeds it.
A free-tier web user sends a question: the factory produces a budget of 10, SSE streaming and no session auto-creation. A pro user on the same endpoint gets 40. The nightly batch, which does not accept an event stream, gets NONE. One week, a prompt change makes the agent loop on search. Free-tier invocations start ending with the budget-exhausted event, the counter alert fires within an hour, and cost is bounded by 10 calls per request instead of 500. Nothing was redeployed to contain it, and the fix was a prompt rollback.
Failure modes
- Silent defaults. Calling the overload without a
RunConfiggives a 500-call budget and default tool execution. Ban it in code review for production paths. - Budget as a bug hider. Raising
maxLlmCallsevery time it trips hides a looping agent and multiplies cost. - Imaginary parallelism. Expecting
NONEorPARALLELto speed up blocking tools; they still run one after another. - Unsafe concurrency.
PARALLEL_SUBSCRIBEwith tools that share a non-thread-safe client or depend on each other's side effects. - Partial events persisted. Enabling
SSEwithout teaching consumers to ignore partial events. - Ghost sessions.
autoCreateSession(true)in a multi-tenant service turns client bugs into empty conversations. - Version drift. Relying on a field or default that does not exist in the adk-java version you ship.
Operating and testing it
Log the effective RunConfig with every invocation id, and emit metrics for model calls per invocation, budget exhaustions per agent and tool latency by execution mode. Those three signals tell you whether the budget and concurrency settings match reality. Unit-test the factory as a pure function, and add an integration test that pins the Runner's behaviour for unknown sessions, so an upgrade that changes it fails in CI.
@Test
void freeTierGetsSmallBudgetAndNoAutoCreate() {
var factory = new RunConfigFactory(new RuntimeDefaults(40, StreamingMode.SSE, ToolExecutionMode.SEQUENTIAL));
RunConfig cfg = factory.forRequest(new RequestContext(Tier.FREE, true, false));
assertEquals(10, cfg.maxLlmCalls());
assertEquals(StreamingMode.SSE, cfg.streamingMode());
assertFalse(cfg.autoCreateSession());
}
@Test
void runnerRejectsUnknownSession() {
var cfg = RunConfig.builder().autoCreateSession(false).build();
runner.runAsync("u1", "no-such-session", msg, cfg)
.test().awaitDone(5, TimeUnit.SECONDS)
.assertError(IllegalArgumentException.class); // delivered down the stream, not thrown
}
What to do next
- Search your code for
runAsynccalls without aRunConfigargument and replace each with a factory call. - Pull a week of traces and compute model calls per invocation per agent; set
maxLlmCallsjust above the 99th percentile. - Add a metric and an alert on
LlmCallsLimitExceededException, and map it to a clear user-facing response. - Classify each tool as blocking or asynchronous and side-effecting or pure, then choose
SEQUENTIAL,PARALLELorPARALLEL_SUBSCRIBEon that basis. - Set
autoCreateSession(false)explicitly and create sessions through a dedicated endpoint. - Decide retention and access control for saved input blobs before enabling
saveInputBlobsAsArtifacts. - Check every field you use against the Javadoc of the adk-java version in your build file.