An ADK Java agent has three kinds of configuration, and most production incidents come from confusing them. Process configuration, such as API keys and model names, is read once when the service starts. Assembly configuration, the agent tree, the session and artifact services and the plugins, is fixed when the Runner is built. And runtime configuration, how one particular invocation should behave, is passed with every call as a RunConfig.

This article is about that third layer. It explains what each RunConfig field does and what it defaults to, how the LLM call budget stops runaway loops, what the streaming and tool execution modes really change, and how to build a RunConfig per request from environment defaults and tenant policy, so it is validated, testable and observable. The field list and semantics below were checked against the current adk-java source; some fields are recent, so confirm them against the version you depend on.

Advertisement

Three layers of configuration

LayerExamplesWhen it changesWhere it is covered
ProcessAPI keys, model id, database URL, feature flagsDeploy or restartenvironment and configuration management
AssemblyAgent tree, tools, session / artifact / memory services, pluginsBuild of the Runnerruntime boot sequence
InvocationLLM call budget, streaming, tool concurrency, session creationEvery callThis article

The distinction matters because the layers have different lifetimes and different owners. An operator changes process configuration. A developer changes assembly. But invocation configuration can legitimately differ between two requests arriving in the same millisecond: a free-tier user gets a smaller LLM budget, a browser that accepts server-sent events gets streaming, a batch job does not. If those decisions leak into process configuration, you end up redeploying to change a per-customer limit. If they are hard-coded at the call site, nobody can see or change them.

Where RunConfig enters the runtime

Process configurationenv vars, secrets, profilesApp assemblyagent tree, services, pluginsRequest contexttenant, tier, endpointRunConfigFactorydefaults from env + overrides from policy -> RunConfig (immutable)Runnerrunner.runAsync(user, session, msg, cfg)or runLive(..., queue, cfg)InvocationContextcounts LLM calls vs maxLlmCallsLLM flowstreamingMode NONE / SSE / BIDIFunction executiontoolExecutionMode
Process defaults and request context feed a factory that produces an immutable RunConfig per call. The Runner passes it into the InvocationContext, where the LLM flow and function execution read it.

The Runner exposes runAsync overloads that take a user id, a session id (or a SessionKey), the new message and a RunConfig, optionally with an initial state delta. For live, bidirectional sessions there is runLive, which takes a LiveRequestQueue instead of a single message. Overloads that omit the config use RunConfig.builder().build(), which means every default described below applies silently.

Inside, the Runner creates an InvocationContext holding the config for the duration of that invocation. The LLM flow reads the streaming mode, the context counts model calls against the budget, and the function-calling code reads the tool execution mode. The config is an immutable value object built with a builder, so it is safe to share across threads and trivial to log.

Advertisement

The fields and their defaults

FieldDefaultWhat it controls
maxLlmCalls500Maximum model calls in one invocation; 0 or negative disables the check
streamingModeNONENONE, SSE or BIDI
toolExecutionModeNONEHow several function calls from one model turn are executed
autoCreateSessionfalseWhether an unknown session id is created or rejected
saveInputBlobsAsArtifactsfalseWhether inline binary parts of the user message are stored as artifacts
responseModalitiesemptyRequested output modalities, for example audio for live agents
speechConfig, inputAudioTranscription, outputAudioTranscriptionunsetVoice and transcription settings for live sessions
customMetadataempty mapCaller-supplied metadata carried with the config

The source also contains fields for avatar settings and an override for how function responses are grouped in history; treat those as advanced and check the Javadoc of your version before relying on them. The two fields that matter most for reliability are the call budget and the tool execution mode, and both have defaults you should override deliberately.

import com.google.adk.agents.RunConfig;
import com.google.adk.agents.RunConfig.StreamingMode;
import com.google.adk.agents.RunConfig.ToolExecutionMode;

RunConfig cfg = RunConfig.builder()
    .maxLlmCalls(25)                                   // default 500; 0 or negative disables the check
    .streamingMode(StreamingMode.SSE)                  // default NONE
    .toolExecutionMode(ToolExecutionMode.PARALLEL)     // default NONE, which behaves like PARALLEL
    .autoCreateSession(false)                          // default false: unknown session is an error
    .saveInputBlobsAsArtifacts(true)                   // default false
    .build();

runner.runAsync(userId, sessionId, userMessage, cfg)
    .blockingForEach(event -> log.info("{} {}", event.author(), event.stringifyContent()));

maxLlmCalls: the budget that stops runaway agents

An agent loop calls the model, the model asks for a tool, the tool result goes back to the model, and so on until the model answers without a tool call. A confused model can loop: calling the same search tool with slightly different arguments, or two agents transferring control back and forth. Each iteration costs latency and money. maxLlmCalls is the circuit breaker.

The semantics are precise. Before each model call, the InvocationContext increments its counter and, only if maxLlmCalls is greater than zero, throws LlmCallsLimitExceededException when the count exceeds the limit. The flow logs the error and propagates it down the event stream, so the invocation ends with an error rather than a final answer. The builder rejects Integer.MAX_VALUE with an IllegalArgumentException and logs a warning for negative values, because a negative value means no limit at all.

The default of 500 is a safety net, not a budget. A typical tool-using conversation turn needs a handful of calls; a planner with sub-agents may need a few dozen. Measure the distribution of calls per invocation in your traces, set the budget a little above the high percentiles for each agent, and treat every budget exhaustion as a signal worth investigating. Never set it to zero to silence an error; that removes the only hard stop on cost. How calls are sequenced inside one step is covered in model call orchestration.

// Map the budget failure to a clear response instead of a 500 with a stack trace.
runner.runAsync(userId, sessionId, msg, factory.forRequest(req))
    .onErrorResumeNext(err -> {
        if (err instanceof LlmCallsLimitExceededException) {
            metrics.counter("adk.llm_budget_exhausted", "agent", agentName).increment();
            return Flowable.just(budgetExhaustedEvent(req));   // "I could not finish; here is what I found"
        }
        return Flowable.error(err);
    })
    .subscribe(sink::send, sink::fail, sink::complete);

Streaming modes

NONE asks the model for a complete response and emits one event per model turn. It is the simplest to reason about and right for batch jobs, tests and back-end callers that only want the final answer.

SSE asks the model for a streamed response and emits partial events as text arrives. Use it when a person is watching a chat interface and time to first token matters. Your transport must then forward partial events, and your consumers must distinguish partial from final events, or they will render text twice or persist fragments.

BIDI is for live sessions driven through runLive, where audio or text flows in both directions over a persistent connection, together with the response modality, speech and transcription fields. It is a different interaction model rather than a faster version of SSE, and it is covered in streaming in ADK Java.

Tool execution modes

When one model turn requests several function calls, toolExecutionMode decides how they run. In the current source the function-calling code behaves as follows:

ModeBehaviour
SEQUENTIALEach tool is subscribed only after the previous one completes. Also used whenever there is only one call.
PARALLEL and NONEAll tools are subscribed eagerly on the caller thread. Asynchronous tools run concurrently; tools that block the subscribing thread still run one after another. Result order is preserved.
PARALLEL_SUBSCRIBELike PARALLEL, but each tool is subscribed on a worker scheduler, so blocking tools also run concurrently. Result order is still preserved.

The trap is that the default, NONE, looks parallel but gives you no speed-up for ordinary blocking tools such as a JDBC query or a synchronous HTTP client. Conversely, switching to PARALLEL_SUBSCRIBE makes blocking tools concurrent, which is only safe if they are thread-safe and do not share mutable state, and if downstream systems can take the extra concurrency. Choose SEQUENTIAL when tools have side effects that must happen in order, such as reserve-then-charge.

Sessions and input blobs

With autoCreateSession(false), the default, a call with a session id that the session service does not know fails with an IllegalArgumentException whose message begins Session not found. With true, the Runner creates a session with that id. Auto-creation is convenient in prototypes and dangerous in multi-tenant services: a typo, a stale client or a guessed id silently starts a fresh, empty conversation instead of surfacing an error. Keep it off and create sessions explicitly, and read session context in depth for what a session carries.

With saveInputBlobsAsArtifacts(true), inline binary parts of the incoming message, such as an uploaded image or PDF, are saved through the artifact service under names of the form artifact_<invocationId>_<index>, and the message passed to the agent carries a reference instead of the bytes. That keeps large payloads out of session history, but it means those files now live in your artifact store, with its retention, access control and cost. See artifacts in ADK Java.

Building RunConfig from defaults and policy

The robust pattern is a small factory. Process configuration supplies validated defaults at start-up; the factory turns a request context into a RunConfig using explicit policy. Policy may lower limits but never raise them above the operator's ceiling, and every field is set explicitly so a library upgrade that changes a default does not change your behaviour.

public record RuntimeDefaults(int maxLlmCalls, StreamingMode streaming, ToolExecutionMode tools) {
    static RuntimeDefaults fromEnv(Map<String, String> env) {
        int calls = Integer.parseInt(env.getOrDefault("ADK_MAX_LLM_CALLS", "40"));
        if (calls < 1 || calls > 200) {
            throw new IllegalStateException("ADK_MAX_LLM_CALLS must be 1..200, got " + calls);
        }
        return new RuntimeDefaults(
            calls,
            StreamingMode.valueOf(env.getOrDefault("ADK_STREAMING_MODE", "NONE")),
            ToolExecutionMode.valueOf(env.getOrDefault("ADK_TOOL_EXECUTION", "SEQUENTIAL")));
    }
}

public final class RunConfigFactory {
    private final RuntimeDefaults defaults;
    public RunConfigFactory(RuntimeDefaults defaults) { this.defaults = defaults; }

    public RunConfig forRequest(RequestContext req) {
        int budget = switch (req.tier()) {            // policy can lower the budget, never raise it
            case FREE -> Math.min(defaults.maxLlmCalls(), 10);
            case PRO  -> defaults.maxLlmCalls();
        };
        StreamingMode streaming = req.acceptsEventStream() ? StreamingMode.SSE : StreamingMode.NONE;
        return RunConfig.builder()
            .maxLlmCalls(budget)
            .streamingMode(defaults.streaming() == StreamingMode.NONE ? StreamingMode.NONE : streaming)
            .toolExecutionMode(defaults.tools())
            .autoCreateSession(false)
            .saveInputBlobsAsArtifacts(req.hasAttachments())
            .build();
    }
}

Validation happens once, at boot, so a bad ADK_MAX_LLM_CALLS stops the deployment instead of the first request, and valueOf fails fast on a misspelled enum name. Because the output is an immutable value, you can log it with the invocation id and replay any incident with exactly the same configuration.

Worked example: one agent, three callers

Take a support agent with a search tool and a ticket tool, served to a web chat, a mobile app and a nightly batch that summarises open tickets. Traces show a median of 4 model calls per invocation and a 99th percentile of 14. The operator sets ADK_MAX_LLM_CALLS=40, ADK_STREAMING_MODE=SSE and ADK_TOOL_EXECUTION=SEQUENTIAL, because the ticket tool must not run concurrently with a search that feeds it.

A free-tier web user sends a question: the factory produces a budget of 10, SSE streaming and no session auto-creation. A pro user on the same endpoint gets 40. The nightly batch, which does not accept an event stream, gets NONE. One week, a prompt change makes the agent loop on search. Free-tier invocations start ending with the budget-exhausted event, the counter alert fires within an hour, and cost is bounded by 10 calls per request instead of 500. Nothing was redeployed to contain it, and the fix was a prompt rollback.

Failure modes

  • Silent defaults. Calling the overload without a RunConfig gives a 500-call budget and default tool execution. Ban it in code review for production paths.
  • Budget as a bug hider. Raising maxLlmCalls every time it trips hides a looping agent and multiplies cost.
  • Imaginary parallelism. Expecting NONE or PARALLEL to speed up blocking tools; they still run one after another.
  • Unsafe concurrency. PARALLEL_SUBSCRIBE with tools that share a non-thread-safe client or depend on each other's side effects.
  • Partial events persisted. Enabling SSE without teaching consumers to ignore partial events.
  • Ghost sessions. autoCreateSession(true) in a multi-tenant service turns client bugs into empty conversations.
  • Version drift. Relying on a field or default that does not exist in the adk-java version you ship.

Operating and testing it

Log the effective RunConfig with every invocation id, and emit metrics for model calls per invocation, budget exhaustions per agent and tool latency by execution mode. Those three signals tell you whether the budget and concurrency settings match reality. Unit-test the factory as a pure function, and add an integration test that pins the Runner's behaviour for unknown sessions, so an upgrade that changes it fails in CI.

@Test
void freeTierGetsSmallBudgetAndNoAutoCreate() {
    var factory = new RunConfigFactory(new RuntimeDefaults(40, StreamingMode.SSE, ToolExecutionMode.SEQUENTIAL));
    RunConfig cfg = factory.forRequest(new RequestContext(Tier.FREE, true, false));
    assertEquals(10, cfg.maxLlmCalls());
    assertEquals(StreamingMode.SSE, cfg.streamingMode());
    assertFalse(cfg.autoCreateSession());
}

@Test
void runnerRejectsUnknownSession() {
    var cfg = RunConfig.builder().autoCreateSession(false).build();
    runner.runAsync("u1", "no-such-session", msg, cfg)
        .test().awaitDone(5, TimeUnit.SECONDS)
        .assertError(IllegalArgumentException.class);   // delivered down the stream, not thrown
}

What to do next

  1. Search your code for runAsync calls without a RunConfig argument and replace each with a factory call.
  2. Pull a week of traces and compute model calls per invocation per agent; set maxLlmCalls just above the 99th percentile.
  3. Add a metric and an alert on LlmCallsLimitExceededException, and map it to a clear user-facing response.
  4. Classify each tool as blocking or asynchronous and side-effecting or pure, then choose SEQUENTIAL, PARALLEL or PARALLEL_SUBSCRIBE on that basis.
  5. Set autoCreateSession(false) explicitly and create sessions through a dedicated endpoint.
  6. Decide retention and access control for saved input blobs before enabling saveInputBlobsAsArtifacts.
  7. Check every field you use against the Javadoc of the adk-java version in your build file.
Key takeaway: RunConfig is ADK Java's per-invocation configuration: the LLM call budget, streaming mode, tool execution mode, session auto-creation and input-blob handling. Defaults are permissive (500 calls, no auto-creation, NONE for streaming and tools), and NONE for tools does not make blocking tools concurrent. Build RunConfig in one factory from validated environment defaults and request policy, set every field explicitly, log it with the invocation id, alert on budget exhaustion, and verify each field against the adk-java version you ship.