Most frameworks have a start-up phase with hooks: an init method, a lifecycle event, a place where configuration is checked and connections are opened. Developers coming to the Agent Development Kit for Java often look for the same thing and assume there is an onStart method on the agent where credentials are fetched and the model is validated. There is not. In ADK Java, booting is ordinary object construction: you build agents, tools, services and a Runner, and almost nothing is checked or contacted until the first request flows through. That design keeps construction cheap and testable, but it moves a whole class of configuration errors from start-up to the first customer conversation unless you add your own checks.
This article follows the sequence as it exists in current google/adk-java source: what each construction step does and does not validate, what happens inside runAsync on the first request, how plugins hook in, and how to add warm-up and readiness steps so a misconfigured deployment fails before it takes traffic. Code uses the public builder APIs; check exact signatures against the version you depend on, because the project moves quickly.
The mental model: assembly, then lazy resolution
An ADK Java application has four kinds of object. Agents (usually LlmAgent, plus workflow agents that compose them) describe behaviour: a name, a model name, an instruction, tools and optional sub-agents. Tools wrap your code so the model can call it. Services hold state: a session service stores conversations and their events, an artifact service stores files and blobs, and a memory service supports recall across sessions. The Runner ties an agent tree to those services and executes requests.
Boot is building these objects in dependency order. The important fact is what is not done at build time. In current source, LlmAgent's build() checks that the name is not empty and little else. The model is stored as a string and resolved to a model client only when the agent first needs it, through LlmRegistry.getLlm, which matches the name against registered patterns such as gemini-.* and caches the result. A misspelled model name therefore builds successfully and fails on the first request with an IllegalArgumentException reading Unsupported model.
Step 1 and 2: building the agent tree and its tools
Agents are immutable once built, so build them once at start-up and share them across requests. Tools are the usual source of boot-time surprises because FunctionTool.create uses reflection to find a method by name and derive a function declaration from its parameters and @Schema annotations.
import com.google.adk.agents.LlmAgent;
import com.google.adk.tools.Annotations.Schema;
import com.google.adk.tools.FunctionTool;
import java.util.Map;
public final class SupportAgents {
public static Map<String, String> lookupOrder(
@Schema(name = "orderId", description = "The order identifier") String orderId) {
// Real code calls the order service; keep tools fast and bounded.
return Map.of("status", "success", "state", "shipped");
}
public static LlmAgent rootAgent(String model) {
return LlmAgent.builder()
.name("support_agent")
.model(model)
.description("Answers order-status questions.")
.instruction("You help customers check order status. Use lookupOrder for facts.")
.tools(FunctionTool.create(SupportAgents.class, "lookupOrder"))
.build();
}
}Keep construction free of I/O. The builder should receive configuration it can use immediately: model name, instruction text, tool objects. Anything that needs a network call, such as fetching a secret or opening a connection pool for a tool, belongs in your application's start-up code before the agent is built, with the resulting client passed into the tool. That keeps a unit test able to build the entire agent tree with no network, which is the cheapest way to catch a renamed tool method or a missing annotation: a test that builds the tree and asserts the tools are present runs in CI on every change.
Instructions can be a static string or a provider function evaluated per request. Placeholders such as {user_name} are filled from session state when a request runs, not at boot, so a missing state key is a per-request problem; seed required keys when sessions are created.
Step 3: services decide what survives a restart
Choosing services is the most consequential boot decision, because it determines what state exists after a restart and whether multiple replicas can share conversations. InMemoryRunner is the convenience entry point: its constructors create an in-memory artifact service, session service and memory service. That is right for development and tests and wrong for any deployment with more than one replica or any expectation of durability, because each process holds its own sessions and a restart loses them.
In production, construct durable implementations of BaseSessionService and the artifact and memory services, and pass them to the runner explicitly. The session service interface is reactive: createSession returns a Single<Session>, getSession a Maybe<Session>, appendEvent a Single<Event> and deleteSession a Completable. The runner appends every non-partial event it produces, so session-store write latency is on the critical path of every turn. Open its connection pool at boot and verify it with a trivial read before declaring the service ready. See the ADK Java memory service for how recall across sessions is wired.
Step 4: the Runner, App and plugins
The Runner is built with Runner.builder(), whose methods include agent, appName, app, sessionService, artifactService, memoryService and plugins. The older multi-argument constructors are deprecated in current source. Alternatively, an App bundles a name, the root agent and plugins; its build() throws if either the name or the root agent is missing and validates the name against the pattern [a-zA-Z_][a-zA-Z0-9_]*, rejecting the reserved name user. A hyphenated application name therefore fails there, at boot, which is one of the few validations that does happen early.
String model = System.getenv().getOrDefault("AGENT_MODEL", "gemini-2.5-flash");
LlmAgent root = SupportAgents.rootAgent(model);
// Fail at boot, not on the first customer request, if no registered
// factory matches the model name ("Unsupported model: ...").
LlmRegistry.getLlm(model);
Runner runner = Runner.builder()
.agent(root)
.appName("support_app")
.sessionService(sessionService) // durable implementation in production
.artifactService(artifactService)
.memoryService(memoryService)
.plugins(new TimingPlugin())
.build();
// Local development only: every service is in memory and is lost on restart.
// Runner dev = new InMemoryRunner(root, "support_app");The explicit LlmRegistry.getLlm(model) call turns the lazy model check into an eager one. It performs pattern matching and client construction, not a network call, so it is cheap. If you register custom model factories with registerLlm, do it before this line.
Plugins are the cross-cutting hook points. The Plugin interface defines callbacks at every stage: onUserMessageCallback, beforeRunCallback, onEventCallback, afterRunCallback, onRunErrorCallback, before and after agent, model and tool callbacks, error callbacks for model and tool failures, and close. Extend the abstract BasePlugin, pass a name to its constructor and override only what you need. Agent-level callbacks, covered in ADK Java callbacks, apply to one agent; plugins apply to every run of the runner.
public final class TimingPlugin extends BasePlugin {
private final ConcurrentMap<String, Long> started = new ConcurrentHashMap<>();
public TimingPlugin() {
super("timing");
}
@Override
public Maybe<Content> beforeRunCallback(InvocationContext ctx) {
started.put(ctx.invocationId(), System.nanoTime());
return Maybe.empty(); // empty: let the run proceed
}
@Override
public Completable afterRunCallback(InvocationContext ctx) {
record(ctx);
return Completable.complete();
}
@Override
public Completable onRunErrorCallback(InvocationContext ctx, Throwable error) {
record(ctx); // without this the map leaks
return Completable.complete();
}
private void record(InvocationContext ctx) {
Long t0 = started.remove(ctx.invocationId());
if (t0 != null) {
metrics.recordRunMillis((System.nanoTime() - t0) / 1_000_000);
}
}
}Two rules keep plugins safe. Clean up state on the error path as well as the success path, or per-invocation maps grow without bound. And never block inside a callback: the callbacks return RxJava types, and a blocking call inside one holds a thread that the rest of the pipeline is waiting on.
The first request, step by step
Calling runner.runAsync(userId, sessionId, content) returns a Flowable<Event>. In current source the runner then does the following, in order. It fetches the session from the session service. It saves any inline binary parts of the user message as artifacts and replaces them with placeholders, so large blobs are not stored in every event. It invokes each plugin's onUserMessageCallback, which may rewrite the message. It builds an InvocationContext holding the session, agent, services and run configuration, and appends the user's message to the session as an event. It calls the root agent's runAsync, which, for an LlmAgent, resolves the model on first use, builds the request from instruction, history and tool declarations, and calls the model. Each event the agent emits is appended to the session unless it is a partial streaming chunk, passed to onEventCallback, and emitted to your subscriber. When the run finishes, afterRunCallback runs.
Two consequences follow. First, every failure, whether a missing session, an unsupported model, a credential error from the model API or an exception in a tool, arrives as an error on the returned stream. Code that subscribes without an error handler will lose it. Second, the first request is slower than the rest because it pays for model-client construction and cold connections. The model-call loop inside the agent is covered in the ADK Java execution loop.
Step 5: warm-up and readiness
Because so much is lazy, a production service needs an explicit readiness step that exercises the real path once before the load balancer sends traffic. The warm-up below creates a throwaway session, runs a one-line prompt through the runner with a timeout, and deletes the session.
static void warmUp(Runner runner) {
String user = "warmup-user";
Session session = runner.sessionService()
.createSession(runner.appName(), user)
.blockingGet();
try {
List<Event> events = runner
.runAsync(user, session.id(), Content.fromParts(Part.fromText("Reply with: ready")))
.timeout(30, TimeUnit.SECONDS)
.toList()
.blockingGet();
if (events.isEmpty()) {
throw new IllegalStateException("warm-up produced no events");
}
} finally {
runner.sessionService()
.deleteSession(runner.appName(), user, session.id())
.blockingAwait();
}
}This proves the session store, credentials, network egress to the model endpoint, model name and the agent's request construction together. It costs one model call per replica start, which is usually acceptable; if it is not, keep the LlmRegistry check and a session-store probe and accept that credentials are first tested by real traffic. Use it to gate initial readiness only, not as a recurring probe: the model endpoint is shared by every replica, so a provider outage would mark them all unready at once, and a liveness check against it would restart them in a loop. For the development UI, AdkWebServer.start(rootAgent) launches a local web server; production services normally embed the runner in their own HTTP layer, as described in ADK Java with Spring.
Worked example: a typo that passes every health check
A team deploys with AGENT_MODEL=gemini-2.5-flsh. Every boot step succeeds: LlmAgent.builder().build() checks only the name, the runner builds, and a health check that returns HTTP 200 once the process is up passes on every replica. The first customer message reaches runAsync, the agent asks LlmRegistry for the model, no registered pattern matches, and the stream fails with Unsupported model: gemini-2.5-flsh. Every request fails the same way until someone reads the logs.
With the one-line LlmRegistry.getLlm(model) call placed before the runner is built, the same deployment throws during start-up, the new replicas never become ready, and the rollout halts with the old version still serving. The fix did not make anything more correct; it moved the failure to the only moment when failing is cheap.
Failure modes and where they surface
| Failure | Surfaces at | Prevention |
|---|---|---|
| Misspelled or unregistered model name | first request, IllegalArgumentException | call LlmRegistry.getLlm at boot |
| Invalid application name in App | boot, App.build() | use letters, digits and underscores |
| Missing API key or cloud credentials | first model call | warm-up call; secret check before build |
| In-memory session service behind a load balancer | conversations lose history between replicas | durable session service in every non-local environment |
| Unknown session ID | IllegalArgumentException on the stream unless the run config auto-creates sessions | create sessions explicitly and handle the error |
| Blocking call in a plugin or tool | latency spikes and stalled streams | offload to a bounded executor; set tool timeouts |
| Plugin state leak on errors | slow memory growth | clean up in onRunErrorCallback |
| Unsubscribed or unhandled stream errors | silent failures | always subscribe with an error handler |
Operating the runtime
Treat configuration as input to boot, not as something agents discover. Resolve the model name, instruction version, tool endpoints and service connection strings from one configuration source, log the resolved values (without secrets) once at start-up, and include the instruction version in every trace. Build exactly one runner per application per process and share it; it is designed to serve concurrent requests, and building one per request repeats every lazy initialisation.
On shutdown, stop accepting requests, let in-flight streams finish or cancel them at a deadline, then call each plugin's close and close service clients, in that order. The draining pattern is covered in graceful shutdown for ADK Java. Record first-request latency separately from steady-state latency so a regression in lazy initialisation is visible.
What to do next
- Remove any code that expects an agent start-up hook, and move I/O into application start-up before the agent tree is built.
- Add a unit test that builds the whole agent tree, including every FunctionTool, with no network.
- Call LlmRegistry.getLlm on every configured model name at boot.
- Replace InMemoryRunner with Runner.builder() and durable services in every shared environment.
- Add a warm-up run with a timeout and connect it to the readiness probe.
- Audit plugins for blocking calls and for cleanup on the error path.
- Make sure every runAsync subscriber handles errors, and track first-request latency as its own metric.