Most frameworks have a start-up phase with hooks: an init method, a lifecycle event, a place where configuration is checked and connections are opened. Developers coming to the Agent Development Kit for Java often look for the same thing and assume there is an onStart method on the agent where credentials are fetched and the model is validated. There is not. In ADK Java, booting is ordinary object construction: you build agents, tools, services and a Runner, and almost nothing is checked or contacted until the first request flows through. That design keeps construction cheap and testable, but it moves a whole class of configuration errors from start-up to the first customer conversation unless you add your own checks.

This article follows the sequence as it exists in current google/adk-java source: what each construction step does and does not validate, what happens inside runAsync on the first request, how plugins hook in, and how to add warm-up and readiness steps so a misconfigured deployment fails before it takes traffic. Code uses the public builder APIs; check exact signatures against the version you depend on, because the project moves quickly.

Advertisement

The mental model: assembly, then lazy resolution

An ADK Java application has four kinds of object. Agents (usually LlmAgent, plus workflow agents that compose them) describe behaviour: a name, a model name, an instruction, tools and optional sub-agents. Tools wrap your code so the model can call it. Services hold state: a session service stores conversations and their events, an artifact service stores files and blobs, and a memory service supports recall across sessions. The Runner ties an agent tree to those services and executes requests.

Boot is building these objects in dependency order. The important fact is what is not done at build time. In current source, LlmAgent's build() checks that the name is not empty and little else. The model is stored as a string and resolved to a model client only when the agent first needs it, through LlmRegistry.getLlm, which matches the name against registered patterns such as gemini-.* and caches the result. A misspelled model name therefore builds successfully and fails on the first request with an IllegalArgumentException reading Unsupported model.

ADK Java: assembly at boot, resolution on first use, then the per-request pathBOOT (your process)1. Agent treeLlmAgent.builder()...build()2. ToolsFunctionTool.create(...)3. Servicessession, artifact, memory4. Runner + pluginsRunner.builder() or App5. Warm-up + readinessresolve model, test callFIRST REQUEST: runner.runAsync(userId, sessionId, content)a. get session from session serviceb. save input blobs as artifactsc. plugins: onUserMessageCallbackd. build InvocationContext, append user evente. agent.runAsync: model resolved on first usef. append each non-partial event to sessiong. plugins: onEventCallback, afterRunCallbackreadyBoot builds objects. Nothing calls a model, and the model name is not checked, until step eunless your own warm-up step forces it. Errors in steps a-g surface on the returned stream.
Left: what your process does at boot. Right: what the Runner does on each request. The model client is created lazily in step e, so the first request also pays that one-time cost unless a warm-up step triggers it earlier.

Step 1 and 2: building the agent tree and its tools

Agents are immutable once built, so build them once at start-up and share them across requests. Tools are the usual source of boot-time surprises because FunctionTool.create uses reflection to find a method by name and derive a function declaration from its parameters and @Schema annotations.

import com.google.adk.agents.LlmAgent;
import com.google.adk.tools.Annotations.Schema;
import com.google.adk.tools.FunctionTool;
import java.util.Map;

public final class SupportAgents {

  public static Map<String, String> lookupOrder(
      @Schema(name = "orderId", description = "The order identifier") String orderId) {
    // Real code calls the order service; keep tools fast and bounded.
    return Map.of("status", "success", "state", "shipped");
  }

  public static LlmAgent rootAgent(String model) {
    return LlmAgent.builder()
        .name("support_agent")
        .model(model)
        .description("Answers order-status questions.")
        .instruction("You help customers check order status. Use lookupOrder for facts.")
        .tools(FunctionTool.create(SupportAgents.class, "lookupOrder"))
        .build();
  }
}

Keep construction free of I/O. The builder should receive configuration it can use immediately: model name, instruction text, tool objects. Anything that needs a network call, such as fetching a secret or opening a connection pool for a tool, belongs in your application's start-up code before the agent is built, with the resulting client passed into the tool. That keeps a unit test able to build the entire agent tree with no network, which is the cheapest way to catch a renamed tool method or a missing annotation: a test that builds the tree and asserts the tools are present runs in CI on every change.

Instructions can be a static string or a provider function evaluated per request. Placeholders such as {user_name} are filled from session state when a request runs, not at boot, so a missing state key is a per-request problem; seed required keys when sessions are created.

Advertisement

Step 3: services decide what survives a restart

Choosing services is the most consequential boot decision, because it determines what state exists after a restart and whether multiple replicas can share conversations. InMemoryRunner is the convenience entry point: its constructors create an in-memory artifact service, session service and memory service. That is right for development and tests and wrong for any deployment with more than one replica or any expectation of durability, because each process holds its own sessions and a restart loses them.

In production, construct durable implementations of BaseSessionService and the artifact and memory services, and pass them to the runner explicitly. The session service interface is reactive: createSession returns a Single<Session>, getSession a Maybe<Session>, appendEvent a Single<Event> and deleteSession a Completable. The runner appends every non-partial event it produces, so session-store write latency is on the critical path of every turn. Open its connection pool at boot and verify it with a trivial read before declaring the service ready. See the ADK Java memory service for how recall across sessions is wired.

Step 4: the Runner, App and plugins

The Runner is built with Runner.builder(), whose methods include agent, appName, app, sessionService, artifactService, memoryService and plugins. The older multi-argument constructors are deprecated in current source. Alternatively, an App bundles a name, the root agent and plugins; its build() throws if either the name or the root agent is missing and validates the name against the pattern [a-zA-Z_][a-zA-Z0-9_]*, rejecting the reserved name user. A hyphenated application name therefore fails there, at boot, which is one of the few validations that does happen early.

String model = System.getenv().getOrDefault("AGENT_MODEL", "gemini-2.5-flash");
LlmAgent root = SupportAgents.rootAgent(model);

// Fail at boot, not on the first customer request, if no registered
// factory matches the model name ("Unsupported model: ...").
LlmRegistry.getLlm(model);

Runner runner = Runner.builder()
    .agent(root)
    .appName("support_app")
    .sessionService(sessionService)     // durable implementation in production
    .artifactService(artifactService)
    .memoryService(memoryService)
    .plugins(new TimingPlugin())
    .build();

// Local development only: every service is in memory and is lost on restart.
// Runner dev = new InMemoryRunner(root, "support_app");

The explicit LlmRegistry.getLlm(model) call turns the lazy model check into an eager one. It performs pattern matching and client construction, not a network call, so it is cheap. If you register custom model factories with registerLlm, do it before this line.

Plugins are the cross-cutting hook points. The Plugin interface defines callbacks at every stage: onUserMessageCallback, beforeRunCallback, onEventCallback, afterRunCallback, onRunErrorCallback, before and after agent, model and tool callbacks, error callbacks for model and tool failures, and close. Extend the abstract BasePlugin, pass a name to its constructor and override only what you need. Agent-level callbacks, covered in ADK Java callbacks, apply to one agent; plugins apply to every run of the runner.

public final class TimingPlugin extends BasePlugin {
  private final ConcurrentMap<String, Long> started = new ConcurrentHashMap<>();

  public TimingPlugin() {
    super("timing");
  }

  @Override
  public Maybe<Content> beforeRunCallback(InvocationContext ctx) {
    started.put(ctx.invocationId(), System.nanoTime());
    return Maybe.empty();                       // empty: let the run proceed
  }

  @Override
  public Completable afterRunCallback(InvocationContext ctx) {
    record(ctx);
    return Completable.complete();
  }

  @Override
  public Completable onRunErrorCallback(InvocationContext ctx, Throwable error) {
    record(ctx);                                // without this the map leaks
    return Completable.complete();
  }

  private void record(InvocationContext ctx) {
    Long t0 = started.remove(ctx.invocationId());
    if (t0 != null) {
      metrics.recordRunMillis((System.nanoTime() - t0) / 1_000_000);
    }
  }
}

Two rules keep plugins safe. Clean up state on the error path as well as the success path, or per-invocation maps grow without bound. And never block inside a callback: the callbacks return RxJava types, and a blocking call inside one holds a thread that the rest of the pipeline is waiting on.

The first request, step by step

Calling runner.runAsync(userId, sessionId, content) returns a Flowable<Event>. In current source the runner then does the following, in order. It fetches the session from the session service. It saves any inline binary parts of the user message as artifacts and replaces them with placeholders, so large blobs are not stored in every event. It invokes each plugin's onUserMessageCallback, which may rewrite the message. It builds an InvocationContext holding the session, agent, services and run configuration, and appends the user's message to the session as an event. It calls the root agent's runAsync, which, for an LlmAgent, resolves the model on first use, builds the request from instruction, history and tool declarations, and calls the model. Each event the agent emits is appended to the session unless it is a partial streaming chunk, passed to onEventCallback, and emitted to your subscriber. When the run finishes, afterRunCallback runs.

Two consequences follow. First, every failure, whether a missing session, an unsupported model, a credential error from the model API or an exception in a tool, arrives as an error on the returned stream. Code that subscribes without an error handler will lose it. Second, the first request is slower than the rest because it pays for model-client construction and cold connections. The model-call loop inside the agent is covered in the ADK Java execution loop.

Step 5: warm-up and readiness

Because so much is lazy, a production service needs an explicit readiness step that exercises the real path once before the load balancer sends traffic. The warm-up below creates a throwaway session, runs a one-line prompt through the runner with a timeout, and deletes the session.

static void warmUp(Runner runner) {
  String user = "warmup-user";
  Session session = runner.sessionService()
      .createSession(runner.appName(), user)
      .blockingGet();
  try {
    List<Event> events = runner
        .runAsync(user, session.id(), Content.fromParts(Part.fromText("Reply with: ready")))
        .timeout(30, TimeUnit.SECONDS)
        .toList()
        .blockingGet();
    if (events.isEmpty()) {
      throw new IllegalStateException("warm-up produced no events");
    }
  } finally {
    runner.sessionService()
        .deleteSession(runner.appName(), user, session.id())
        .blockingAwait();
  }
}

This proves the session store, credentials, network egress to the model endpoint, model name and the agent's request construction together. It costs one model call per replica start, which is usually acceptable; if it is not, keep the LlmRegistry check and a session-store probe and accept that credentials are first tested by real traffic. Use it to gate initial readiness only, not as a recurring probe: the model endpoint is shared by every replica, so a provider outage would mark them all unready at once, and a liveness check against it would restart them in a loop. For the development UI, AdkWebServer.start(rootAgent) launches a local web server; production services normally embed the runner in their own HTTP layer, as described in ADK Java with Spring.

Worked example: a typo that passes every health check

A team deploys with AGENT_MODEL=gemini-2.5-flsh. Every boot step succeeds: LlmAgent.builder().build() checks only the name, the runner builds, and a health check that returns HTTP 200 once the process is up passes on every replica. The first customer message reaches runAsync, the agent asks LlmRegistry for the model, no registered pattern matches, and the stream fails with Unsupported model: gemini-2.5-flsh. Every request fails the same way until someone reads the logs.

With the one-line LlmRegistry.getLlm(model) call placed before the runner is built, the same deployment throws during start-up, the new replicas never become ready, and the rollout halts with the old version still serving. The fix did not make anything more correct; it moved the failure to the only moment when failing is cheap.

Failure modes and where they surface

FailureSurfaces atPrevention
Misspelled or unregistered model namefirst request, IllegalArgumentExceptioncall LlmRegistry.getLlm at boot
Invalid application name in Appboot, App.build()use letters, digits and underscores
Missing API key or cloud credentialsfirst model callwarm-up call; secret check before build
In-memory session service behind a load balancerconversations lose history between replicasdurable session service in every non-local environment
Unknown session IDIllegalArgumentException on the stream unless the run config auto-creates sessionscreate sessions explicitly and handle the error
Blocking call in a plugin or toollatency spikes and stalled streamsoffload to a bounded executor; set tool timeouts
Plugin state leak on errorsslow memory growthclean up in onRunErrorCallback
Unsubscribed or unhandled stream errorssilent failuresalways subscribe with an error handler

Operating the runtime

Treat configuration as input to boot, not as something agents discover. Resolve the model name, instruction version, tool endpoints and service connection strings from one configuration source, log the resolved values (without secrets) once at start-up, and include the instruction version in every trace. Build exactly one runner per application per process and share it; it is designed to serve concurrent requests, and building one per request repeats every lazy initialisation.

On shutdown, stop accepting requests, let in-flight streams finish or cancel them at a deadline, then call each plugin's close and close service clients, in that order. The draining pattern is covered in graceful shutdown for ADK Java. Record first-request latency separately from steady-state latency so a regression in lazy initialisation is visible.

What to do next

  1. Remove any code that expects an agent start-up hook, and move I/O into application start-up before the agent tree is built.
  2. Add a unit test that builds the whole agent tree, including every FunctionTool, with no network.
  3. Call LlmRegistry.getLlm on every configured model name at boot.
  4. Replace InMemoryRunner with Runner.builder() and durable services in every shared environment.
  5. Add a warm-up run with a timeout and connect it to the readiness probe.
  6. Audit plugins for blocking calls and for cleanup on the error path.
  7. Make sure every runAsync subscriber handles errors, and track first-request latency as its own metric.
Key takeaway: ADK Java has no agent start-up lifecycle: booting is building agents, tools, services and a Runner, and model resolution, credential checks and session access all happen lazily on the first request. Make those failures happen at boot instead: resolve models through LlmRegistry, use durable services, add a warm-up run behind the readiness probe, keep plugins non-blocking and clean on errors, and share one runner per process.