Most server frameworks teach middleware as an onion: each layer receives a request and a next handle, may change the request, calls next, may change the response and returns it. It is tempting to look for that shape in ADK for Java, and some write-ups even show an AgentMiddleware interface with an AgentChain. Neither type exists in ADK. What ADK has instead are hooks: per-agent callbacks registered on the agent builder, and plugins registered once on the Runner that apply to every agent, model call and tool call in it.

Hooks can do everything middleware does, but they compose by different rules, and the differences cause real bugs: metering that silently undercounts, redaction that runs on some responses but not others, timers that leak memory on the error path. This article reads the composition rules from the ADK source, shows how to build a genuine middleware pipeline on top of them, and gives an ordering discipline for a production plugin stack. The plugin contract itself is introduced in extending the runtime.

The hook points

ADK exposes hooks at four levels: the run, the agent, the model call and the tool call. Plugins see all of them; agent callbacks see the last three for their own agent. Every intercepting hook returns an RxJava Maybe (or an Optional for the Sync callback variants). Empty means carry on; a value means use this instead.

Hook pointPlugin methodA returned value means
User messageonUserMessageCallbackReplaces the user message
Run startbeforeRunCallbackHalts the run with that content
Each eventonEventCallbackReplaces the event
Run end / errorafterRunCallback, onRunErrorCallbackNothing: these return Completable
AgentbeforeAgentCallback, afterAgentCallbackSkips the agent / replaces its output
ModelbeforeModelCallback, afterModelCallbackSkips the call / replaces the response
Model erroronModelErrorCallbackUsed as the response instead of failing
ToolbeforeToolCallback, afterToolCallbackSkips the tool / replaces its result map
Tool erroronToolErrorCallbackUsed as the result instead of failing

Signatures differ between the two families, so copy them from the source rather than from memory. A plugin's beforeToolCallback(BaseTool, Map<String,Object>, ToolContext) has no InvocationContext parameter, while the agent callback type Callbacks.BeforeToolCallback receives one first. A plugin's afterToolCallback gets the result as a Map<String,Object>; the agent callback gets an Object. A plugin's onModelErrorCallback receives the LlmRequest.Builder; the agent variant receives the built LlmRequest.

How a chain is evaluated

How ADK evaluates one before-model hook pointPlugin 1registration orderPlugin 2Agent callback 1list orderAgent callback 2Model callBaseLlmall emptyempty: continueempty: continueempty: continueFirst non-empty result winsreturned as the response; later hooks skippedMutating the shared LlmRequest.Builder and returning empty composes.Returning a value short-circuits everything registered after you.
Plugins run in registration order, then agent callbacks in list order. The first hook to return a value decides the outcome; nothing after it runs.

Three rules, all visible in the source, decide what runs. First, the PluginManager runs plugins in registration order with concatMapMaybe and takes firstElement(): the first plugin to return a value ends the chain, and later plugins are never called for that hook. Second, plugins run before agent callbacks. In BaseLlmFlow the agent's callbacks are attached with pluginResult.switchIfEmpty(...), so if any plugin answered, no agent callback runs; BaseAgent does the same for agent-level hooks by putting the plugin callback at the front of the list. Third, inside an agent's own callback list the same first-non-empty rule applies.

// Shape of BaseLlmFlow.handleBeforeModelCallback, simplified
Maybe<LlmResponse> pluginResult =
    context.pluginManager().beforeModelCallback(callbackContext, llmRequestBuilder);
Maybe<LlmResponse> agentResult = Maybe.defer(() ->
    Flowable.fromIterable(agent.canonicalBeforeModelCallbacks())
        .concatMapMaybe(cb -> cb.call(callbackContext, llmRequestBuilder))
        .firstElement());
return pluginResult.switchIfEmpty(agentResult);   // agent callbacks only if plugins were silent

The after-model hook follows the same pattern and ends with defaultIfEmpty(llmResponse). Note what that implies: every after-model hook receives the original response, not the output of the previous hook. There is no piping of results. Hooks are a chain of first responders, not an onion.

Agent callbacks are registered on the builder, either one at a time or as a list whose order is the evaluation order. The list overloads take the base marker types, so asynchronous callbacks can share a list; the Sync variants, which return Optional and are wrapped into the asynchronous form internally, are registered singly.

LlmAgent refunds = LlmAgent.builder()
    .name("refunds")
    .model("gemini-2.5-flash")
    .instruction("Handle refund requests.")
    .beforeToolCallback(List.of(normaliseCurrency, requireOrderOwnership))  // evaluated in order
    .afterModelCallback(stripInternalIds)
    .build();

Here normaliseCurrency rewrites the argument map in place and returns empty, so requireOrderOwnership sees the normalised amount; if ownership fails it returns a refusal map and the tool never runs. Both are skipped entirely if a plugin's beforeToolCallback answered first.

Mutate or answer: two ways to compose

There are therefore two ways a hook can influence behaviour, and they compose very differently. A before-model hook receives the mutable LlmRequest.Builder. If it appends an instruction, adjusts the GenerateContentConfig or trims contents and then returns empty, the next hook sees the change and the chain continues. Mutations compose. Returning a value, by contrast, answers the hook point for everyone registered after you.

Here is the bug this produces. A team registers two plugins: a RedactionPlugin that returns a copy of the model response with card numbers masked when it finds any, and a UsageMeterPlugin that reads token counts from every response and returns empty. Registered in that order, the meter never runs on a response that contained a card number, because the redaction plugin's non-empty Maybe ended the chain. Billing undercounts exactly the conversations involving payments, and nothing errors. Reversing the order fixes metering, but now a third transformer added later, say a citation formatter, will either starve or be starved by the redactor.

The general rule: observers return empty and go first; at most one hook per hook point should return a transformed value; if you need several transformations, put them inside one plugin that runs them as a pipeline.

Building a real pipeline on top

A small pipeline plugin restores onion-like composition where you need it. Each stage is a function from response to optional replacement; the plugin threads the current value through every stage and returns a value only if something changed.

public final class ResponsePipelinePlugin extends BasePlugin {
  /** A stage returns a replacement, or empty to leave the response unchanged. */
  public interface Stage { Optional<LlmResponse> apply(CallbackContext ctx, LlmResponse in); }

  private final List<Stage> stages;

  public ResponsePipelinePlugin(List<Stage> stages) {
    super("response-pipeline");
    this.stages = List.copyOf(stages);
  }

  @Override
  public Maybe<LlmResponse> afterModelCallback(CallbackContext ctx, LlmResponse response) {
    LlmResponse current = response;
    for (Stage stage : stages) {
      current = stage.apply(ctx, current).orElse(current);   // each stage sees the previous output
    }
    return current == response ? Maybe.empty() : Maybe.just(current);
  }
}

Runner runner = Runner.builder()
    .agent(rootAgent)
    .appName("support")
    .plugins(
        new TracingPlugin(),                 // observers first: always return empty
        new UsageMeterPlugin(meter),
        new AdmissionPlugin(policy),         // may short-circuit with a refusal
        new ResponseCachePlugin(cache),      // may short-circuit with a cached answer
        new ResponsePipelinePlugin(List.of(redactCards, formatCitations)))  // the only transformer
    .build();

Registration order is now the policy, so make it reviewable: build the list in one place, and add a unit test that asserts it. Plugin names must be unique; registering a duplicate throws IllegalArgumentException. Also note that Runner.Builder rejects plugins(...) when an App has been set; with an App, plugins come from the app definition instead.

Timing without an around-hook

Middleware frameworks make timing trivial because one layer wraps the call. ADK has no around-hook, so a timer is a before and after pair that must find each other. Key the start time by invocation, branch and agent, and clear it on both the success and the error path.

public final class ModelLatencyPlugin extends BasePlugin {
  private final ConcurrentHashMap<String, Long> starts = new ConcurrentHashMap<>();
  private final Timer timer;

  public ModelLatencyPlugin(Timer timer) { super("model-latency"); this.timer = timer; }

  private static String key(CallbackContext c) {
    return c.invocationId() + "|" + c.branch().orElse("") + "|" + c.agentName();
  }

  @Override
  public Maybe<LlmResponse> beforeModelCallback(CallbackContext c, LlmRequest.Builder req) {
    starts.put(key(c), System.nanoTime());
    return Maybe.empty();
  }

  @Override
  public Maybe<LlmResponse> afterModelCallback(CallbackContext c, LlmResponse r) {
    Long t0 = starts.remove(key(c));
    if (t0 != null) timer.record(System.nanoTime() - t0, TimeUnit.NANOSECONDS);
    return Maybe.empty();
  }

  @Override
  public Maybe<LlmResponse> onModelErrorCallback(CallbackContext c, LlmRequest.Builder req, Throwable e) {
    starts.remove(key(c));                    // without this the map grows on every failure
    return Maybe.empty();
  }
}

Register timing plugins first. If a cache or admission plugin earlier in the list answers the before-model hook, your timer never starts, which is correct: no model call happened. Streaming responses can reach after-model hooks more than once per call, so decide whether you are measuring time to first chunk or time to completion and check partial() accordingly. As a backstop, evict entries older than your longest timeout in afterRunCallback.

Errors inside hooks

The PluginManager does not catch exceptions from your hooks. It logs the plugin and callback name and lets the error propagate, which fails the invocation. A null pointer in a metrics plugin is therefore an outage for every agent behind that Runner. Decide for each hook whether it fails open or closed, and make that explicit:

static Maybe<LlmResponse> failOpen(String name, Supplier<Maybe<LlmResponse>> hook) {
  return Maybe.defer(hook::get)
      .onErrorResumeNext(e -> { log.warn("hook {} failed open", name, e); return Maybe.empty(); });
}
// Observers and transformers: failOpen. Admission, authorization and redaction: fail closed
// (let the error propagate, or return an explicit refusal response).

close() is the exception: the manager closes every plugin and reports errors afterwards, so one failing close does not leak the others' resources.

Testing the chain

Test the chain, not just each hook. A recording plugin that appends its name to a shared list, a scripted fake model and three assertions catch most ordering bugs: which hooks ran, in what order, and whether the agent callbacks ran at all.

@Test
void cacheHitSkipsMeterAndAgentCallbacks() {
  List<String> trace = new CopyOnWriteArrayList<>();
  Runner runner = Runner.builder().agent(agentWithRecordingCallback(trace)).appName("t")
      .plugins(new Recording("trace", trace), new AlwaysHitCache(), new Recording("meter", trace))
      .build();
  run(runner, "hello");
  assertEquals(List.of("trace:beforeModel"), trace);   // meter and agent callback never ran
}

A fake model makes this deterministic; see testing custom LLMs.

Failure modes

  • Starved hooks. A transformer returns a value and every later plugin and every agent callback is skipped. Audit which hooks can return non-empty.
  • Assumed piping. Two after-model hooks each edit the original response; only the first edit survives.
  • Leaking correlation state. Before and after pairs that forget the error hook grow without bound.
  • Hook exceptions as outages. An observer throws and the invocation fails.
  • Global by accident. Logic meant for one agent is written as a plugin and starts rewriting requests for sub-agents too. Check agentName() or move it to an agent callback.
  • Blocking in hooks. Synchronous database calls inside a hook block the flow's thread; return a deferred Maybe backed by an async client.
  • Eager work in the returned Maybe. Building an expensive Maybe eagerly, outside Maybe.defer or fromCallable, does the work even when an earlier hook has already answered. Keep side effects inside the deferred body.
  • Shared mutable state across invocations. A plugin is one instance per Runner and serves concurrent invocations. Any field it mutates must be thread-safe and keyed by invocation, or two users' requests will read each other's values.

Trade-offs

MechanismScopeUse it for
Agent callbackOne agentAgent-specific guards, state shaping, tool argument fixes
PluginEvery agent and tool in the RunnerTracing, metering, policy, caching, redaction
Pipeline pluginOne hook point, ordered stagesSeveral transformations of the same object
BaseLlm decoratorEvery call to one modelRetries, hedging, provider-specific headers
Custom agentControl flowAnything that changes which agent runs next

Policy enforcement built from these pieces is covered in governance and guardrails.

What to do next

  1. List every plugin and callback you run and mark each as observer (always empty), short-circuiter or transformer.
  2. Reorder so observers come first, then admission and caching, then a single transformer per hook point.
  3. Collapse multiple transformers into one pipeline plugin with an explicit stage order.
  4. Add the error-path cleanup to every before and after pair, plus an eviction backstop.
  5. Wrap each hook in an explicit fail-open or fail-closed policy.
  6. Write a chain-order test with recording plugins and a fake model, and run it in CI.
Key takeaway: ADK for Java has no middleware chain; it has hooks that run as first responders. Plugins run before agent callbacks, both run in order, and the first hook to return a value ends the chain, while mutations of the request builder compose. Put observers first, allow one transformer per hook point, build pipelines explicitly when you need them, clean up on the error path, and test the order.