Most server frameworks teach middleware as an onion: each layer receives a request and a next handle, may change the request, calls next, may change the response and returns it. It is tempting to look for that shape in ADK for Java, and some write-ups even show an AgentMiddleware interface with an AgentChain. Neither type exists in ADK. What ADK has instead are hooks: per-agent callbacks registered on the agent builder, and plugins registered once on the Runner that apply to every agent, model call and tool call in it.
Hooks can do everything middleware does, but they compose by different rules, and the differences cause real bugs: metering that silently undercounts, redaction that runs on some responses but not others, timers that leak memory on the error path. This article reads the composition rules from the ADK source, shows how to build a genuine middleware pipeline on top of them, and gives an ordering discipline for a production plugin stack. The plugin contract itself is introduced in extending the runtime.
The hook points
ADK exposes hooks at four levels: the run, the agent, the model call and the tool call. Plugins see all of them; agent callbacks see the last three for their own agent. Every intercepting hook returns an RxJava Maybe (or an Optional for the Sync callback variants). Empty means carry on; a value means use this instead.
| Hook point | Plugin method | A returned value means |
|---|---|---|
| User message | onUserMessageCallback | Replaces the user message |
| Run start | beforeRunCallback | Halts the run with that content |
| Each event | onEventCallback | Replaces the event |
| Run end / error | afterRunCallback, onRunErrorCallback | Nothing: these return Completable |
| Agent | beforeAgentCallback, afterAgentCallback | Skips the agent / replaces its output |
| Model | beforeModelCallback, afterModelCallback | Skips the call / replaces the response |
| Model error | onModelErrorCallback | Used as the response instead of failing |
| Tool | beforeToolCallback, afterToolCallback | Skips the tool / replaces its result map |
| Tool error | onToolErrorCallback | Used as the result instead of failing |
Signatures differ between the two families, so copy them from the source rather than from memory. A plugin's beforeToolCallback(BaseTool, Map<String,Object>, ToolContext) has no InvocationContext parameter, while the agent callback type Callbacks.BeforeToolCallback receives one first. A plugin's afterToolCallback gets the result as a Map<String,Object>; the agent callback gets an Object. A plugin's onModelErrorCallback receives the LlmRequest.Builder; the agent variant receives the built LlmRequest.
How a chain is evaluated
Three rules, all visible in the source, decide what runs. First, the PluginManager runs plugins in registration order with concatMapMaybe and takes firstElement(): the first plugin to return a value ends the chain, and later plugins are never called for that hook. Second, plugins run before agent callbacks. In BaseLlmFlow the agent's callbacks are attached with pluginResult.switchIfEmpty(...), so if any plugin answered, no agent callback runs; BaseAgent does the same for agent-level hooks by putting the plugin callback at the front of the list. Third, inside an agent's own callback list the same first-non-empty rule applies.
// Shape of BaseLlmFlow.handleBeforeModelCallback, simplified
Maybe<LlmResponse> pluginResult =
context.pluginManager().beforeModelCallback(callbackContext, llmRequestBuilder);
Maybe<LlmResponse> agentResult = Maybe.defer(() ->
Flowable.fromIterable(agent.canonicalBeforeModelCallbacks())
.concatMapMaybe(cb -> cb.call(callbackContext, llmRequestBuilder))
.firstElement());
return pluginResult.switchIfEmpty(agentResult); // agent callbacks only if plugins were silentThe after-model hook follows the same pattern and ends with defaultIfEmpty(llmResponse). Note what that implies: every after-model hook receives the original response, not the output of the previous hook. There is no piping of results. Hooks are a chain of first responders, not an onion.
Agent callbacks are registered on the builder, either one at a time or as a list whose order is the evaluation order. The list overloads take the base marker types, so asynchronous callbacks can share a list; the Sync variants, which return Optional and are wrapped into the asynchronous form internally, are registered singly.
LlmAgent refunds = LlmAgent.builder()
.name("refunds")
.model("gemini-2.5-flash")
.instruction("Handle refund requests.")
.beforeToolCallback(List.of(normaliseCurrency, requireOrderOwnership)) // evaluated in order
.afterModelCallback(stripInternalIds)
.build();Here normaliseCurrency rewrites the argument map in place and returns empty, so requireOrderOwnership sees the normalised amount; if ownership fails it returns a refusal map and the tool never runs. Both are skipped entirely if a plugin's beforeToolCallback answered first.
Mutate or answer: two ways to compose
There are therefore two ways a hook can influence behaviour, and they compose very differently. A before-model hook receives the mutable LlmRequest.Builder. If it appends an instruction, adjusts the GenerateContentConfig or trims contents and then returns empty, the next hook sees the change and the chain continues. Mutations compose. Returning a value, by contrast, answers the hook point for everyone registered after you.
Here is the bug this produces. A team registers two plugins: a RedactionPlugin that returns a copy of the model response with card numbers masked when it finds any, and a UsageMeterPlugin that reads token counts from every response and returns empty. Registered in that order, the meter never runs on a response that contained a card number, because the redaction plugin's non-empty Maybe ended the chain. Billing undercounts exactly the conversations involving payments, and nothing errors. Reversing the order fixes metering, but now a third transformer added later, say a citation formatter, will either starve or be starved by the redactor.
The general rule: observers return empty and go first; at most one hook per hook point should return a transformed value; if you need several transformations, put them inside one plugin that runs them as a pipeline.
Building a real pipeline on top
A small pipeline plugin restores onion-like composition where you need it. Each stage is a function from response to optional replacement; the plugin threads the current value through every stage and returns a value only if something changed.
public final class ResponsePipelinePlugin extends BasePlugin {
/** A stage returns a replacement, or empty to leave the response unchanged. */
public interface Stage { Optional<LlmResponse> apply(CallbackContext ctx, LlmResponse in); }
private final List<Stage> stages;
public ResponsePipelinePlugin(List<Stage> stages) {
super("response-pipeline");
this.stages = List.copyOf(stages);
}
@Override
public Maybe<LlmResponse> afterModelCallback(CallbackContext ctx, LlmResponse response) {
LlmResponse current = response;
for (Stage stage : stages) {
current = stage.apply(ctx, current).orElse(current); // each stage sees the previous output
}
return current == response ? Maybe.empty() : Maybe.just(current);
}
}
Runner runner = Runner.builder()
.agent(rootAgent)
.appName("support")
.plugins(
new TracingPlugin(), // observers first: always return empty
new UsageMeterPlugin(meter),
new AdmissionPlugin(policy), // may short-circuit with a refusal
new ResponseCachePlugin(cache), // may short-circuit with a cached answer
new ResponsePipelinePlugin(List.of(redactCards, formatCitations))) // the only transformer
.build();Registration order is now the policy, so make it reviewable: build the list in one place, and add a unit test that asserts it. Plugin names must be unique; registering a duplicate throws IllegalArgumentException. Also note that Runner.Builder rejects plugins(...) when an App has been set; with an App, plugins come from the app definition instead.
Timing without an around-hook
Middleware frameworks make timing trivial because one layer wraps the call. ADK has no around-hook, so a timer is a before and after pair that must find each other. Key the start time by invocation, branch and agent, and clear it on both the success and the error path.
public final class ModelLatencyPlugin extends BasePlugin {
private final ConcurrentHashMap<String, Long> starts = new ConcurrentHashMap<>();
private final Timer timer;
public ModelLatencyPlugin(Timer timer) { super("model-latency"); this.timer = timer; }
private static String key(CallbackContext c) {
return c.invocationId() + "|" + c.branch().orElse("") + "|" + c.agentName();
}
@Override
public Maybe<LlmResponse> beforeModelCallback(CallbackContext c, LlmRequest.Builder req) {
starts.put(key(c), System.nanoTime());
return Maybe.empty();
}
@Override
public Maybe<LlmResponse> afterModelCallback(CallbackContext c, LlmResponse r) {
Long t0 = starts.remove(key(c));
if (t0 != null) timer.record(System.nanoTime() - t0, TimeUnit.NANOSECONDS);
return Maybe.empty();
}
@Override
public Maybe<LlmResponse> onModelErrorCallback(CallbackContext c, LlmRequest.Builder req, Throwable e) {
starts.remove(key(c)); // without this the map grows on every failure
return Maybe.empty();
}
}Register timing plugins first. If a cache or admission plugin earlier in the list answers the before-model hook, your timer never starts, which is correct: no model call happened. Streaming responses can reach after-model hooks more than once per call, so decide whether you are measuring time to first chunk or time to completion and check partial() accordingly. As a backstop, evict entries older than your longest timeout in afterRunCallback.
Errors inside hooks
The PluginManager does not catch exceptions from your hooks. It logs the plugin and callback name and lets the error propagate, which fails the invocation. A null pointer in a metrics plugin is therefore an outage for every agent behind that Runner. Decide for each hook whether it fails open or closed, and make that explicit:
static Maybe<LlmResponse> failOpen(String name, Supplier<Maybe<LlmResponse>> hook) {
return Maybe.defer(hook::get)
.onErrorResumeNext(e -> { log.warn("hook {} failed open", name, e); return Maybe.empty(); });
}
// Observers and transformers: failOpen. Admission, authorization and redaction: fail closed
// (let the error propagate, or return an explicit refusal response).close() is the exception: the manager closes every plugin and reports errors afterwards, so one failing close does not leak the others' resources.
Testing the chain
Test the chain, not just each hook. A recording plugin that appends its name to a shared list, a scripted fake model and three assertions catch most ordering bugs: which hooks ran, in what order, and whether the agent callbacks ran at all.
@Test
void cacheHitSkipsMeterAndAgentCallbacks() {
List<String> trace = new CopyOnWriteArrayList<>();
Runner runner = Runner.builder().agent(agentWithRecordingCallback(trace)).appName("t")
.plugins(new Recording("trace", trace), new AlwaysHitCache(), new Recording("meter", trace))
.build();
run(runner, "hello");
assertEquals(List.of("trace:beforeModel"), trace); // meter and agent callback never ran
}A fake model makes this deterministic; see testing custom LLMs.
Failure modes
- Starved hooks. A transformer returns a value and every later plugin and every agent callback is skipped. Audit which hooks can return non-empty.
- Assumed piping. Two after-model hooks each edit the original response; only the first edit survives.
- Leaking correlation state. Before and after pairs that forget the error hook grow without bound.
- Hook exceptions as outages. An observer throws and the invocation fails.
- Global by accident. Logic meant for one agent is written as a plugin and starts rewriting requests for sub-agents too. Check
agentName()or move it to an agent callback. - Blocking in hooks. Synchronous database calls inside a hook block the flow's thread; return a deferred
Maybebacked by an async client. - Eager work in the returned Maybe. Building an expensive
Maybeeagerly, outsideMaybe.deferorfromCallable, does the work even when an earlier hook has already answered. Keep side effects inside the deferred body. - Shared mutable state across invocations. A plugin is one instance per Runner and serves concurrent invocations. Any field it mutates must be thread-safe and keyed by invocation, or two users' requests will read each other's values.
Trade-offs
| Mechanism | Scope | Use it for |
|---|---|---|
| Agent callback | One agent | Agent-specific guards, state shaping, tool argument fixes |
| Plugin | Every agent and tool in the Runner | Tracing, metering, policy, caching, redaction |
| Pipeline plugin | One hook point, ordered stages | Several transformations of the same object |
BaseLlm decorator | Every call to one model | Retries, hedging, provider-specific headers |
| Custom agent | Control flow | Anything that changes which agent runs next |
Policy enforcement built from these pieces is covered in governance and guardrails.
What to do next
- List every plugin and callback you run and mark each as observer (always empty), short-circuiter or transformer.
- Reorder so observers come first, then admission and caching, then a single transformer per hook point.
- Collapse multiple transformers into one pipeline plugin with an explicit stage order.
- Add the error-path cleanup to every before and after pair, plus an eviction backstop.
- Wrap each hook in an explicit fail-open or fail-closed policy.
- Write a chain-order test with recording plugins and a fake model, and run it in CI.