An ADK Java agent has two places to put cross-cutting code. Callbacks belong to one agent: you pass them to LlmAgent.builder() and they fire only for that agent. Plugins belong to the application: you register them once on the Runner or App, and every agent, model call and tool call in every invocation passes through them. That difference is the whole reason plugins exist. Logging, token budgets, policy, analytics export and context trimming are properties of the deployment, not of one agent, and writing them as per-agent callbacks means copying them into every agent and forgetting one.

This article reads the plugin layer as it ships in ADK Java 1.11.0, checked against the compiled classes rather than assumed from the Python docs. It covers the contract, how the plugin manager runs the hooks, the run-level hooks that only plugins have, the plugins that ship in the jar, how to keep per-invocation state safely, and how to test and package a plugin. For how a plugin chain composes with agent callbacks hook by hook, read Agent Hooks and the Middleware Pattern; this page is about the plugin as a unit you build and operate.

The contract

The contract is the interface com.google.adk.plugins.Plugin. It has one abstract method, getName(), and fourteen default methods that do nothing. BasePlugin is the abstract class you usually extend; its constructor takes the name. Every hook returns an RxJava type, and the type tells you what the hook can do:

HookReturnsA value means
onUserMessageCallbackMaybe<Content>replace the incoming user message
beforeRunCallbackMaybe<Content>answer now and end the run
onEventCallbackMaybe<Event>replace the event the caller receives
afterRunCallbackCompletableside effect only
onRunErrorCallbackCompletableside effect only
closeCompletablerelease resources at shutdown
beforeAgentCallbackMaybe<Content>skip the agent, use this content
afterAgentCallbackMaybe<Content>replace the agent's output
beforeModelCallbackMaybe<LlmResponse>skip the model call
afterModelCallbackMaybe<LlmResponse>replace the model response
onModelErrorCallbackMaybe<LlmResponse>recover from a model error
beforeToolCallbackMaybe<Map>skip the tool, use this result
afterToolCallbackMaybe<Map>replace the tool result
onToolErrorCallbackMaybe<Map>recover from a tool error

Two rules follow from these types. First, Maybe.empty() is how a plugin says "carry on", and it is the right answer for anything that only observes. Second, the before-model hook receives a mutable LlmRequest.Builder, so a plugin can rewrite the request in place and still return empty. That is how the shipped instruction and context plugins work: they edit, they do not answer.

Registration and identity

Plugins attach at the top of the application. Both builders accept them:

App app = App.builder()
    .name("support")
    .rootAgent(rootAgent)
    .plugins(new TokenBudgetPlugin(40_000), new LoggingPlugin())
    .build();

Runner runner = Runner.builder()
    .app(app)
    .sessionService(sessionService)
    .build();

// or directly on the runner
Runner direct = Runner.builder()
    .agent(rootAgent)
    .appName("support")
    .plugins(List.of(new TokenBudgetPlugin(40_000)))
    .build();

The runner wraps them in a PluginManager, which is itself a plugin and is exposed by runner.pluginManager(). Registration enforces one rule: names are unique. Registering a second plugin with a name already present throws IllegalArgumentException at construction, which is what you want, because getPlugin(name) returns the first match and two plugins called "audit" would make lookups ambiguous. If you need two instances of one class, for example two context filters with different settings, give them distinct names through the constructor or builder.

There is one manager per runner and one instance of each plugin. Every concurrent invocation calls the same object, so a plugin is a shared, long-lived component and must be thread-safe. That is the first design constraint, and most plugin bugs come from forgetting it.

How the plugin manager runs hooks

The manager runs every Maybe hook the same way: it iterates the plugins in registration order with concatMapMaybe and takes firstElement(). The first plugin to return a value wins and the rest are not called. Each call is wrapped with doOnError, which logs [name] Error during callback and lets the error continue, so a plugin that throws fails the operation it was guarding. Nothing in the manager swallows errors for you.

Then the call sites decide what happens next. In BaseAgent, BaseLlmFlow and Functions the plugin manager is asked first, and the agent's own callbacks are reached only through switchIfEmpty. So a plugin that returns a value from afterModelCallback silently skips every agent-level afterModelCallback. This is the behaviour to design around: a plugin that answers takes precedence over every agent below it.

The completable hooks differ. afterRunCallback and onRunErrorCallback run plugins in order with concatMapCompletable, so the first failure stops the rest; a flaky exporter registered first can prevent a cleanup plugin registered after it from running. close() uses concatMapCompletableDelayError, so every plugin's close runs and errors are reported at the end.

Where plugins sit in one invocationCallerrunner.runAsyncRunner + PluginManageronUserMessage, beforeRun, onEvent, afterRunmessageAgent hooksplugins, then agent callbacksModel hooksbefore / after / onErrorTool hooksbefore / after / onErrorSession serviceappendEventeventOne instanceshared by every agentregistration order = run orderFirst plugin to return a value wins; agent callbacks run only if every plugin returned empty.onEvent sees an event after it is stored; its replacement reaches the caller, not the session.
One plugin manager per runner; run-level hooks live on the runner, agent, model and tool hooks in the flow, and plugins are consulted before agent callbacks at every level.

Run-level hooks only plugins have

The run-level hooks exist only on plugins, because only the runner can call them.

  • onUserMessageCallback: the runner takes its result with defaultIfEmpty(original), so returning a Content replaces the user message before anything sees it. Use it for normalisation or for stripping data you must not store, keeping in mind that the replaced message is what lands in the session.
  • beforeRunCallback: a returned Content is turned into an event and the agent never runs. This is the place for a maintenance switch, a per-user rate limit or a tenant that is suspended.
  • onEventCallback: for a non-partial event the runner calls appendEvent on the session service first and offers the event to plugins afterwards. A replacement event reaches the caller's stream but not the stored session. That makes it the right hook for shaping what a client sees, such as removing internal metadata, and the wrong hook for redacting what you persist.
  • afterRunCallback and onRunErrorCallback: the end of an invocation, successful or not. These are where per-invocation state is released and buffered telemetry is flushed.

The plugins that ship with ADK Java

The 1.11.0 jar ships four plugins. Read them before writing your own, because two of them are templates for the most common shapes.

PluginHooksWhat it does
LoggingPlugin12 of 14logs every hook except onRunError and close at INFO through SLF4J, default name logging_plugin; prints text parts, function calls and responses
GlobalInstructionPluginbefore modeladds an instruction, fixed or computed from the callback context, to the system instruction of every model call
ContextFilterPluginbefore modeltrims request contents; builder takes numInvocationsToKeep, customFilter and name
BigQueryAgentAnalyticsPluginmanybatches event rows to BigQuery with its own retry config and optional GCS offload

LoggingPlugin is a debugging aid, not a production logger: it writes message text and tool arguments at INFO, which is a privacy problem in most deployments. GlobalInstructionPlugin shows the edit-in-place pattern, and ContextFilterPlugin shows how to change what the model sees without touching the stored session.

Building a plugin with state

A useful plugin with real state: cap the tokens one invocation may spend. The state is keyed by invocationId(), which every callback context carries, held in a concurrent map because invocations run in parallel, and released at the end of the run.

public final class TokenBudgetPlugin extends BasePlugin {
  private final long limit;
  private final ConcurrentHashMap<String, AtomicLong> spent = new ConcurrentHashMap<>();

  public TokenBudgetPlugin(long limit) {
    super("token_budget");
    this.limit = limit;
  }

  @Override
  public Maybe<LlmResponse> beforeModelCallback(CallbackContext ctx, LlmRequest.Builder req) {
    AtomicLong used = spent.get(ctx.invocationId());
    if (used == null || used.get() < limit) {
      return Maybe.empty();                       // carry on
    }
    return Maybe.just(LlmResponse.builder()       // answer instead of calling the model
        .content(Content.fromParts(Part.fromText(
            "This request reached its processing budget. Please narrow the question.")))
        .build());
  }

  @Override
  public Maybe<LlmResponse> afterModelCallback(CallbackContext ctx, LlmResponse resp) {
    resp.usageMetadata()
        .flatMap(u -> u.totalTokenCount())
        .ifPresent(n -> spent.computeIfAbsent(ctx.invocationId(), k -> new AtomicLong())
                             .addAndGet(n));
    return Maybe.empty();                         // observe only: agent callbacks still run
  }

  @Override
  public Completable afterRunCallback(InvocationContext ctx) {
    return Completable.fromAction(() -> spent.remove(ctx.invocationId()));
  }

  @Override
  public Completable onRunErrorCallback(InvocationContext ctx, Throwable error) {
    return Completable.fromAction(() -> spent.remove(ctx.invocationId()));
  }
}

Three choices matter. The after-model hook returns empty, so it never displaces an agent's own after-model callbacks. Cleanup is in both end hooks, because a run ends one way or the other. And the budget answer is an ordinary model response, so the flow, the session and the caller all handle it without special cases. In production add a periodic sweep that drops entries older than your longest plausible invocation, because a cancelled subscription or a cleanup plugin stopped by an earlier failure leaves entries behind.

Worked example: a budget that trips mid-turn

Walk one invocation with a 40,000-token limit. The user asks for a reconciliation across three accounts. The root agent's first model call reports 9,200 total tokens; the plugin records 9,200. The model asks for a ledger lookup, and the follow-up call, which resends the history plus the tool result, reports 14,100: total 23,300. A second lookup and follow-up report 16,800: total 40,100. That call started at 23,300, under the limit, so it ran and crossed the line. The model asks for a third lookup; the tool runs, and the model call that would read its result finds 40,100 already spent, so the plugin answers it. The agent produces the budget message as its final response and afterRunCallback removes the entry.

Two lessons fall out. A before-call check bounds the number of calls past the limit to one, not the tokens: if a single call can be large, also cap the request size with ContextFilterPlugin or a max-output setting. And total tokens grow roughly quadratically with tool rounds because each round resends the history, which is why budgets set from single-call measurements are always too generous.

Testing and packaging

Test a plugin at two levels. Unit-test the hooks directly: they are plain methods returning RxJava types, so plugin.beforeModelCallback(ctx, builder).test() gives you a TestObserver to assert empty or a value. Then run one end-to-end test through a real Runner with an in-memory session service and a fake model that returns scripted responses with usage metadata. The end-to-end test is the one that catches ordering mistakes: register your plugin alongside the others in production order and assert both the final event and that agent callbacks you expect did run.

Package plugins as a small library with no dependency on your agents. Configuration comes in through the constructor, the plugin owns its name, and the application decides order. Publish a short note with each plugin listing which hooks it answers in, because that is what its users need to place it correctly. ADK Java Audit Logging is a complete worked plugin built this way.

Failure modes

  • Answering where you meant to observe. A metrics plugin that returns the response it received from afterModelCallback instead of empty disables every agent-level after-model callback. Observers return empty.
  • Order inversion. A cache plugin registered before a policy plugin answers from cache before policy runs. Put deny-capable plugins first, as ADK Java Guardrails lays out.
  • State leaks. Per-invocation maps with no cleanup grow until the heap fills, or keep growing when an earlier end-hook failure stops yours from running.
  • Redacting in the wrong hook. Changing an event in onEventCallback does not change the stored session; redact before the data is produced.
  • Blocking I/O in a hook. Hooks run on the invocation's thread; a slow synchronous export adds its latency to every call it observes.
  • LoggingPlugin in production. Message text at INFO ends up in log storage with weaker access controls than your session store; see ADK Java observability for a metrics-and-spans plugin instead.

Trade-offs

ChoiceGainCost
Plugin instead of agent callbackapplies everywhere, cannot be forgottenapplies everywhere, including agents that need an exception
Answering hookcan block or short-circuithides agent callbacks below it
Edit the request builderno precedence effectsevery later hook sees the edited request
State in the pluginsimple, fastper-process only; lost on restart, not shared across replicas
Async export in afterRunno added latencydata loss on crash unless buffered durably

What to do next

  1. List every per-agent callback you have copied into more than one agent; each is a plugin candidate.
  2. Write down the intended plugin order with a one-line reason for each position.
  3. Make every observing hook return Maybe.empty() and add a test proving agent callbacks still run.
  4. Key per-invocation state by invocationId(), clean it in both end hooks and add a sweep.
  5. Replace LoggingPlugin in production with a structured logger that omits content.
  6. Run one end-to-end test through a real runner with production plugin order.
Key takeaway: A plugin is one shared, thread-safe object that every agent, model call and tool call passes through. Plugins run in registration order, the first one to return a value wins and agent callbacks run only when every plugin returned empty. Observers return empty, deny-capable plugins go first, per-invocation state is keyed by invocation id and released in both end hooks, and redaction happens before data is stored, not in onEvent.