An output content filter is code that inspects what an agent is about to say and removes, rewrites or withholds the parts that must not leave: credentials, card numbers, internal hostnames, personal data the user is not entitled to, or text that breaks a rule your business has. In most stacks the question is only what to detect. In the Agent Development Kit for Java there is a second question that decides whether the filter works at all: at which hook does it run? ADK keeps three copies of every answer, the one the caller is shown, the one the session stores, and the one fed back to the model on the next call, and the hook you choose decides which of those copies the filter actually changes.

This page builds an output filter as an ADK plugin, explains exactly which copies each hook affects, adds a holdback buffer so streamed text cannot leak a secret before the filter sees all of it, and shows how to test the result. The hook behaviour described here was read from the adk-java main branch in October 2026 (Runner, BaseLlmFlow and the Plugin interface); confirm it against the release you run. For the wider guardrail picture, including input checks and tool policy, see the guardrails page linked at the end.

Three copies of every answer

Follow one model response through a turn. The model returns an LlmResponse. ADK runs the after-model callbacks on it, plugins first and then the agent's own callbacks, and builds an Event from whatever they return. The runner then appends that event to the session through the session service, but only if it is not a partial streaming chunk. After the append, the runner passes the stored event to each plugin's onEventCallback, and emits whatever that returns to the caller. On the next model call in the same session, the conversation history is rebuilt from the stored events.

So there are three copies: shown (what onEventCallback emits), stored (what appendEvent persisted) and context (what the model reads next time, which is the stored copy). A filter in the after-model hook changes all three. A filter in onEventCallback changes only the shown copy; the raw text stays in the session, reappears when the user reloads the conversation, is visible to anyone who reads the session store, and goes back into the model's context, where the model can repeat it in a later answer that the filter then has to catch again.

One model response, three copies: where each filter placement takes effectModelLlmResponse streamafterModelCallbackplugins, then agentEvent builtfrom callback resultappendEventnon-partial onlyonEventCallbackplugin may replaceCaller / UIwhat is shownSession storewhat is storedNext model callhistory = storedFilter in afterModelCallback: shown, stored and next-turn context all see the filtered text.Filter in onEventCallback: only the caller sees it; the session keeps the raw text and feeds it back.Partial (streamed) events skip appendEvent but still pass through both hooks.
The after-model hook runs before the event exists, so its result is what gets stored. onEventCallback runs after appendEvent, so its replacement reaches only the caller.

Which hook changes which copy

PlacementShownStoredNext-turn contextUse it for
Plugin afterModelCallbackFilteredFilteredFilteredContent rules: secrets, PII, policy text
Agent afterModelCallbackFilteredFilteredFilteredRules for one agent only; runs only if no plugin returned a value
Plugin onEventCallbackFilteredRawRawPresentation only: formatting, display-only masking
Gateway in front of the runnerFilteredRawRawLast-resort net; channel-specific rules

The rule that follows is short. Anything that must not be stored or repeated belongs in the after-model hook. Use onEventCallback when you genuinely want the stored copy to differ, for example masking a card number for a support agent's screen while an access-controlled audit store keeps the original, and write that decision down, because it looks like a bug to the next engineer.

Two exceptions in the current flow code matter. If a before-model callback returns a response, the model is skipped and so are the after-model callbacks, so canned replies from an input guard are not filtered; keep them static and reviewed. And runLive, the bidirectional live path, does not call the after-model callbacks at all, so a live voice or video agent needs its filter elsewhere and its own test.

Detection rules and actions

Keep detection in a plain class with no ADK types, so it can be unit-tested with strings. Each rule has a pattern, an action and an optional confirmation step that cuts false positives: a 16-digit number is only a card number if it passes the Luhn checksum. Actions form a ladder. PASS leaves the text alone. REDACT replaces just the matched span and keeps the rest of the answer useful. BLOCK replaces the whole answer, for findings where partial output is itself a leak, such as a private key header or a cloud access key.

public enum Action { PASS, REDACT, BLOCK }
public record Finding(String rule, int start, int end, Action action) {}
public record Verdict(Action action, String text, List<String> rules) {}

public final class OutputFilter {
  private record Rule(String name, Pattern pattern, Action action, Predicate<String> confirm) {}

  private static final List<Rule> RULES = List.of(
      new Rule("aws_access_key_id", Pattern.compile("\\b(?:AKIA|ASIA)[0-9A-Z]{16}\\b"), Action.BLOCK, m -> true),
      new Rule("private_key", Pattern.compile("-----BEGIN [A-Z ]*PRIVATE KEY-----"), Action.BLOCK, m -> true),
      new Rule("card_number", Pattern.compile("\\b\\d(?:[ -]?\\d){12,18}\\b"), Action.REDACT, OutputFilter::luhn),
      new Rule("internal_host", Pattern.compile("\\b[a-z0-9-]+\\.corp\\.example\\.com\\b"), Action.REDACT, m -> true));

  public List<Finding> scan(String text) {
    List<Finding> out = new ArrayList<>();
    for (Rule r : RULES) {
      Matcher m = r.pattern().matcher(text);
      while (m.find()) {
        if (r.confirm().test(m.group())) out.add(new Finding(r.name(), m.start(), m.end(), r.action()));
      }
    }
    out.sort(Comparator.comparingInt(Finding::start));
    return out;
  }

  /** Redacts findings that lie wholly inside [from, to) and returns that slice. */
  public String redact(String text, List<Finding> found, int from, int to) {
    StringBuilder sb = new StringBuilder();
    int pos = from;
    for (Finding f : found) {
      if (f.start() < pos || f.end() > to) continue;
      sb.append(text, pos, f.start()).append("[redacted:").append(f.rule()).append(']');
      pos = f.end();
    }
    return sb.append(text, pos, to).toString();
  }

  public Verdict apply(String text) {
    List<Finding> found = scan(text);
    List<String> rules = found.stream().map(Finding::rule).distinct().toList();
    if (found.isEmpty()) return new Verdict(Action.PASS, text, rules);
    if (found.stream().anyMatch(f -> f.action() == Action.BLOCK)) {
      return new Verdict(Action.BLOCK,
          "I can't include part of that answer. A support engineer has been notified.", rules);
    }
    return new Verdict(Action.REDACT, redact(text, found, 0, text.length()), rules);
  }

  static boolean luhn(String s) {
    String d = s.replaceAll("[ -]", "");
    int sum = 0;
    for (int i = 0; i < d.length(); i++) {
      int n = d.charAt(d.length() - 1 - i) - '0';
      if (i % 2 == 1) { n *= 2; if (n > 9) n -= 9; }
      sum += n;
    }
    return sum % 10 == 0;
  }
}

The patterns are examples: AKIA or ASIA plus 16 uppercase letters or digits is the shape of AWS access key IDs, and the hostname rule stands in for your own naming scheme. Add rules from your incidents and a secret scanner's rule set.

Streaming: the holdback buffer

With StreamingMode.SSE the model's text arrives as partial responses and the after-model hook sees each one separately, followed by one non-partial aggregate holding the whole text. A filter that judges each chunk alone misses a key split across two chunks, and one that only judges the aggregate lets every chunk reach the screen first. The holdback buffer solves both: per stream, keep the raw text so far, scan all of it, and release only text that is at least HOLDBACK characters behind the end, since anything a rule could still match there might be completed by the next chunk. If a finding straddles the release point, pull the point back to the finding's start.

public final class OutputFilterPlugin extends BasePlugin {
  private static final int HOLDBACK = 64;    // at least the longest span a rule must see whole
  private static final class Stream { final StringBuilder raw = new StringBuilder(); int released; boolean blocked; }

  private final OutputFilter filter;
  private final FilterMetrics metrics;
  private final Map<String, Stream> streams = new ConcurrentHashMap<>();

  public OutputFilterPlugin(OutputFilter filter, FilterMetrics metrics) {
    super("output_filter");
    this.filter = filter; this.metrics = metrics;
  }

  @Override
  public Maybe<LlmResponse> afterModelCallback(CallbackContext ctx, LlmResponse resp) {
    List<Part> parts = resp.content().flatMap(Content::parts).orElse(List.of());
    if (parts.stream().anyMatch(p -> p.functionCall().isPresent())) return Maybe.empty();
    String key = ctx.invocationId() + "/" + ctx.agentName();
    if (resp.partial().orElse(false)) return Maybe.just(holdback(key, textOf(parts), resp));

    streams.remove(key);                          // the aggregate is the record
    Verdict v = filter.apply(textOf(parts));
    metrics.record(ctx.agentName(), v);
    return v.action() == Action.PASS ? Maybe.empty() : Maybe.just(withText(resp, v.text()));
  }

  private LlmResponse holdback(String key, String chunk, LlmResponse resp) {
    Stream s = streams.computeIfAbsent(key, k -> new Stream());
    synchronized (s) {
      if (s.blocked) return withText(resp, "");
      s.raw.append(chunk);
      String raw = s.raw.toString();
      List<Finding> found = filter.scan(raw);
      if (found.stream().anyMatch(f -> f.action() == Action.BLOCK)) {
        s.blocked = true;                         // stop the stream; the aggregate decides the record
        return withText(resp, "");
      }
      int end = Math.max(s.released, raw.length() - HOLDBACK);
      for (Finding f : found) {                   // never release half of a finding
        if (f.start() < end && f.end() > end) end = Math.max(s.released, f.start());
      }
      String out = filter.redact(raw, found, s.released, end);
      s.released = end;
      return withText(resp, out);
    }
  }

  @Override
  public Completable afterRunCallback(InvocationContext ctx) {
    streams.keySet().removeIf(k -> k.startsWith(ctx.invocationId() + "/"));   // runs that errored
    return Completable.complete();
  }

  private static LlmResponse withText(LlmResponse resp, String text) {
    // toBuilder() keeps the partial flag and usage metadata; check it exists in your version.
    return resp.toBuilder().content(Content.fromParts(Part.fromText(text))).build();
  }
}

Three details make this correct. The aggregate is filtered in full, independently of what the partials did, and it is the only copy stored, so the record is right even if the holdback logic has a bug. The stream key includes the agent name, because a sub-agent's chunks can follow the coordinator's inside one invocation, and the aggregate removes the key, so the next model call in a tool loop starts a fresh buffer. Finally, afterRunCallback clears buffers from runs that ended in an error before an aggregate arrived, which would otherwise leak memory slowly.

The cost is latency: nothing is visible until HOLDBACK characters exist, typically well under a second. Clients that replace their draft with the final message need no change.

Worked example: a key split across two chunks

A support agent answers a question about a failed deploy. The model streams two chunks: the first ends with the text deploy key AKIA, and the second begins IOSFODNN7EXAMPLE, AWS's documented sample key. After chunk one, the scan finds no rule match: AKIA alone is not a key. The release point sits 64 characters behind the end, so the tail containing AKIA stays held and the user sees only the opening words. After chunk two, the buffer contains the full 20-character key, the BLOCK rule fires, the stream is marked blocked and every later partial is released as empty text.

Then the aggregate arrives. The filter runs on the whole text, finds the key again and returns the fallback message. ADK builds the event from that message, appends it to the session and emits it; the client replaces its draft with it. The aws_access_key_id metric increments. Had the rule lived in onEventCallback, the session would hold the key, a reload would show it, and the next deploy question would hand it back to the model.

Testing the filter

Unit-test OutputFilter with plain strings, including near misses such as card-like numbers that fail Luhn. Then test the plugin through a real Runner with a stub model. A stub BaseLlm must emit both the partial chunks and the aggregate itself, because the aggregation that Gemini's streaming client performs does not happen for a stub, and the test should split the secret across the chunk boundary, since that is the case the holdback exists for.

@Test
void keySplitAcrossChunksReachesNeitherClientNorSession() {
  // A stub BaseLlm must emit the aggregate itself; Gemini's aggregator is not involved.
  BaseLlm stub = new ScriptedLlm(
      partial("Rotated the deploy key AKIA"),
      partial("IOSFODNN7EXAMPLE this morning."),
      aggregate("Rotated the deploy key AKIAIOSFODNN7EXAMPLE this morning."));
  Runner runner = runnerWith(stub, new OutputFilterPlugin(new OutputFilter(), metrics));

  List<Event> events = runner.runAsync(USER, sessionId, userMessage("status?"), sseConfig())
      .toList().blockingGet();

  assertThat(allText(events)).doesNotContain("AKIAIOSFODNN7EXAMPLE");
  assertThat(allText(storedEvents(sessionId))).doesNotContain("AKIA");
  assertThat(metrics.count("aws_access_key_id")).isEqualTo(1);
}

The second assertion is the one most teams forget: read the stored events back from the session service and check them, not just the stream.

Failure modes

  • Filtered on screen, raw in the session. The filter sits in onEventCallback or a gateway, and the model repeats the leaked value two turns later. Move content rules to the after-model hook.
  • Split secrets. Per-chunk scanning without a buffer misses values that span chunks.
  • Function calls rewritten. A filter that replaces a function-call response with text breaks the tool loop; skip responses carrying function calls, and use the tool hooks for arguments.
  • Bypassed paths. Canned before-model replies and runLive skip the after-model hook.
  • Over-redaction. A broad email rule removes the user's own address from a confirmation. Allow values the session already proved belong to the user, read from state written by a tool.
  • A throwing filter. Exceptions in plugin callbacks propagate and fail the run. That is fail-closed, which suits a leak filter, but a regex bug is then an outage, so test patterns against long and adversarial inputs.

Trade-offs

ChoiceBenefitCost
BLOCK versus REDACTBLOCK leaks nothing partialLoses the useful rest of the answer
Larger holdbackCatches longer patterns wholeLater first visible text
Disable streaming for sensitive agentsNo partial text ever escapesWorse perceived latency
Regex and checksum rulesFast, predictable, testableMiss paraphrased or novel leaks
Classifier model in the hookCatches fuzzy policy violationsLatency on every response, its own errors

What to do next

  1. Find every output filter you run today and note which hook it uses; move any content rule that sits in onEventCallback or a gateway into afterModelCallback.
  2. Put detection in a plain class with unit tests and measure each rule's precision on a sample of real responses.
  3. For streaming agents, add the holdback buffer and a test that splits a secret across chunks.
  4. Add an assertion that reads stored session events back and checks them.
  5. List agents that use runLive and decide how their output is filtered.
  6. Read ADK Java guardrails for input and tool layers, ADK Java streaming events for partial and aggregate events, output guardrails for detection pipelines and thresholds, and secret scanning in AI outputs for the rule set.
Key takeaway: In ADK Java an answer exists as three copies: what the caller is shown, what the session stores, and what the model reads on the next call. A filter in afterModelCallback runs before the event is built, so it changes all three. A filter in onEventCallback runs after the event is persisted, so it changes only what the caller sees, and the raw text comes back on reload and in the model's context. Put content rules in the after-model hook, keep detection in a plain tested class, use a holdback buffer for streamed partial responses, remember that canned before-model replies and runLive skip the after-model hook, and test the stored session as well as the stream.