An output content filter is code that inspects what an agent is about to say and removes, rewrites or withholds the parts that must not leave: credentials, card numbers, internal hostnames, personal data the user is not entitled to, or text that breaks a rule your business has. In most stacks the question is only what to detect. In the Agent Development Kit for Java there is a second question that decides whether the filter works at all: at which hook does it run? ADK keeps three copies of every answer, the one the caller is shown, the one the session stores, and the one fed back to the model on the next call, and the hook you choose decides which of those copies the filter actually changes.
This page builds an output filter as an ADK plugin, explains exactly which copies each hook affects, adds a holdback buffer so streamed text cannot leak a secret before the filter sees all of it, and shows how to test the result. The hook behaviour described here was read from the adk-java main branch in October 2026 (Runner, BaseLlmFlow and the Plugin interface); confirm it against the release you run. For the wider guardrail picture, including input checks and tool policy, see the guardrails page linked at the end.
Three copies of every answer
Follow one model response through a turn. The model returns an LlmResponse. ADK runs the after-model callbacks on it, plugins first and then the agent's own callbacks, and builds an Event from whatever they return. The runner then appends that event to the session through the session service, but only if it is not a partial streaming chunk. After the append, the runner passes the stored event to each plugin's onEventCallback, and emits whatever that returns to the caller. On the next model call in the same session, the conversation history is rebuilt from the stored events.
So there are three copies: shown (what onEventCallback emits), stored (what appendEvent persisted) and context (what the model reads next time, which is the stored copy). A filter in the after-model hook changes all three. A filter in onEventCallback changes only the shown copy; the raw text stays in the session, reappears when the user reloads the conversation, is visible to anyone who reads the session store, and goes back into the model's context, where the model can repeat it in a later answer that the filter then has to catch again.
Which hook changes which copy
| Placement | Shown | Stored | Next-turn context | Use it for |
|---|---|---|---|---|
Plugin afterModelCallback | Filtered | Filtered | Filtered | Content rules: secrets, PII, policy text |
Agent afterModelCallback | Filtered | Filtered | Filtered | Rules for one agent only; runs only if no plugin returned a value |
Plugin onEventCallback | Filtered | Raw | Raw | Presentation only: formatting, display-only masking |
| Gateway in front of the runner | Filtered | Raw | Raw | Last-resort net; channel-specific rules |
The rule that follows is short. Anything that must not be stored or repeated belongs in the after-model hook. Use onEventCallback when you genuinely want the stored copy to differ, for example masking a card number for a support agent's screen while an access-controlled audit store keeps the original, and write that decision down, because it looks like a bug to the next engineer.
Two exceptions in the current flow code matter. If a before-model callback returns a response, the model is skipped and so are the after-model callbacks, so canned replies from an input guard are not filtered; keep them static and reviewed. And runLive, the bidirectional live path, does not call the after-model callbacks at all, so a live voice or video agent needs its filter elsewhere and its own test.
Detection rules and actions
Keep detection in a plain class with no ADK types, so it can be unit-tested with strings. Each rule has a pattern, an action and an optional confirmation step that cuts false positives: a 16-digit number is only a card number if it passes the Luhn checksum. Actions form a ladder. PASS leaves the text alone. REDACT replaces just the matched span and keeps the rest of the answer useful. BLOCK replaces the whole answer, for findings where partial output is itself a leak, such as a private key header or a cloud access key.
public enum Action { PASS, REDACT, BLOCK }
public record Finding(String rule, int start, int end, Action action) {}
public record Verdict(Action action, String text, List<String> rules) {}
public final class OutputFilter {
private record Rule(String name, Pattern pattern, Action action, Predicate<String> confirm) {}
private static final List<Rule> RULES = List.of(
new Rule("aws_access_key_id", Pattern.compile("\\b(?:AKIA|ASIA)[0-9A-Z]{16}\\b"), Action.BLOCK, m -> true),
new Rule("private_key", Pattern.compile("-----BEGIN [A-Z ]*PRIVATE KEY-----"), Action.BLOCK, m -> true),
new Rule("card_number", Pattern.compile("\\b\\d(?:[ -]?\\d){12,18}\\b"), Action.REDACT, OutputFilter::luhn),
new Rule("internal_host", Pattern.compile("\\b[a-z0-9-]+\\.corp\\.example\\.com\\b"), Action.REDACT, m -> true));
public List<Finding> scan(String text) {
List<Finding> out = new ArrayList<>();
for (Rule r : RULES) {
Matcher m = r.pattern().matcher(text);
while (m.find()) {
if (r.confirm().test(m.group())) out.add(new Finding(r.name(), m.start(), m.end(), r.action()));
}
}
out.sort(Comparator.comparingInt(Finding::start));
return out;
}
/** Redacts findings that lie wholly inside [from, to) and returns that slice. */
public String redact(String text, List<Finding> found, int from, int to) {
StringBuilder sb = new StringBuilder();
int pos = from;
for (Finding f : found) {
if (f.start() < pos || f.end() > to) continue;
sb.append(text, pos, f.start()).append("[redacted:").append(f.rule()).append(']');
pos = f.end();
}
return sb.append(text, pos, to).toString();
}
public Verdict apply(String text) {
List<Finding> found = scan(text);
List<String> rules = found.stream().map(Finding::rule).distinct().toList();
if (found.isEmpty()) return new Verdict(Action.PASS, text, rules);
if (found.stream().anyMatch(f -> f.action() == Action.BLOCK)) {
return new Verdict(Action.BLOCK,
"I can't include part of that answer. A support engineer has been notified.", rules);
}
return new Verdict(Action.REDACT, redact(text, found, 0, text.length()), rules);
}
static boolean luhn(String s) {
String d = s.replaceAll("[ -]", "");
int sum = 0;
for (int i = 0; i < d.length(); i++) {
int n = d.charAt(d.length() - 1 - i) - '0';
if (i % 2 == 1) { n *= 2; if (n > 9) n -= 9; }
sum += n;
}
return sum % 10 == 0;
}
}The patterns are examples: AKIA or ASIA plus 16 uppercase letters or digits is the shape of AWS access key IDs, and the hostname rule stands in for your own naming scheme. Add rules from your incidents and a secret scanner's rule set.
Streaming: the holdback buffer
With StreamingMode.SSE the model's text arrives as partial responses and the after-model hook sees each one separately, followed by one non-partial aggregate holding the whole text. A filter that judges each chunk alone misses a key split across two chunks, and one that only judges the aggregate lets every chunk reach the screen first. The holdback buffer solves both: per stream, keep the raw text so far, scan all of it, and release only text that is at least HOLDBACK characters behind the end, since anything a rule could still match there might be completed by the next chunk. If a finding straddles the release point, pull the point back to the finding's start.
public final class OutputFilterPlugin extends BasePlugin {
private static final int HOLDBACK = 64; // at least the longest span a rule must see whole
private static final class Stream { final StringBuilder raw = new StringBuilder(); int released; boolean blocked; }
private final OutputFilter filter;
private final FilterMetrics metrics;
private final Map<String, Stream> streams = new ConcurrentHashMap<>();
public OutputFilterPlugin(OutputFilter filter, FilterMetrics metrics) {
super("output_filter");
this.filter = filter; this.metrics = metrics;
}
@Override
public Maybe<LlmResponse> afterModelCallback(CallbackContext ctx, LlmResponse resp) {
List<Part> parts = resp.content().flatMap(Content::parts).orElse(List.of());
if (parts.stream().anyMatch(p -> p.functionCall().isPresent())) return Maybe.empty();
String key = ctx.invocationId() + "/" + ctx.agentName();
if (resp.partial().orElse(false)) return Maybe.just(holdback(key, textOf(parts), resp));
streams.remove(key); // the aggregate is the record
Verdict v = filter.apply(textOf(parts));
metrics.record(ctx.agentName(), v);
return v.action() == Action.PASS ? Maybe.empty() : Maybe.just(withText(resp, v.text()));
}
private LlmResponse holdback(String key, String chunk, LlmResponse resp) {
Stream s = streams.computeIfAbsent(key, k -> new Stream());
synchronized (s) {
if (s.blocked) return withText(resp, "");
s.raw.append(chunk);
String raw = s.raw.toString();
List<Finding> found = filter.scan(raw);
if (found.stream().anyMatch(f -> f.action() == Action.BLOCK)) {
s.blocked = true; // stop the stream; the aggregate decides the record
return withText(resp, "");
}
int end = Math.max(s.released, raw.length() - HOLDBACK);
for (Finding f : found) { // never release half of a finding
if (f.start() < end && f.end() > end) end = Math.max(s.released, f.start());
}
String out = filter.redact(raw, found, s.released, end);
s.released = end;
return withText(resp, out);
}
}
@Override
public Completable afterRunCallback(InvocationContext ctx) {
streams.keySet().removeIf(k -> k.startsWith(ctx.invocationId() + "/")); // runs that errored
return Completable.complete();
}
private static LlmResponse withText(LlmResponse resp, String text) {
// toBuilder() keeps the partial flag and usage metadata; check it exists in your version.
return resp.toBuilder().content(Content.fromParts(Part.fromText(text))).build();
}
}Three details make this correct. The aggregate is filtered in full, independently of what the partials did, and it is the only copy stored, so the record is right even if the holdback logic has a bug. The stream key includes the agent name, because a sub-agent's chunks can follow the coordinator's inside one invocation, and the aggregate removes the key, so the next model call in a tool loop starts a fresh buffer. Finally, afterRunCallback clears buffers from runs that ended in an error before an aggregate arrived, which would otherwise leak memory slowly.
The cost is latency: nothing is visible until HOLDBACK characters exist, typically well under a second. Clients that replace their draft with the final message need no change.
Worked example: a key split across two chunks
A support agent answers a question about a failed deploy. The model streams two chunks: the first ends with the text deploy key AKIA, and the second begins IOSFODNN7EXAMPLE, AWS's documented sample key. After chunk one, the scan finds no rule match: AKIA alone is not a key. The release point sits 64 characters behind the end, so the tail containing AKIA stays held and the user sees only the opening words. After chunk two, the buffer contains the full 20-character key, the BLOCK rule fires, the stream is marked blocked and every later partial is released as empty text.
Then the aggregate arrives. The filter runs on the whole text, finds the key again and returns the fallback message. ADK builds the event from that message, appends it to the session and emits it; the client replaces its draft with it. The aws_access_key_id metric increments. Had the rule lived in onEventCallback, the session would hold the key, a reload would show it, and the next deploy question would hand it back to the model.
Testing the filter
Unit-test OutputFilter with plain strings, including near misses such as card-like numbers that fail Luhn. Then test the plugin through a real Runner with a stub model. A stub BaseLlm must emit both the partial chunks and the aggregate itself, because the aggregation that Gemini's streaming client performs does not happen for a stub, and the test should split the secret across the chunk boundary, since that is the case the holdback exists for.
@Test
void keySplitAcrossChunksReachesNeitherClientNorSession() {
// A stub BaseLlm must emit the aggregate itself; Gemini's aggregator is not involved.
BaseLlm stub = new ScriptedLlm(
partial("Rotated the deploy key AKIA"),
partial("IOSFODNN7EXAMPLE this morning."),
aggregate("Rotated the deploy key AKIAIOSFODNN7EXAMPLE this morning."));
Runner runner = runnerWith(stub, new OutputFilterPlugin(new OutputFilter(), metrics));
List<Event> events = runner.runAsync(USER, sessionId, userMessage("status?"), sseConfig())
.toList().blockingGet();
assertThat(allText(events)).doesNotContain("AKIAIOSFODNN7EXAMPLE");
assertThat(allText(storedEvents(sessionId))).doesNotContain("AKIA");
assertThat(metrics.count("aws_access_key_id")).isEqualTo(1);
}The second assertion is the one most teams forget: read the stored events back from the session service and check them, not just the stream.
Failure modes
- Filtered on screen, raw in the session. The filter sits in
onEventCallbackor a gateway, and the model repeats the leaked value two turns later. Move content rules to the after-model hook. - Split secrets. Per-chunk scanning without a buffer misses values that span chunks.
- Function calls rewritten. A filter that replaces a function-call response with text breaks the tool loop; skip responses carrying function calls, and use the tool hooks for arguments.
- Bypassed paths. Canned before-model replies and
runLiveskip the after-model hook. - Over-redaction. A broad email rule removes the user's own address from a confirmation. Allow values the session already proved belong to the user, read from state written by a tool.
- A throwing filter. Exceptions in plugin callbacks propagate and fail the run. That is fail-closed, which suits a leak filter, but a regex bug is then an outage, so test patterns against long and adversarial inputs.
Trade-offs
| Choice | Benefit | Cost |
|---|---|---|
| BLOCK versus REDACT | BLOCK leaks nothing partial | Loses the useful rest of the answer |
| Larger holdback | Catches longer patterns whole | Later first visible text |
| Disable streaming for sensitive agents | No partial text ever escapes | Worse perceived latency |
| Regex and checksum rules | Fast, predictable, testable | Miss paraphrased or novel leaks |
| Classifier model in the hook | Catches fuzzy policy violations | Latency on every response, its own errors |
What to do next
- Find every output filter you run today and note which hook it uses; move any content rule that sits in
onEventCallbackor a gateway intoafterModelCallback. - Put detection in a plain class with unit tests and measure each rule's precision on a sample of real responses.
- For streaming agents, add the holdback buffer and a test that splits a secret across chunks.
- Add an assertion that reads stored session events back and checks them.
- List agents that use
runLiveand decide how their output is filtered. - Read ADK Java guardrails for input and tool layers, ADK Java streaming events for partial and aggregate events, output guardrails for detection pipelines and thresholds, and secret scanning in AI outputs for the rule set.