An ADK agent session is an append-only log of events: user messages, model turns, function calls and function responses. Every model call replays the relevant history. A support agent that runs for forty turns, with a few large tool responses along the way, eventually costs more per turn than it is worth, and later hits the model's context window. Context compression is how you keep a long session usable. You replace older stretches of history with a summary while keeping the recent turns verbatim.

The ideas are covered in general terms in context compaction for long-running agents and, for ADK as a whole, in ADK context compaction. This article is about the Java implementation: the com.google.adk.summarizer package, the two compactors, how to configure them on an App, how to write a summarizer that fails safely, and how to tell in production whether compression is helping. API names are taken from the ADK Java reference documentation as of October 2026. ADK Java is now in its 1.x line and the compaction docs list it from 0.2.0, so check the Javadoc for the release you run before copying a builder method.

The moving parts

Six types do the work, and it helps to know which one owns what:

TypeRole
EventsCompactionConfigA record holding compactionInterval, overlapSize, summarizer, tokenThreshold and eventRetentionSize. Built with EventsCompactionConfig.builder() and set on the App.
SlidingWindowEventCompactorCompacts once enough new user-initiated invocations have accumulated, with an overlap into the previous window.
TailRetentionEventCompactorCompacts when the prompt token count exceeds a threshold, keeping the most recent retentionSize events raw.
BaseEventSummarizerOne method, Maybe<Event> summarizeEvents(List<Event> events). An empty Maybe means no compaction happened.
LlmEventSummarizerThe provided implementation: an LLM call. Constructors take a BaseLlm and optionally a custom prompt template.
EventCompactionAttached to the summary event through EventActions.compaction(): startTimestamp, endTimestamp and compactedContent.
Where compaction sits in an ADK Java invocationUser messagenew invocationRunnerApp + root agentContents processorbuilds LLM requestModelsees summary + tailSession event log (append-only)raw events + compaction events, nothing deletedreadappendCompactorsliding window or tail retentionBaseEventSummarizerMaybe of summary EventSummarizer LLMcheap model, own prompttrigger metappend summaryA compaction event carries EventActions.compaction(): start and end timestamps plus compactedContent.When the next prompt is built, events inside that range are skipped and the summary is inserted in their place.
Compaction never deletes events. It appends a summary event covering a timestamp range, and the request builder substitutes that summary for the covered events.

The key design choice is in the last row. A compaction is itself an event in the session log. The compactors' compact(Session, BaseSessionService) method appends it through the session service. When the next prompt is built, events whose timestamps fall inside the compaction's range are left out and its compactedContent goes in their place. The raw history stays in storage. That matters for audits and debugging, and it means compaction shrinks the prompt, not the database. If storage growth is your problem, you need a retention policy on the session store as well. See ADK Java events for the rest of the EventActions contract.

Sliding-window compaction, worked through

Sliding-window compaction is the predictable option. It counts invocations, meaning one user message and everything the agent did in response, rather than tokens. This is the configuration from the ADK documentation, with a cheaper model doing the summarising:

Gemini summarizationLlm = Gemini.builder()
    .model("gemini-flash-latest")
    .build();

LlmEventSummarizer summarizer = new LlmEventSummarizer(summarizationLlm);

App app = App.builder()
    .name("support-agent")
    .rootAgent(rootAgent)
    .eventsCompactionConfig(EventsCompactionConfig.builder()
        .compactionInterval(3)   // compact after every 3 new user-initiated invocations
        .overlapSize(1)          // re-include the last invocation of the previous window
        .summarizer(summarizer)
        .build())
    .build();

Work through what happens over nine invocations with an interval of 3 and an overlap of 1. After invocation 3 completes, invocations 1 to 3 are summarised into compaction S1. After invocation 6, the new block is 4 to 6. The overlap reaches back one invocation, so S2 covers 3 to 6. After invocation 9, S3 covers 6 to 9. At turn 10 the model sees S1, S2 and S3 followed by the raw events of invocation 10. The overlap exists so that a topic spanning a window boundary appears whole in at least one summary. The price is that the boundary invocation is summarised twice.

Choose the interval by turn size. With short text turns, an interval of 5 to 10 keeps summarisation calls rare. With turns that carry large tool responses, a smaller interval stops a few heavy turns from dominating the prompt. Keep the overlap at 1 unless you see summaries losing threads across boundaries, because each extra overlapping invocation is paid for in every window.

Token-triggered compaction with a retained tail

Sliding windows ignore size: three invocations that each return a 40,000-token document are treated the same as three greetings. The token-based fields cover that case. tokenThreshold is the prompt token count above which compaction triggers, and eventRetentionSize is how many recent events are kept raw. They map onto TailRetentionEventCompactor, whose constructor takes exactly a summarizer, a retention size and a token threshold:

EventsCompactionConfig config = EventsCompactionConfig.builder()
    .summarizer(summarizer)
    .tokenThreshold(60_000)      // compact when the prompt passes about 60k tokens
    .eventRetentionSize(12)      // always keep the 12 most recent events verbatim
    .build();

Its behaviour differs from the sliding window in one important way: the summary is rolling. When it triggers, it compacts every older event not already compacted, including the most recent compaction event. The new summary therefore absorbs the previous one and supersedes it. You end up with a single summary plus a raw tail, rather than a growing list of window summaries.

The ADK documentation says that when both styles are configured, the token-based trigger takes precedence. Treat it as a safety net over a sliding window rather than a replacement for one. Choose the threshold as a fraction of the window of the model you actually call, not the summariser's model. Leave room for the response and for one more large tool result: 60 to 70 percent of the window is a reasonable start. Count retention in events, not invocations, and remember that a single tool-using invocation produces several events (call, response, final answer). A retention of 12 events can be as few as three turns. Token counting across models in ADK Java explains where trustworthy prompt token counts come from.

A summarizer that fails safely

LlmEventSummarizer is fine as a starting point, but production systems usually want three more things: a prompt that preserves domain facts, a time limit, and a guarantee that a failed summary leaves the session untouched. The interface makes that easy, because an empty Maybe means "no compaction". A decorator can add all three without building events by hand:

public final class GuardedSummarizer implements BaseEventSummarizer {
    private static final String PROMPT = """
        Summarize the conversation events below for an agent that will continue it.
        Keep verbatim: order IDs, amounts, dates, names, and every commitment the
        agent made. Record decisions and open questions. Drop greetings and
        tool payloads that were already acted on. Do not invent facts.

        {conversation_history}
        """;   // LlmEventSummarizer replaces this placeholder with the events

    private final BaseEventSummarizer delegate;
    private final Duration timeout;

    public GuardedSummarizer(BaseLlm model, Duration timeout) {
        this.delegate = new LlmEventSummarizer(model, PROMPT);
        this.timeout = timeout;
    }

    @Override
    public Maybe<Event> summarizeEvents(List<Event> events) {
        if (events.size() < 4) {
            return Maybe.empty();                 // not worth a model call
        }
        long started = System.nanoTime();         // Metrics is your own wrapper
        return delegate.summarizeEvents(events)
            .timeout(timeout.toMillis(), TimeUnit.MILLISECONDS)
            .doOnSuccess(e -> Metrics.compactionLatency(System.nanoTime() - started))
            .doOnError(Metrics::compactionFailed)
            .onErrorComplete();                   // failure = keep raw history
    }
}

The important line is onErrorComplete(). If the summariser model is slow, rate-limited or down, the compactor gets an empty result, appends nothing, and the agent keeps working on full history. The worst case is one expensive turn, not a broken session. A custom template must keep the {conversation_history} placeholder; without it the model receives instructions and no events. The prompt says what to keep as well as what to drop: generic summaries reliably lose identifiers and numbers, and those are exactly what a support or operations agent needs later.

One more design rule: facts the agent must never lose do not belong in conversation history at all. Write them to session state through the tool or callback context, as described in session context in ADK Java. State is not summarised, so an order ID stored there survives any number of compactions.

Operating it: what to measure

Compaction trades tokens for fidelity, and you only know the trade is good if you measure both sides:

SignalWhere to get itWhat it tells you
Prompt tokens per turnUsage metadata on model response eventsShould rise, drop at each compaction, and settle into a sawtooth below your threshold
Compactions per sessionCount events where actions().compaction() is presentSeveral per turn means the threshold sits below the size of a single turn
Summary lengthSize of compactedContent partsA rolling summary growing without limit is turning into a second transcript
Summariser latency and failuresThe decorator's metricsAdded time on the turns that compact; failure rate of the fallback
Task success after compactionYour eval set, replayed with compaction on and offThe only real measure of what the summary lost

Build that last evaluation on purpose. Take twenty long recorded sessions, cut each at a point after a compaction, and ask a question whose answer was only in the compacted range: the refund amount, the chosen plan, the error code. Run it with compaction on and off. If accuracy drops, fix the summary prompt or move the fact into state before you lower the thresholds any further.

Also watch how compaction interacts with prompt caching. Providers that cache a stable prompt prefix give the biggest discount when the start of the prompt does not change between turns. Each compaction rewrites that start, so the first turn after it pays full price for the whole prompt again. Sliding windows that compact every few turns can therefore save fewer tokens on the bill than the raw counts suggest. Compare billed cost, not just prompt tokens, before and after enabling compaction, and prefer fewer, larger compactions when caching matters. Finally, read the summaries. Sampling a dozen compaction events a week and comparing them with the raw events they cover catches prompt problems that aggregate metrics miss.

Failure modes

  • Summary drift. With tail retention, each summary summarises the previous one. Over many rounds, details erode and the summary's own phrasing hardens into "fact". Mitigate with a prompt that copies identifiers verbatim, and with state for critical facts.
  • Thrashing. A token threshold close to the size of one turn compacts on almost every turn. You pay the summariser each time, and you invalidate any prompt cache prefix each time.
  • Orphaned tool context. A retained tail that starts in the middle of a tool exchange, a response without its call, confuses some models. Size retention generously in events.
  • Expensive summariser. Using the main model for summaries can cost more than the tokens saved. Use a cheaper model and confirm the net saving from usage data.
  • Assuming storage shrinks. Compaction appends; raw events remain. Session stores still need their own retention and deletion policy, including for privacy requests.
  • Silent version mismatch. The builder methods your code calls must exist in the ADK Java release you ship. Pin the version and keep a test that runs a session past a compaction.

What to do next

  1. Measure prompt tokens per turn on real long sessions before enabling anything, and note the turn at which cost or latency becomes a problem.
  2. Enable sliding-window compaction with interval 3 to 5 and overlap 1, using a cheaper summariser model.
  3. Wrap the summariser in a guarded decorator with a timeout, metrics and onErrorComplete().
  4. Move must-not-forget facts (IDs, commitments, user preferences) into session state.
  5. Build the compacted-range evaluation set and compare task success with compaction on and off.
  6. If some sessions carry very large tool outputs, add a token threshold at 60 to 70 percent of the window as a safety net, and watch for thrashing.
  7. Add an integration test that drives a session through several compactions on your pinned ADK Java version.
Key takeaway: ADK Java compresses context by appending summary events that cover a range of older events, which the request builder then substitutes for that range; raw history stays in the session. Start with sliding-window compaction and a cheap, guarded summarizer, keep critical facts in session state, add a token threshold as a safety net, and prove with an evaluation set that the summaries keep what your agent needs.