Memory compression is deciding what an agent may forget. A long conversation does not fit in a context window forever, and even when it fits, paying for forty thousand tokens of old tool output on every call is waste. So something has to become shorter. Every compression of natural-language memory is lossy, though, and the losses are rarely random: summaries drop exact numbers, identifiers and the commitments the agent made, which are the facts users notice first when they go missing.
This page treats compression as an engineering problem with a measurable error. It covers where the tokens in an ADK Java session actually come from, compressing at the source before anything is summarized, what ADK Java 1.11.0's built-in compaction does and does not do, why rolling summaries drift, a summarizer that pins facts, and an evaluation that tells you how much a configuration forgets. The compaction API itself is covered in depth in context compression in ADK Java; here it appears only as one of three tools.
Where the tokens come from
Before compressing, measure. Every model event in ADK Java can carry usageMetadata(), and its prompt token count is the real size of what the model read. Log it per turn, then attribute it by walking the events and estimating each one's share. In a typical support agent the result looks something like this (illustrative, from a hypothetical twenty-turn session):
| Source | Share of prompt | Compressible how |
|---|---|---|
| Tool results (search hits, records, logs) | about 60% | At source: trim, cap, offload |
| System instruction and tool declarations | about 15% | Not by compaction; shorten the text |
| User and agent messages | about 20% | Summarize old turns |
| State injected into the instruction | about 5% | Keep state small and structured |
The lesson generalizes: the largest item is usually tool output the model already acted on, and the instruction plus tool declarations are a fixed cost that compaction never touches. If your numbers look like this, summarizing conversation turns is the third thing to fix, not the first. Token counting across models explains why estimates and billed counts differ.
Compress at the source first
Compression at the source is lossless where it matters, because you choose what to drop while you still know what it means. Three techniques do most of the work.
- Return less from tools. A search tool that returns ten full documents could return ten titles, ids and two-sentence snippets, with a second tool to fetch one document. The model pays for detail only when it asks.
- Offload large payloads to artifacts. Save a 200 KB log or CSV with the artifact service and return its name plus a summary the tool computed: row count, columns, error count. The bytes stay retrievable and out of the prompt. See ADK Java artifacts.
- Keep state structured. Session state injected into the instruction should be facts, such as
{"plan":"pro","open_ticket":4471}, not transcripts. Overwrite values instead of appending history to them.
These are deterministic, cheap and testable with ordinary unit tests, which is why they come first. Short-term memory in ADK Java shows how to trim tool output with callbacks.
What built-in event compaction does
For conversation history, ADK Java 1.11.0 provides event compaction, configured with EventsCompactionConfig on the App. Its behaviour, read from the bytecode, comes down to a few facts that decide what you lose:
- Token-triggered (tail retention). Runs before each model call, only when both
tokenThresholdandeventRetentionSizeare set. It compares the latest prompt token count fromusageMetadatawith the threshold, falling back to characters divided by four if there is none. Over the threshold, it keeps the lasteventRetentionSizeevents verbatim and summarizes everything older, including the previous summary. - Interval-triggered (sliding window). Runs after an invocation finishes, when
compactionIntervalis above zero andoverlapSizeis set. Once enough new invocations have accumulated, it summarizes them plusoverlapSizeearlier ones for continuity. - Nothing is deleted. A compaction is a new event whose
EventActions.compaction()holds start and end timestamps and the summary content. The request builder substitutes it for the events it covers. The prompt shrinks; the stored session grows. - Default summarizer. Without one, the runner builds an
LlmEventSummarizerfrom the root agent's model, and fails if the root is not anLlmAgent. Its default prompt asks for a concise summary of key information, decisions and unresolved questions.
App app = App.builder()
.name("support")
.rootAgent(supportAgent)
.eventsCompactionConfig(EventsCompactionConfig.builder()
.tokenThreshold(24_000) // compact when the last prompt exceeded this
.eventRetentionSize(12) // always keep the newest 12 events verbatim
.summarizer(new PinningSummarizer(summaryModel))
.build())
.build();
Runner runner = Runner.builder().app(app)
.sessionService(sessions).artifactService(artifacts).build();Summary drift
Tail-retention compaction is recursive: the next summary is written from the previous summary plus newer events. After five compactions, a fact from the first turn has been paraphrased five times. Each paraphrase is reasonable on its own and the sequence is not. "Refund of 412.50 EUR approved by Dana on 3 October, pending finance" becomes "a refund was approved" and then disappears, because later summaries judge it settled. This is summary drift, and it grows with session length, which is exactly where compaction is used.
The fix is to stop paraphrasing what must not change. Separate the summary into two parts: pinned facts copied verbatim from one generation to the next, and narrative that may be rewritten freely.
A summarizer that pins facts
The summarizer below decorates LlmEventSummarizer. It asks the model for a summary that ends with a facts block, then merges that block with the facts from the previous compaction so a fact survives even if the model omits it. It builds the result with the public EventCompaction and EventActions builders.
public final class PinningSummarizer implements BaseEventSummarizer {
private static final String PROMPT = """
Summarize this agent conversation for the agent that will continue it.
Then output a line FACTS: followed by one fact per line, verbatim:
ids, amounts with currency, dates, names, and every commitment made.
{conversation_history}
""";
private static final String MARK = "FACTS:";
private final LlmEventSummarizer delegate;
public PinningSummarizer(BaseLlm model) {
this.delegate = new LlmEventSummarizer(model, PROMPT);
}
@Override
public Maybe<Event> summarizeEvents(List<Event> events) {
Set<String> carried = new LinkedHashSet<>();
for (Event e : events) { // facts from older summaries
e.actions().compaction().ifPresent(cmp -> carried.addAll(facts(text(cmp.compactedContent()))));
}
return delegate.summarizeEvents(events).map(ev -> {
EventCompaction old = ev.actions().compaction().orElseThrow();
String body = text(old.compactedContent());
Set<String> all = new LinkedHashSet<>(carried);
all.addAll(facts(body)); // union: never drop a pinned fact
String narrative = body.contains(MARK) ? body.substring(0, body.indexOf(MARK)) : body;
Content merged = Content.builder().role("model")
.parts(List.of(Part.fromText(narrative.strip() + "\n" + MARK + "\n"
+ String.join("\n", all))))
.build();
return ev.toBuilder().actions(EventActions.builder()
.compaction(EventCompaction.builder()
.startTimestamp(old.startTimestamp())
.endTimestamp(old.endTimestamp())
.compactedContent(merged).build())
.build()).build();
});
}
// text(): concatenate text parts; facts(): lines after MARK, trimmed, non-empty
}The union only grows, so cap it: expire facts tied to closed tickets, or keep the newest fifty. The cap is a policy decision, and making it explicit is the point. Facts that matter beyond the session belong in state or long-term memory, not in a summary at all.
Measuring what compression loses
A compression setting is only as good as what the agent can still answer afterwards. Measure it with a fact-retention evaluation: replay recorded sessions, compact them with the candidate configuration, and ask probe questions whose answers sit in the compacted part.
record Probe(String question, String expected) {}
double retention(List<RecordedSession> corpus, Function<List<Event>, List<Event>> compress,
Agent answerer) {
int asked = 0, kept = 0;
for (RecordedSession s : corpus) {
List<Event> history = compress.apply(s.events()); // candidate configuration
for (Probe p : s.probes()) { // written from early turns only
asked++;
String a = answerer.ask(history, p.question());
if (Normalize.contains(a, p.expected())) kept++; // exact ids and amounts
}
}
return (double) kept / asked;
}Write probes before you look at any summary, from the raw early turns, or you will unconsciously write questions the summary happens to answer. Ask them through the same agent and instruction you run in production, because retention depends on how the answering model reads a summary as much as on the summary itself. Keep the corpus fixed between runs and version it, so a change in score means a change in configuration and not in data. Two hundred probes over twenty sessions is enough to separate configurations that differ by ten points; for smaller differences, add sessions rather than probes per session, since probes from one session are correlated.
Run it three times: with no compression (the ceiling, which is rarely 100 percent), with the default summarizer, and with your candidate. Report retention next to the prompt size reduction, and pick the configuration on that curve rather than by intuition. Write probes for the facts users notice: amounts, ids, dates, promises. A summary that keeps the gist and loses the order number scores badly here, which is the correct result.
Compressing long-term memory
Long-term memory has the same problem with a longer horizon. BaseMemoryService.addSessionToMemory stores what you give it, and a whole session is a poor memory. Distil first: extract durable facts and preferences with their source session and date, drop the chit-chat, and store those. Memory summarization strategies covers what a finished session should become. The same retention evaluation applies: probe a fresh session with questions answerable only from memory.
Worked example: a billing agent
A billing agent averages 31,000 prompt tokens by turn twenty, and users complain that it forgets refunds it promised. Measurement shows 58 percent of the prompt is invoice payloads from a getInvoice tool. Step one returns invoice headers and line counts only, with the full invoice saved as an artifact; the average drops to 14,000 tokens with no measured retention loss, because the model never needed the line items twice. Step two enables tail-retention compaction at 12,000 tokens with twelve retained events and the default summarizer; prompts fall to about 9,000, but probe retention drops from 0.94 to 0.71, mostly lost refund amounts. Step three swaps in the pinning summarizer; retention returns to 0.92 at about 9,500 tokens. (Figures are illustrative of the method, not a benchmark.)
Failure modes
- Compacting before trimming. Summarizing huge tool payloads is slow, costly and still loses detail. Trim at source first.
- Summary drift. Recursive summaries lose early facts. Pin them, or move them into state.
- Threshold above the window. The check reads the prompt token count stored on the latest event, which comes from the previous call, so compaction always lags one call behind. A threshold close to the model's limit fires too late, and the call that crosses the limit fails. Leave headroom of at least one large tool result.
- Retention size too small. Keeping two events can split a function call from its response in the verbatim tail. Retain enough events to cover a full tool round trip.
- Expecting storage to shrink. Compaction adds events. Retire old sessions separately.
- Summarizer failures. Tail-retention compaction runs inside the request processor and the model request waits for it, so a slow summary delays the turn and a failed one fails the turn's model request. Add a timeout and return empty on error so the session simply stays uncompacted.
- No evaluation. Without probes, you learn what was forgotten from users.
Trade-offs
Every lever trades tokens for fidelity, latency or engineering time. Source trimming is cheap and nearly lossless but needs per-tool work. Event compaction is generic but lossy, and each compaction is an extra model call at an unpredictable moment. Pinning facts raises fidelity at the cost of a longer summary that only grows. A larger retained tail protects recent context and saves less. Choose with the retention curve in hand, and revisit when the model or tools change, because both move the curve.
What to do next
- Log prompt token counts from
usageMetadataper turn and attribute them to instruction, tools, messages and tool results. - Trim the largest tool results at source and offload big payloads to artifacts.
- Keep session state as small, overwritten facts.
- Write probe questions for twenty recorded long sessions, focused on ids, amounts, dates and commitments.
- Enable tail-retention compaction with headroom below the context window and a retained tail that covers a tool round trip.
- Measure retention against prompt size for the default summarizer, then for a fact-pinning one.
- Add a timeout and empty-on-error behaviour to the summarizer, and alert on compaction failures.
- Distil durable facts before writing to long-term memory, and test them with the same probes.