An ADK Java agent has two kinds of memory. The session is the conversation in progress: an append-only list of events that the runner replays into each model call. Long-term memory is what survives after the session ends, held by a BaseMemoryService and searched in later sessions. Between them sits one method, addSessionToMemory(Session session), and one decision that the framework leaves entirely to you: what should a finished conversation turn into before it is stored?
The answer is a summarization strategy, and it matters more than the choice of vector database. Store raw events and later searches return greetings, tool scaffolding and half-finished thoughts. Store a single paragraph and the one identifier you needed is paraphrased away. Store facts without keys and the user who changed their mind last week gets both answers back. This article compares five strategies for the memory write path, builds a summarizing memory service on ADK Java's real interfaces, and shows how to test a summary before it becomes something the agent believes.
Two neighbours are deliberately out of scope. Shrinking the prompt inside a running session is compaction, covered in ADK Java context compression. Building a dedicated agent whose job is producing summaries is covered in the ADK Java summary agent. Here the question is narrower: what goes into long-term memory, in what shape, and how it is kept true over time.
The memory write path in ADK Java
Start with the interface, because it constrains every design. In ADK Java, BaseMemoryService has two methods: Completable addSessionToMemory(Session session) and Single<SearchMemoryResponse> searchMemory(String appName, String userId, String query). A search returns a list of MemoryEntry objects, each holding a Content, an optional author and an optional timestamp string that is forwarded to the model. The bundled InMemoryMemoryService stores every event with text and matches searches by keyword overlap; its own documentation calls it a prototype. It is a useful baseline precisely because it does no summarization at all.
The Javadoc on addSessionToMemory adds one sentence that shapes the whole design: a session may be added multiple times during its lifetime. You might ingest when a session goes idle, again when it resumes and closes, and again after a crash recovery. Any strategy that appends on every call will store the same fact three times. Every write must therefore be idempotent per session: replace what this session previously contributed rather than adding to it.
Five summarization strategies
There are five strategies in practical use. They differ in what they keep, what they lose and what they cost per session.
| Strategy | What is stored | Good at | Loses |
|---|---|---|---|
| Raw events | Every text event, verbatim | Exact wording, zero model cost | Precision; recall drowns in small talk |
| Session abstract | One paragraph per session | Cheap, readable, good for 'what did we do' | Identifiers, numbers, minority topics |
| Structured facts | Typed, keyed facts with evidence | Preferences and decisions, exact recall | Narrative and reasoning |
| Episodic digest | Short abstract plus a pointer to the session | 'Last Tuesday we...' recall | Needs the session store to stay around |
| Consolidated profile | Facts merged across sessions | Bounded size, current truth | History of how it changed |
Raw events suit only small prototypes. A session abstract is what most teams build first: render the transcript, ask a model for a paragraph, store it. It reads well and it fails quietly, because a model writing prose optimises for the main thread of the conversation and drops the side detail, often the order number or the deadline.
Structured fact extraction asks the model for a list of typed records instead of prose. Each record has a kind (preference, decision, fact about the user's world, open item), a stable key such as deploy.region, a one-sentence statement and the id of the event that supports it. The key is what makes later supersession possible; the evidence id is what makes validation possible. The cost is that reasoning disappears: you remember that the user chose plan B, not why.
An episodic digest stores a short abstract plus the session id, so the agent can find the right past conversation and, when it needs detail, reload it from the session service. Consolidation is not a separate extraction method but a second pass over stored facts: periodically merge, deduplicate and retire superseded entries per user so memory stays bounded. Production systems usually combine three of these: structured facts for things the agent must get exactly right, an episodic digest for continuity, and consolidation to stop growth.
A summarizing memory service
The service below implements the combined strategy. It renders the session, asks a cheap model for facts and a digest, validates the result, and replaces the session's previous contribution in one store call. The Extractor and FactStore interfaces are yours; back the first with any ADK model call and the second with a relational table, a vector index or both.
public record Fact(String kind, String key, String text, Instant observedAt,
String sessionId, String evidenceEventId) {}
public record Extraction(String digest, List<Fact> facts) {}
public final class SummarizingMemoryService implements BaseMemoryService {
private final Extractor extractor; // cheap model, returns JSON parsed into Extraction
private final FactStore store; // replaceForSession, search, enqueueRetry
public SummarizingMemoryService(Extractor extractor, FactStore store) {
this.extractor = extractor;
this.store = store;
}
@Override
public Completable addSessionToMemory(Session session) {
String scope = session.appName() + "/" + session.userId();
Map<String, Event> byId = session.events().stream()
.collect(Collectors.toMap(Event::id, e -> e, (a, b) -> a));
return extractor.extract(render(session.events()))
.timeout(30, TimeUnit.SECONDS)
.map(x -> Validator.check(x, byId, session.id()))
.flatMapCompletable(x -> store.replaceForSession(scope, session.id(), x))
.onErrorResumeNext(err -> store.enqueueRetry(scope, session.id(), err));
}
@Override
public Single<SearchMemoryResponse> searchMemory(String appName, String userId, String query) {
return store.search(appName + "/" + userId, query, 8)
.map(facts -> SearchMemoryResponse.builder()
.memories(facts.stream().map(SummarizingMemoryService::toEntry).toList())
.build());
}
private static MemoryEntry toEntry(Fact f) {
return MemoryEntry.builder()
.content(Content.fromParts(Part.fromText(f.kind() + " (" + f.key() + "): " + f.text())))
.author("memory")
.timestamp(f.observedAt())
.build();
}
static String render(List<Event> events) {
StringBuilder sb = new StringBuilder();
for (Event e : events) {
if (e.partial().orElse(false)) continue; // streaming fragments
if (!e.functionResponses().isEmpty()) { // tool bodies are not memories
sb.append("[").append(e.id()).append("] tool result omitted\n");
continue;
}
String t = e.stringifyContent();
if (!t.isBlank()) sb.append("[").append(e.id()).append("] ")
.append(e.author()).append(": ").append(t).append('\n');
}
return sb.toString();
}
}Three details carry the design. The renderer tags every line with the event id, so the model can cite evidence and the validator can check it. Tool responses are replaced by a marker, because a 40 KB search result is not something the user said and storing it turns memory into a cache of stale data. And failure goes to a retry queue instead of storing a partial result: a missing memory is recoverable, a wrong one is acted on.
The validator is short and does most of the safety work:
static Extraction check(Extraction x, Map<String, Event> byId, String sessionId) {
List<Fact> ok = new ArrayList<>();
for (Fact f : x.facts()) {
Event ev = byId.get(f.evidenceEventId());
if (ev == null) continue; // cited a line that does not exist
if (!KINDS.contains(f.kind()) || !KEY.matcher(f.key()).matches()) continue;
if (f.kind().equals("preference") && !"user".equals(ev.author())) continue;
ok.add(new Fact(f.kind(), f.key(), f.text(),
Instant.ofEpochMilli(ev.timestamp()), sessionId, ev.id()));
}
return new Extraction(x.digest(), ok);
}Note the preference rule. A preference must be supported by something the user said, not by something the agent or a tool said. That single check removes a whole class of laundering, where text inside a retrieved document claims to be the user's preference and becomes permanent memory. The timestamp is taken from the evidence event, not from the model's output, so it cannot be invented.
Keeping memory true over time
Facts change. The structured strategy handles that with keys. When the store receives a fact whose (scope, key) already exists from an earlier session, it marks the old one superseded and keeps it for audit, but search returns only the newest. Without keys, two contradictory statements are equally relevant to a query and the model picks one by chance.
Keys need a small controlled vocabulary or they fragment: deploy.region and deployment.region.preferred will never supersede each other. Give the extraction prompt a list of known keys for your domain and let it propose new ones under an other. prefix that you review and promote.
Consolidation runs as a scheduled job per user, not on the request path. It loads the user's active facts, merges near-duplicates that share a key, drops open items that a later decision closed, and caps the total. If a user has accumulated 400 facts, retrieval quality falls because ranking has too many near-ties; a cap of a few dozen active facts per key family is a reasonable start. Summarizing summaries is where drift creeps in, so consolidation should only merge records and delete superseded ones, never rewrite the text of a fact it cannot trace back to an evidence event.
Deletion requests interact with all of this. Because every fact carries its session id and evidence id, forgetting a session means deleting rows by session id and re-running consolidation. Abstract-only strategies cannot do this cleanly: once a sentence from a deleted session is blended into a paragraph, it is gone from nowhere. The details are in memory privacy and the right to forget.
Worked example: three sessions, one preference
Follow one user of an internal deployment assistant across three sessions.
- Session A. The user says they deploy to eu-west and never use us-east. The extractor returns
preference deploy.region: deploys to eu-west, avoids us-eastciting the user's event. The digest reads 'Set up first service; region preference stated.' Ingestion runs at idle and again at close; the second call replaces the first, so there is still one fact. - Session B. A search tool returns a runbook stating 'all teams must use us-central'. The model, asked for facts, proposes
preference deploy.region: us-centralciting the tool's event. The validator drops it, because the evidence author is not the user. - Session C. The user says they moved everything to us-east. The extractor emits a new fact under the same key. The store supersedes session A's fact; search now returns only the us-east preference, with session C's timestamp, so the model can also say when it changed.
A fourth session asks the agent to scaffold a deployment. The agent searches memory, receives one entry, and defaults to us-east. With a session-abstract strategy the same search would have returned three paragraphs, two of them mentioning eu-west, and the result would depend on ranking. Retrieval behaviour on the read side is covered in long-term memory retrieval in ADK Java.
Evaluating a summary before trusting it
A summary is a lossy function and you should measure the loss before trusting it. Build a probe set from real sessions: for each session, write three to five questions whose answers appear only once in the transcript, such as an order number, a stated deadline or the option the user rejected. Ingest the session, then answer each question using only searchMemory output. Score exact-match recall per strategy.
for (Probe p : probes) {
memory.addSessionToMemory(p.session()).blockingAwait();
SearchMemoryResponse r = memory.searchMemory(APP, p.userId(), p.question()).blockingGet();
String recalled = r.memories().stream()
.map(m -> m.content().text()).collect(Collectors.joining("\n"));
results.add(p.id(), recalled.contains(p.expectedAnswer()));
}Track three numbers per strategy: recall on the probes, the false-memory rate (sampled facts a human reviewer cannot find in the transcript) and entries stored per session. A strategy with high recall and a false-memory rate above zero is worse than one with lower recall, because the agent will repeat invented facts with confidence. Re-run the probe set whenever you change the extraction prompt or the model.
Failure modes
- Duplicate memories. Appending on every
addSessionToMemorycall stores each fact once per ingestion. Replace by session id. - Identifier erosion. Prose summaries paraphrase numbers and ids. Extract them as keyed facts and check exact-match recall.
- Stale truth. Unkeyed facts never supersede each other. Key them and return only the newest per key.
- Memory laundering. Text from tools or documents becomes a 'user preference'. Require user-authored evidence for preferences and treat recalled memory as data, not instruction.
- Summary-of-summary drift. Consolidation that rewrites text slowly invents facts. Merge and delete; do not paraphrase.
- Blocking the turn. Running extraction synchronously at the end of a turn adds model latency to the user's wait. Ingest in the background and accept that memory is eventually consistent.
- Scope leaks. A store keyed only by user id mixes applications. Always key by app and user, as the interface's own search signature does.
Trade-offs
Abstracts cost one small model call per ingestion and lose detail. Structured facts cost the same call plus a validator and a key vocabulary, and in exchange give exact recall, supersession and clean deletion. If you can afford only one model call per session, ask it for facts and a one-line digest in the same JSON response rather than choosing between them. If facts the agent must never lose show up during a session, write them to session state as they happen; memory ingestion is a backstop, not the primary record. The service wiring itself is covered in the ADK Java memory service.
What to do next
- List what your agent actually needs to remember across sessions and group it into kinds and keys.
- Measure your current strategy with a probe set of real sessions before changing anything.
- Implement
addSessionToMemoryas replace-by-session-id, and test it by ingesting the same session twice. - Render transcripts with event ids, omit tool bodies and partial events, and require evidence ids in the extraction output.
- Add the validator rules: known kinds, key format, user-authored evidence for preferences, timestamps from events.
- Key facts and supersede by key; add a nightly consolidation job that merges and deletes but never rewrites.
- Track probe recall, false-memory rate and entries per session on every prompt or model change.