The most common memory bug report for an ADK Java agent is short: "it doesn't remember". The user said something in Monday's session, and on Thursday the agent acts as if it never heard it. Behind that report is a pipeline of four steps: something must write the session into memory, under the right key, then the model must ask for it with a query that the store's matching rules accept, and the content must be in a form the store indexes. If any step fails, the result is the same: LoadMemoryResponse with an empty memories list. An empty list does not tell you which step failed.
This article is about telling those cases apart. It explains how the memory classes in ADK Java actually behave, builds an inspecting wrapper and a tracing plugin, works through a debugging session, and ends with tests that keep the bug fixed. The ideas behind what to remember are in cross-session memory, and ranking and recall measurement are in long-term memory retrieval. All class and method names were checked against the google/adk-java source on the main branch in October 2026. Confirm them against the version you build with.
The pieces and how they connect
The contract is the BaseMemoryService interface in com.google.adk.memory, with two methods:
public interface BaseMemoryService {
Completable addSessionToMemory(Session session); // may be called many times per session
Single<SearchMemoryResponse> searchMemory(String appName, String userId, String query);
}The core module ships one implementation, InMemoryMemoryService, which InMemoryRunner wires in by default. Agents reach memory through LoadMemoryTool, a FunctionTool whose loadMemory(String query, ToolContext toolContext) calls toolContext.searchMemory(query) and wraps the result in LoadMemoryResponse, a record holding a List<MemoryEntry>. The tool also adds an instruction to every request telling the model that it has memory and should call loadMemory when a question needs it. ToolContext.searchMemory fills in the app name and user id from the current session and throws IllegalStateException("Memory service is not initialized.") if the runner has no memory service. Each MemoryEntry carries the event's content, its author and an ISO-8601 timestamp string.
Checks 1 and 2: written, and under the right key
Check 1: was anything written? The Runner accepts a memory service through its builder or constructor and passes it into the invocation context, but nothing in the runner calls addSessionToMemory. Writing memory is your application's job. If no code calls it, every search returns nothing, and the code that reads memory looks fine. Decide when a session is worth remembering, for example at the end of a conversation or after each completed turn, and add the call explicitly:
// after a turn completes: reload the session so it includes every appended event
Session session = sessionService
.getSession(appName, userId, sessionId, Optional.empty())
.blockingGet(); // Maybe<Session>; null if absent
if (session != null) {
memoryService.addSessionToMemory(session).blockingAwait();
}Re-reading the session matters. The in-memory service stores session.immutableEvents() as they are at the moment of the call, and calling it again for the same session id replaces that session's entry rather than appending to it. That makes repeated calls safe, but a call made with a stale Session object quietly stores an old copy of the conversation.
Check 2: is it under the right key? The store is a ConcurrentHashMap keyed by appName + "/" + userId, then by session id. A search for support-bot/u-42 sees nothing written under support_bot/u-42 or support-bot/U-42. This happens when the app name passed to the runner differs from the one used by a back-office job that writes memory, or when one code path uses an email address as the user id and another uses a database id. Log the exact key on both sides.
Checks 3 and 4: match rules and indexed content
Check 3: do the match rules accept the query? InMemoryMemoryService is a keyword matcher, and the two sides are tokenised differently. The query is lowercased and split on whitespace only. The text of each event is tokenised with the regular expression [A-Za-z]+ and lowercased. An event matches if at least one query token equals one event word. Every matching event is returned, with no score and no limit, and the order across sessions follows a hash map, so it is not defined. Suppose the stored event is "My favourite coffee is a flat white. Order 4417 shipped." Its words are my, favourite, coffee, is, a, flat, white, order and shipped. Now try some queries:
| Query the model sends | Query tokens | Result | Why |
|---|---|---|---|
favourite coffee | favourite, coffee | match | both are event words |
coffee? | coffee? | no match | punctuation stays attached to query tokens |
flat-white | flat-white | no match | event side split it into flat and white |
4417 | 4417 | no match | digits are never event words |
café | café | no match | non-ASCII letters are not in [A-Za-z] |
what is my order | what, is, my, order | match, plus noise | is and my match almost every event |
The last row is the opposite failure. Common words match nearly everything, so a broad query returns the user's whole history, unranked, and the model may pick the wrong memory. Both failures are properties of this implementation, not of your agent. It is a reference implementation for development, not a search engine. The fix in production is a real retrieval backend behind the same interface, as described in semantic memory with vector stores.
Check 4: is the content indexed at all? When writing, events whose content has no parts are dropped. When searching, only part.text() is read. Function calls, function responses and inline data have no text, so a fact that exists only inside a tool result, such as an order status returned by an API, is stored but can never be found. The same applies to session state: values in session.state() are not part of any event's text. If a fact should be remembered, make sure it appears in text, for example in the model's own reply, or write a summary event before calling addSessionToMemory.
An inspecting memory service
To inspect memory without changing the agent, wrap the real service in a decorator. It implements the same interface, so the runner cannot tell the difference. It logs keys and counts, never content, and offers an explain method that repeats the in-memory tokenisation so you can see why a query missed:
public final class InspectingMemoryService implements BaseMemoryService {
private static final Logger log = LoggerFactory.getLogger(InspectingMemoryService.class);
private static final Pattern WORD = Pattern.compile("[A-Za-z]+"); // mirrors InMemoryMemoryService
private final BaseMemoryService delegate;
public InspectingMemoryService(BaseMemoryService delegate) { this.delegate = delegate; }
@Override
public Completable addSessionToMemory(Session session) {
long textEvents = session.immutableEvents().stream() // events the matcher can ever see
.filter(e -> e.content().flatMap(Content::parts).orElse(List.of()).stream()
.anyMatch(part -> !part.text().orElse("").isEmpty()))
.count();
return delegate.addSessionToMemory(session)
.doOnComplete(() -> log.info("memory.write key={}/{} session={} events={} textEvents={}",
session.appName(), session.userId(), session.id(),
session.immutableEvents().size(), textEvents));
}
@Override
public Single<SearchMemoryResponse> searchMemory(String appName, String userId, String query) {
long start = System.nanoTime();
return delegate.searchMemory(appName, userId, query)
.doOnSuccess(r -> log.info("memory.search key={}/{} tokens={} hits={} ms={}",
appName, userId, query.toLowerCase(Locale.ROOT).split("\\s+").length,
r.memories().size(), (System.nanoTime() - start) / 1_000_000));
}
/** Which query tokens can never match a text-only, [A-Za-z]+ index. */
public static List<String> explain(String query) {
List<String> dead = new ArrayList<>();
for (String t : query.toLowerCase(Locale.ROOT).split("\\s+")) {
if (!WORD.matcher(t).matches()) dead.add(t);
}
return dead;
}
}Register it with Runner.builder().memoryService(new InspectingMemoryService(new InMemoryMemoryService())), plus your agent, app name and session service. The textEvents count reads only part.text(), like the store does. Do not use Event.stringifyContent() for it: that helper also renders function calls and responses, so it would count events the matcher can never see. The two log lines answer checks 1 and 2 at once: a search for a key that no write line ever mentioned is a key mismatch, and no write lines at all means memory is never written.
Tracing what the model asks for
The decorator shows what the store did. To see what the model asked for, trace the tool calls. A plugin sees every tool invocation across all agents in the runner. Plugin.beforeToolCallback(BaseTool, Map<String,Object>, ToolContext) receives the arguments and afterToolCallback receives them again with the result. Returning Maybe.empty() from both leaves behaviour unchanged:
public final class MemoryTracePlugin extends BasePlugin {
private static final Logger log = LoggerFactory.getLogger(MemoryTracePlugin.class);
private final Map<String, Long> started = new ConcurrentHashMap<>();
public MemoryTracePlugin() { super("memory_trace"); }
@Override
public Maybe<Map<String, Object>> beforeToolCallback(
BaseTool tool, Map<String, Object> args, ToolContext ctx) {
if ("loadMemory".equals(tool.name())) {
ctx.functionCallId().ifPresent(id -> started.put(id, System.nanoTime()));
log.info("memory.query q={} dead={}", args.get("query"),
InspectingMemoryService.explain(String.valueOf(args.get("query"))));
}
return Maybe.empty();
}
@Override
public Maybe<Map<String, Object>> afterToolCallback(
BaseTool tool, Map<String, Object> args, ToolContext ctx, Map<String, Object> result) {
if ("loadMemory".equals(tool.name())) {
ctx.functionCallId().map(started::remove).ifPresent(t0 ->
log.info("memory.tool ms={}", (System.nanoTime() - t0) / 1_000_000));
}
return Maybe.empty();
}
}The tool's registered name comes from its method, so check tool.name() in a debugger once rather than trusting the string above. Log the query text only in development, or redact it, because it often contains personal data. Finally, if the plugin never logs a loadMemory call, the problem is earlier still: the agent does not have LoadMemoryTool in its tools list, or the model decided not to call it. The second case needs instruction or prompt work, not memory work.
Worked example: the forgotten coffee order
A support agent is told on Monday, "My favourite coffee is a flat white." On Thursday the user asks "What coffee do I usually get?" and the agent says it does not know. Here is the debugging session, using the logs from the decorator and the plugin:
- The plugin logs
memory.query q=coffee preference? dead=[preference?]. The model did call the tool. One token can never match, butcoffeecan, so the query is not the cause. - The decorator logs
memory.search key=support/u-42 tokens=2 hits=0. The search ran undersupport/u-42and found nothing. - Searching the logs for
memory.writefinds Monday's write underkey=support/42. The writing job built the user id from a numeric database id, while the chat endpoint usesu-42. This is a key mismatch, check 2. - After the user id is normalised in one shared function, the search logs
hits=1. A second test, "Where's order 4417?", still fails. The order number only appeared inside a function response and the query token is digits. That is check 4 plus check 3, and it is fixed by moving to a real retrieval backend and by having the agent restate the order status in text.
Tests that keep it fixed
Once a bug is fixed, pin it with a test that runs through the real interface, so it also guards a future backend swap. With the in-memory service and a hand-built session this needs no model at all:
@Test
void recallsAcrossSessionsUnderSameKey() {
BaseMemoryService memory = new InMemoryMemoryService();
Session monday = Session.builder("s-1").appName("support").userId("u-42")
.events(List.of(Event.builder().author("user")
.content(Content.fromParts(Part.fromText("My favourite coffee is a flat white.")))
.build()))
.build();
memory.addSessionToMemory(monday).blockingAwait();
assertEquals(1, memory.searchMemory("support", "u-42", "coffee").blockingGet().memories().size());
assertEquals(0, memory.searchMemory("support", "42", "coffee").blockingGet().memories().size());
assertEquals(0, memory.searchMemory("support", "u-42", "coffee?").blockingGet().memories().size());
}The second and third assertions record known behaviour, not desired behaviour. When you move to a backend that normalises punctuation, the third will fail, and that failure is a useful sign that the behaviour changed. Add an end-to-end test per release that runs two sessions through the runner and asserts that the second one's loadMemory result contains the fact.
Failure modes and production concerns
| Failure mode | Signal | Fix |
|---|---|---|
| Memory never written | No write lines in logs | Call addSessionToMemory explicitly at a defined point |
| Key mismatch | Writes and searches use different keys | One function to build appName and userId |
| Stale session written | Last turns missing from recall | Re-fetch the session before writing |
| Query punctuation or digits | Plugin shows dead tokens | Real retrieval backend; query shaping |
| Fact only in tool output | Write count fine, no text events | Restate facts in text or write summaries |
| Replicas disagree | Recall works on one pod only | Shared backend; in-memory is process-local |
| Noisy recall | Large unranked hit lists | Ranking, caps, stop words in a real backend |
Two production points sit behind the table. InMemoryMemoryService lives in one JVM: with three replicas behind a load balancer, a write lands on one of them and later searches are spread across all three, and everything is lost on restart. It also grows without limit, because nothing evicts old sessions. Use it in tests and local runs, and put a durable, shared, ranked store behind BaseMemoryService for anything else. Inspection should also respect privacy: log keys, counts and timings by default, and gate content logging behind a debug flag, in line with the right-to-forget design.
What to do next
- Search your code for
addSessionToMemory. If there are no calls, decide when sessions are written and add one. - Put app name and user id construction in one shared function and use it in every writer and in the runner setup.
- Wrap your memory service in the inspecting decorator and register the trace plugin in development and staging.
- Replay a failed conversation and walk through the four checks in order: written, key, match rules, indexed content.
- Add the three-assertion unit test, plus one two-session end-to-end test through the runner.
- If you still run
InMemoryMemoryServiceoutside tests, plan the move to a shared, ranked backend and measure recall before and after. - Default inspection logs to keys and counts only, with content behind a flag.