An agent with long-term memory has two jobs that are easy to confuse. Writing memory means deciding what to keep from finished sessions and storing it somewhere durable. Retrieving memory means deciding, in the middle of a turn, which few stored items belong in the model's context right now. Most teams spend their effort on the first job and then find that the agent remembers everything and recalls almost nothing useful.
This article is about the second job in ADK for Java: the framework's contract, a retrieval policy built as a decorator over any memory service, and how to measure recall. For the service overview see ADK Java MemoryService architecture, and for a vector-backed store and its schema see Semantic memory and vector stores.
The four decisions retrieval makes
Every recall, whoever triggers it, answers four questions. When: on every turn, or only when the model asks? With what query: the user's raw words, or a query rewritten to match how memories are stored? How many: how many items, and how many tokens, may enter the context? In what form: raw text, or text labelled with its date and source so the model can weigh it?
None of these questions is answered by the store. A vector index answers a different one: which stored items are nearest to this query vector. That is why a well-built store can still produce a poor agent. Treat retrieval as a policy with its own code, tests and metrics, not as a single call to a similarity search.
What ADK Java gives you
The contract is the BaseMemoryService interface in com.google.adk.memory, and it has two methods. addSessionToMemory(Session session) returns an RxJava Completable and is the write path. searchMemory(String appName, String userId, String query) returns Single<SearchMemoryResponse> and is the read path. A SearchMemoryResponse holds an immutable list of MemoryEntry values. Each entry has a Content, an optional author and an optional timestamp, and the timestamp is a String, not an Instant.
The model reaches memory through LoadMemoryTool. It exposes a function with one query string argument, adds an instruction telling the model that it has memory and should call the function when a question needs it, and calls ToolContext.searchMemory(query), which fills in the current app and user. Its result is a LoadMemoryResponse containing the memories, and the tool currently uses only the text parts of each entry. At the time of writing, adk-java's core has no ready-made tool that preloads memory into every turn, so the model decides when to recall unless you add that yourself.
The reference implementation is InMemoryMemoryService. It keys sessions by app and user and stores their non-empty events. Its search is deliberately simple, and it is worth reading exactly what it does:
// What InMemoryMemoryService does on searchMemory, paraphrased from the adk-java source:
// query words = query.toLowerCase().split("\\s+")
// event words = every [A-Za-z]+ run in the event's text, lowercased
// match = the two sets share at least one word
// No ranking, no limit, no stemming, and digits are never words.
memoryService.searchMemory("support-app", "user-17", "status of order 4521")
// -> every remembered event that contains "status", "of" or "order",
// in storage order. "4521" cannot match anything.That is fine for tests and a trap elsewhere. There is no ranking, stop words such as "of" match nearly everything, and numbers, often the most precise part of a question, are ignored. The retrieval policy has to be your own.
Pull or push
With pull, the model calls the memory tool when it judges that a question needs history. That is what LoadMemoryTool gives you. It costs nothing on turns that need no memory, and the model writes the query, which is often better than the raw user message. Its weakness is that the model has to know it has forgotten something. Questions that depend on a stored preference without mentioning it, such as "book me a flight" when the user has a stored seat preference, never trigger a call.
With push, your code searches memory before every model call and adds the top few items to the request. Recall stops depending on the model's judgment, but every turn pays for a search and some tokens, and irrelevant memories distract the model. In ADK Java you would do this in a before-model callback. See ADK Java callbacks for the hooks. Because core has no preload tool, this is code you own.
A common split: push a handful of short, high-importance facts, and leave open-ended history to pull, with both going through the same searchMemory policy.
A ranking decorator over any memory service
The cleanest place for a retrieval policy is a class that implements BaseMemoryService itself and wraps the real stores. The runner and LoadMemoryTool never know the difference, and the policy can be unit-tested without a model. The class below is our own code, not part of ADK. It forwards writes to one store, queries several scoped sources in parallel, and fuses, ranks, deduplicates and budgets the results.
// Our own code, not part of ADK: a retrieval policy layered over any BaseMemoryService.
public final class RankedMemoryService implements BaseMemoryService {
private static final int RRF_K = 60;
private final BaseMemoryService writeStore; // where sessions are ingested
private final List<BaseMemoryService> sources; // candidate generators, each scoped by app+user
private final Duration halfLife; // e.g. Duration.ofDays(30)
private final int maxEntries; // e.g. 8
private final int maxChars; // e.g. 4000
private final Clock clock;
public RankedMemoryService(BaseMemoryService writeStore, List<BaseMemoryService> sources,
Duration halfLife, int maxEntries, int maxChars, Clock clock) {
this.writeStore = writeStore;
this.sources = List.copyOf(sources);
this.halfLife = halfLife;
this.maxEntries = maxEntries;
this.maxChars = maxChars;
this.clock = clock;
}
@Override
public Completable addSessionToMemory(Session session) {
return writeStore.addSessionToMemory(session);
}
@Override
public Single<SearchMemoryResponse> searchMemory(String appName, String userId, String query) {
String shaped = QueryShaper.shape(query); // our code: keeps ids, adds known aliases
List<Single<List<MemoryEntry>>> calls = new ArrayList<>();
for (BaseMemoryService source : sources) {
calls.add(source.searchMemory(appName, userId, shaped)
.map(r -> (List<MemoryEntry>) r.memories())
.timeout(300, TimeUnit.MILLISECONDS) // one slow source must not stall the turn
.onErrorReturnItem(List.of())); // degrade to fewer sources, never fail the turn
}
return Single.zip(calls, this::rank)
.map(ranked -> SearchMemoryResponse.builder().memories(ranked).build());
}
@SuppressWarnings("unchecked")
private List<MemoryEntry> rank(Object[] perSource) {
Map<String, MemoryEntry> byText = new LinkedHashMap<>();
Map<String, Double> fused = new HashMap<>();
for (Object o : perSource) {
List<MemoryEntry> list = (List<MemoryEntry>) o;
for (int rank = 0; rank < list.size(); rank++) {
String key = normalize(text(list.get(rank)));
if (key.isEmpty()) continue;
byText.putIfAbsent(key, list.get(rank)); // exact duplicates collapse here
fused.merge(key, 1.0 / (RRF_K + rank + 1), Double::sum);
}
}
Instant now = clock.instant();
List<String> keys = new ArrayList<>(byText.keySet());
keys.sort(Comparator.comparingDouble(
(String k) -> fused.get(k) * recencyWeight(byText.get(k), now)).reversed());
List<MemoryEntry> out = new ArrayList<>();
int used = 0;
for (String k : keys) {
if (out.size() == maxEntries) break;
if (used + k.length() > maxChars) continue; // a smaller entry may still fit
out.add(byText.get(k));
used += k.length();
}
return out;
}
private double recencyWeight(MemoryEntry e, Instant now) {
try {
double days = Duration.between(Instant.parse(e.timestamp()), now).toHours() / 24.0;
return 0.5 + 0.5 * Math.pow(0.5, days / halfLife.toDays()); // never below 0.5
} catch (RuntimeException unparsedOrMissing) {
return 0.5; // unknown age: treat as old, not as absent
}
}
static String text(MemoryEntry e) {
return e.content().parts().orElse(List.of()).stream()
.map(p -> p.text().orElse(""))
.collect(Collectors.joining(" ")).trim();
}
static String normalize(String s) {
return s.toLowerCase(Locale.ROOT).replaceAll("\\s+", " ");
}
}Several choices here are deliberate. Each source gets a timeout and an empty-list fallback, so a slow vector database degrades recall instead of failing the turn. Fusion uses reciprocal rank fusion, which scores an item by the sum of 1/(k + rank) over the sources that returned it. It needs only ranks, so it can combine a lexical score and a cosine similarity that are not on comparable scales. Deduplication on normalized text means a fact remembered in three sessions takes one slot. The budget is in characters, calibrated once against your tokenizer.
Scoping is not done in the decorator. appName and userId are passed to every source, and each source must filter on them inside its own query. If you fuse first and filter afterwards, one bug in the filter leaks another user's memories into the prompt.
Worked example: scoring three candidates
Suppose a user asks, "What did we decide about the Frankfurt deploy?", and the model calls the memory tool with the query "Frankfurt deploy decision". The lexical source and the vector source each return a ranked list, and three candidates matter. Use k = 60 and a 30-day half-life.
| Memory | Lexical rank | Vector rank | RRF score | Age | Recency weight | Final |
|---|---|---|---|---|---|---|
| A: "Chose eu-central for the Frankfurt deploy" | 1 | 3 | 1/61 + 1/63 = 0.0323 | 90 days | 0.5 + 0.5 x 0.125 = 0.5625 | 0.0181 |
| B: "Frankfurt office wants a demo" | - | 1 | 1/61 = 0.0164 | 2 days | 0.977 | 0.0160 |
| C: "Moved the Frankfurt deploy to next sprint" | 2 | 2 | 2/62 = 0.0323 | 1 day | 0.989 | 0.0319 |
C ranks first. It appears high in both lists and is recent. A ranks second, because agreement between two sources outweighs its age. B, a strong semantic match about the wrong subject, ranks last despite being top of the vector list. The floor of 0.5 on the recency weight is what keeps A in the answer. Without a floor, a 90-day-old decision would be scored at an eighth of its relevance, and the agent would forget decisions just because they were settled long ago.
A and C may conflict. Retrieval should return both with their timestamps and let the model reconcile them; silently dropping the older one produces confident, wrong answers.
Shaping the query
A few cheap transformations of the query improve recall more than index tuning does. Keep identifiers intact and route them to a lexical source, because embeddings match "4521" poorly. Expand known aliases from a small table, so "FRA" and "Frankfurt" match. On push, build the query from the last user message plus the current task, not the whole transcript. Add a model-based rewrite only if evaluation shows it helps: it costs a model call, and a hallucinated entity retrieves confidently wrong memories.
Measuring recall
A retrieval policy without a metric is tuned by anecdote. Build a small labelled set from real, consented sessions: the question the model would ask, and the IDs of memories a human judged relevant. Fifty to a hundred cases are enough to see large regressions. Then measure how often a relevant memory lands in the top k, and how high it lands:
# Offline recall evaluation: run against a copy of real, consented memory data.
cases = [
# (app, user, question the model would ask, ids of memories a human marked relevant)
("support-app", "user-17", "status of order 4521", {"m-903"}),
("support-app", "user-17", "which deploy region did we pick", {"m-411", "m-412"}),
]
def evaluate(search, cases, k=8):
hit_at_k, rr = 0, 0.0
for app, user, question, relevant in cases:
ids = [m.id for m in search(app, user, question)[:k]]
if relevant & set(ids):
hit_at_k += 1
first = min(ids.index(i) for i in relevant if i in ids)
rr += 1.0 / (first + 1)
return {"recall@k": hit_at_k / len(cases), "mrr": rr / len(cases)}Run it whenever you change the shaper, a source, the half-life or the budget. Tag each case by type (identifier, paraphrase, preference, stale fact), because an average can hide a change that fixed paraphrases and broke identifiers.
Failure modes
| Failure | Symptom | Mitigation |
|---|---|---|
| Scope filter applied after retrieval | Another user's memory appears in a response. | Filter on app and user inside every source query; test with two users who share vocabulary. |
| Unbounded results | Prompt tokens spike and answers drift off topic. | Enforce max entries and max characters in the policy, not in the prompt. |
| Identifiers ignored | "Order 4521" retrieves every order. | A lexical source that keeps digits; query shaping that preserves ids. |
| Stale facts win | The agent repeats a preference the user changed. | Recency weighting, return conflicts together, and update or delete facts at write time. |
| Memory poisoning | Text that a user or a tool once said, stored and later retrieved as an instruction. | Label memories as data in the prompt, and filter instruction-like text at ingestion. |
| Empty result invites invention | The model "remembers" things that were never stored. | Tell the model explicitly when recall returned nothing, and instruct it to say so. |
Operating it
Log every recall with its query, the IDs returned by each source, the fused order, what was cut by the budget and the latency per source. Watch the fraction of turns that call the memory tool; a sudden drop usually means an instruction change made the model stop asking. Connect these logs to your traces, as described in ADK Java observability, so one slow recall can be found inside one slow turn.
What to do next
- Replace
InMemoryMemoryServiceoutside tests, and write down which of the four decisions (when, query, how many, what form) your agent currently makes, and where. - Wrap your store in a decorator like
RankedMemoryServiceso the policy is one class with unit tests. - Add a lexical source that keeps digits next to your vector source, and fuse them with RRF.
- Set a hard entry and character budget, and a recency weight with a floor.
- Build a labelled recall set of at least fifty questions and record recall@k and MRR before you change anything else.
- Test scope with two users who share vocabulary, and assert that neither ever sees the other's memories.
- Decide on push, pull or a split. Implement push in a before-model callback only for a short list of high-value facts.