"Give the agent memory" is one of the most common requests in agent projects and one of the least precise. A model has no memory between calls at all: every request starts from nothing but the text you send. Everything an agent appears to remember is something your code chose to store and chose to put back into a later request. The questions that matter are therefore which facts to keep, for how long, for whom, in what form, and how they come back.
This article gives a working taxonomy of agent memory, organised by lifetime and by how the memory is read back, and maps every kind onto a concrete construct in ADK for Java. It ends with a routing rule you can put in code and a worked example. The details of each construct live on their own pages: state prefixes and session scope in session context in ADK Java, the memory service in ADK Java MemoryService architecture, and recall policy in long-term memory retrieval. This page owns the classification decision that comes before all of them.
Why a taxonomy is worth having
Teams that skip classification tend to end up with one of two designs. In the first, everything goes into the conversation history, the context fills, costs rise with every turn and the model starts ignoring material in the middle of long prompts. In the second, everything goes into a vector store, and exact facts such as a seat preference or an order number come back as fuzzy near-matches, or not at all. Both designs fail because they treat memory as one thing.
The useful axes are few. Lifetime: does the fact die with this model call, this turn, this session, or outlive all of them? Scope: does it belong to one session, one user, or everyone? Read pattern: is it looked up by an exact key, or found by searching for something similar? Mutability: is there one current value that replaces the old one, or an accumulating history? Classify a fact on those four axes and its home is usually obvious.
The kinds, from shortest-lived to longest
The names below follow common usage in the agent literature, which borrows them loosely from cognitive psychology. The borrowing is only an analogy; what matters is the engineering behaviour of each kind.
| Kind | Lifetime and scope | Read by | ADK Java home |
|---|---|---|---|
| Working memory | One model request | The model, directly | The request the framework builds: instruction, relevant session events, tool results |
| Scratch state | One invocation (turn) | Your callbacks and tools | State keys with the temp: prefix |
| Conversational | One session | The model, as history | The Session's event list |
| Session state | One session | Exact key | State keys without a prefix |
| User profile | Across sessions, one user | Exact key | State keys with the user: prefix |
| Long-term episodic and semantic | Across sessions, one user | Search | A BaseMemoryService implementation, reached through LoadMemoryTool or your own code |
| Shared knowledge | All users of an app | Exact key or search | app: state for small values; a retrieval tool for documents |
| Procedural | Until you redeploy | Always present | The agent's instruction, its tools and your code |
| Parametric | Until the model changes | Implicit | The model weights; not writable from an agent |
Two distinctions in the long-term row deserve care. Episodic memory stores what happened: past conversations or events, with their time. Semantic memory stores what is true: facts distilled from those episodes, such as "the user is vegetarian". ADK's memory service contract does not distinguish them. Its write path, addSessionToMemory(Session), accepts a whole finished session, which is episodic by default, and an implementation is free to extract facts from it instead. The reference InMemoryMemoryService stores the session's events as they are and searches them by word overlap, so it is purely episodic, unranked, and, because it only matches alphabetic words, cannot find a number such as an order id.
The routing rule
Every candidate memory, whether the model proposes it or your code extracts it, should pass through the same decision. Ask the questions in this order, because the earlier ones override the later ones.
- Is it sensitive (credentials, health data, payment details, anything you could not justify storing)? Drop it, or store it only in a system built for that data, not in agent memory.
- Would storing it change how the agent behaves for everyone or for all future turns, such as "always skip the confirmation step"? That is procedural memory. Route it to a human who edits the instruction or code, never straight into a store the agent reads back as authority.
- Is it needed in future sessions? If it is an exact, single-valued fact about this user, put it in user: state, where it is overwritten when it changes. Otherwise it belongs in long-term memory, found by search.
- Is it needed later in this session only? Session state.
- Otherwise it is scratch: temp: state, or simply a local variable.
// Our code, not part of ADK: decide where a candidate memory belongs before writing it.
public enum MemoryHome { DROP, TURN_SCRATCH, SESSION_STATE, USER_STATE, LONG_TERM, PROCEDURAL_REVIEW }
public record Candidate(String text, boolean exactValue, boolean aboutUser,
boolean neededLaterThisSession, boolean neededInFutureSessions,
boolean changesAgentBehaviour, boolean sensitive) {}
public final class MemoryRouter {
public static MemoryHome route(Candidate c) {
if (c.sensitive()) return MemoryHome.DROP; // never persist by default
if (c.changesAgentBehaviour()) return MemoryHome.PROCEDURAL_REVIEW; // a human edits instructions
if (c.neededInFutureSessions()) {
return (c.aboutUser() && c.exactValue())
? MemoryHome.USER_STATE // one current value, overwritten on change
: MemoryHome.LONG_TERM; // fuzzy, many, recalled by search
}
if (c.neededLaterThisSession()) return MemoryHome.SESSION_STATE;
return MemoryHome.TURN_SCRATCH;
}
}The classes above are ours, not ADK's. They exist to make the decision testable: you can feed in a table of example facts and assert where each one lands, which is far cheaper than discovering the answer from a confused agent in production. The booleans are usually filled in by a small extraction step, either rules or a model call with a strict output schema, run at the end of a turn or at the end of a session.
Writing each kind in ADK Java
Exact state is written through the context object your tool receives. A put on toolContext.state() is visible immediately in the same turn and is also recorded in the stateDelta of the event the step produces; the session service applies that delta by prefix when the event is appended. A write that never reaches an event is lost. Long-term memory is written by handing a finished session to the memory service.
// Our tool: store an exact, user-scoped preference. The key prefix decides the lifetime.
public static Map<String, Object> savePreference(
@Schema(name = "name", description = "Preference name, e.g. seat or diet") String name,
@Schema(name = "value", description = "The new value") String value,
@Schema(name = "toolContext") ToolContext toolContext) {
if (!ALLOWED_PREFS.contains(name)) { // a fixed vocabulary, not free text
return Map.of("status", "rejected", "reason", "unknown preference");
}
toolContext.state().put("user:pref_" + name, value); // recorded in this event's stateDelta
toolContext.state().put("user:pref_" + name + "_at", Instant.now().toString());
return Map.of("status", "ok");
}
// After the conversation ends: hand the finished session to long-term memory.
Session done = sessionService
.getSession("travel", userId, sessionId, Optional.empty())
.blockingGet(); // Maybe: null if it no longer exists
if (done != null) {
memoryService.addSessionToMemory(done).blockingAwait(); // write path of BaseMemoryService
}Note the fixed vocabulary in the tool. Letting the model invent state keys produces near-duplicates such as user:diet, user:dietary_pref and user:food, each holding a different answer. A small allowed list keeps exact memory exact. Storing the time of each write alongside the value is cheap and lets you resolve conflicts with long-term memory later, because a profile value and a remembered episode can disagree.
Reading follows the same split. Exact state is read by key in a tool or callback, or written into the request by your own code. Long-term memory is pulled by the model through LoadMemoryTool, which calls ToolContext.searchMemory(query) for the current app and user, or pushed by your own code before a model call. Which of pull or push to use, and how to rank results, is the subject of the retrieval article; the point here is that exact facts should not need either, because they are not searched for at all.
Worked example: one conversation, eight facts
A travel assistant has this conversation with a returning user. The table shows each fact that appears, how it classifies, and where it goes.
| Fact in the conversation | Lifetime and read pattern | Home |
|---|---|---|
| "I'm vegetarian" | Future sessions, exact, single value | user:pref_diet |
| "Window seat, always" | Future sessions, exact, single value | user:pref_seat |
| Booking reference AB12CD for this trip | This session, exact | Session state booking_ref |
| Raw JSON from the flight search API | This turn only | temp: key, or a local variable |
| "Last time the Lisbon hotel was too noisy" | Future sessions, fuzzy, recalled when relevant | Long-term memory, via the finished session |
| The user's passport number | Sensitive | Dropped; the booking system holds it |
| "Stop asking me to confirm prices" | Changes agent behaviour | Procedural review; at most a user-scoped preference the instruction explicitly honours |
| Visa rules for Portugal | Shared, changes over time | A retrieval tool over maintained documents, not memory |
Three outcomes are worth noticing. The diet and seat preferences never touch the vector store, so they come back exactly, every time, at no search cost. The noisy-hotel remark does go to long-term memory, because it is only useful when a future question is about hotels in Lisbon, which is exactly the situation search is good at. And the request to stop confirming prices is not stored as free text that the agent will later read as an instruction; if the business agrees to it, it becomes a named preference that the instruction already knows how to interpret. That last rule is also a security control, discussed below.
Costs: what each kind charges you
Every kind of memory has a cost profile, and the right mix depends on it. Working memory is paid in tokens on every model call, so a session that replays its whole event history grows more expensive and slower with every turn; long sessions need summarisation or windowing. Exact state is nearly free to read but is only as good as the key discipline behind it. Long-term memory costs a search on each recall, embedding and storage on each write, and, more expensively, the context tokens of whatever it returns. A recall that returns ten loosely related memories can cost more than it helps. Procedural memory costs nothing at run time but is changed only through review and deployment, which is the point.
Deletion has a cost too. A user who asks to be forgotten has data in at least three places: their user: state, every session the session service still holds, and every entry the memory service extracted from those sessions. If you cannot enumerate those places from your design, you cannot honour the request. The schema choices for a durable session store are discussed in the Postgres session table schema.
Failure modes
- Everything in history. The context grows until the model loses track of early facts or the request exceeds the window. Fix: move exact facts to state and summarise or window old events.
- Exact facts in a vector store. "Order 4521" retrieves order 4512, or nothing, because the reference in-memory service ignores digits and embeddings blur numbers. Fix: exact facts go in state, keyed.
- Contradictions across stores. user:pref_seat says aisle, a remembered episode says window. Fix: state is the current truth, episodes are history; label recalled memories with their date so the model can weigh them, and prefer the newer timestamp.
- Memory poisoning. Text a user or a tool result put into memory comes back later and is followed as an instruction, possibly in another session. Fix: never route behaviour-changing content to memory, and frame recalled text in the request as quoted past content, not as instructions.
- Leakage across users. A custom memory service that forgets to filter by app and user returns another user's history. Fix: enforce the filter in the store's query and test it, as in semantic memory on pgvector.
- Relying on temp: across turns. Its handling differs between the base service and the in-memory one. Fix: treat temp: values as gone when the turn ends.
What to do next
- List every fact your agent needs to remember and classify each on lifetime, scope, read pattern and mutability.
- Write the routing rule as code with a table-driven unit test, using real examples from your transcripts.
- Move every exact, single-valued user fact into user: state with a fixed key vocabulary and a written-at timestamp.
- Replace the in-memory memory service before production with one that ranks results, matches numbers and filters by app and user.
- Decide pull or push recall per kind of fact, and cap how many tokens recalled memory may add.
- Route behaviour-changing requests to human review, and treat recalled text as untrusted content.
- Document every place a user's data lives so that deletion is a script, not an investigation.