Computer memory is a hierarchy: registers, caches, RAM and disk, each larger, slower and cheaper than the one above, with hardware and the operating system moving data between levels. Agent memory has the same shape. Tokens in the context window are the fastest and most expensive. Session state is cheap to read but scoped to one conversation. A memory service holds everything, but every read costs a search and may return the wrong thing. Raw archives hold the full record and are almost never read.
Most agent memory bugs are hierarchy bugs: a fact held in two tiers that disagree, a profile that grows until it crowds out the conversation, a recall that duplicates what is already in the prompt. This article treats ADK Java memory as a hierarchy you design on purpose. It defines the tiers, gives each a budget, and sets the rules that move items up and down and keep the copies consistent. For the kinds of memory (episodic, semantic, procedural) see the memory taxonomy. This page is about where items live and how they move.
Five tiers
Five tiers cover almost every ADK Java agent. The state prefixes are explained in session context in depth. Here they matter only as tier boundaries.
| Tier | ADK mechanism | Scope | Read cost | Capacity |
|---|---|---|---|---|
| T0 Context window | Instruction plus session events in the request | One model call | Tokens on every call | Model limit, minus output |
| T1 Session state | temp: and unprefixed keys | Invocation or session | Map lookup | Small: kilobytes |
| T2 User and app state | user: and app: keys | All of a user's or app's sessions | Map lookup, injected as tokens | Small, pinned |
| T3 Memory service | BaseMemoryService.searchMemory | User, across sessions | A search, plus tokens for hits | Large |
| T4 Archive | Stored sessions, artifacts, event export | Everything | Batch jobs only | Unbounded |
Two properties decide where an item belongs. The first is how often the next turn needs it: constantly (the user's name, the task in hand), sometimes (last month's order) or almost never (a transcript from a year ago). The second is how expensive it is to be wrong about having it. A dietary restriction that the agent fails to recall has a high cost, so it is pinned in T2 even though it is rarely needed. A small talk detail with a low cost can wait in T3 until a search finds it.
Two tiers need a caution. How long temp: keys survive depends on the session service and ADK version, so check yours before treating T1 as more than a scratchpad for the current invocation. Keep anything that must last the whole session under an unprefixed key. T4 is not useless just because the hot path never reads it. It is the source you rebuild T3 from when you change the extraction prompt or the embedding model, the corpus for evaluation sets, and the record you consult when a user disputes what the agent remembered. Keep it complete, cheap and out of the request path.
The read path and its budgets
On each model call the agent reads the hierarchy top-down, and each tier gets a token budget. The window holds the instruction and recent turns. A before-model callback adds a pinned block from T2 and, if the turn warrants it, recalled items from T3. T4 is never read on the hot path. The numbers below are an example budget for a 32k-token window, not measurements. Size your own from traces, as in short-term memory and the context window.
| Slot | Source tier | Example budget | Loaded |
|---|---|---|---|
| Instruction and tool declarations | Agent definition | 4,000 | Always |
| Pinned profile | T2 | 600 | Always, capped |
| Session summary | T1 | 1,200 | After compaction |
| Recent turns | T0 events | 18,000 | Always, newest first |
| Recalled memories | T3 | 2,000 | When the turn needs history |
| Headroom for output and tool results | 6,200 |
static Maybe<LlmResponse> assemble(CallbackContext ctx, LlmRequest.Builder req) {
List<String> blocks = new ArrayList<>();
blocks.add(Pinned.render(ctx.state(), 600)); // T2, capped by tokens
Object summary = ctx.state().get("session_summary"); // T1
if (summary instanceof String s) blocks.add("Earlier in this conversation: " + s);
req.appendInstructions(blocks); // pinned tiers first
String query = Recall.queryFor(ctx); // empty if no recall needed
if (query.isEmpty()) return Maybe.empty();
InvocationContext inv = ctx.invocationContext();
if (inv.memoryService() == null) return Maybe.empty(); // runner built without memory
return inv.memoryService().searchMemory(inv.appName(), inv.userId(), query)
.map(r -> Recall.render(r.memories(), Pinned.keys(ctx.state()), 2000)) // T3
.doOnSuccess(recalled -> req.appendInstructions(List.of(recalled)))
.ignoreElement()
.onErrorComplete() // recall is best effort
.andThen(Maybe.<LlmResponse>empty());
}Two lines carry the hierarchy logic. Recall.queryFor decides whether T3 is worth a search on this turn, so a greeting does not trigger one. Recall.render receives the pinned keys so it can drop recalled items that duplicate T2, which is the agent equivalent of not caching a line twice. Recall failure is treated as a cache miss, not an error: the turn proceeds without it, and the pinned blocks were appended before the search so a failed recall cannot lose them.
Promotion and demotion
Promotion moves an item to a faster tier because it is being used. Demotion moves it down because the faster tier is full or the item has gone cold. Write both as explicit rules, run by your code at known points, never as side effects of a prompt.
| Move | Trigger | Rule |
|---|---|---|
| T3 to T2 | Recalled in 3 of the last 5 sessions | Pin under a typed user: key, within the profile cap |
| T0 to T1 | Window passes 70% of budget | Summarize the oldest turns into session_summary |
| T1 to T3 | Session idle or closed | Call addSessionToMemory with the session |
| T2 to T3 | Profile over its cap | Unpin the least recently used key; it stays searchable |
| T3 to T4 | Not recalled for a year | Move to the archive; remove it from the search index |
The session-to-memory step deserves emphasis because ADK Java does not do it for you. Nothing in the framework calls addSessionToMemory. Your application decides when a session is finished, usually after an idle timeout, and calls it, ideally through a job that also writes the T1 summary, as in memory summarization strategies. Promotion into T2 needs a cap and a typed key set. A profile that any turn can write to grows until it is the largest block in the prompt.
Write policy and consistency
When a fact changes, which tier is written first? Caches answer this with write-through (update every level now) and write-back (update the fast level, flush later). Agents need a rule per kind of item:
- Pinned facts: write-through. When the user changes their dietary restriction, update the
user:key in the same event and write the change to the memory service straight away. A stale T3 copy would be recalled next to the new pinned value and contradict it. - Conversation content: write-back. Turns live in T0 and T1 and reach T3 when the session is consolidated. Writing every turn to the memory service doubles cost and indexes half-finished thoughts.
- One owner per fact. Each fact has exactly one authoritative tier. Other tiers hold copies stamped with the version they were copied from. On conflict, the owner wins, and rendering code drops a copy older than the owner's version.
// T2 is the owner for pinned keys; T3 copies carry the version they saw.
void setPinned(State state, String key, String value, MemoryWriter t3) {
long version = System.currentTimeMillis();
state.put("user:" + key, Map.of("value", value, "v", version)); // owner, same event
t3.upsertFact(key, value, version); // write-through copy
}
boolean superseded(MemoryEntry hit, Map<String, Long> pinnedVersions) {
return Facts.keyOf(hit).map(k -> pinnedVersions.getOrDefault(k, 0L) > Facts.versionOf(hit))
.orElse(false);
}Because MemoryEntry carries only content, author and timestamp, the key and version have to travel inside the content, for example as a short structured header your own service writes and Facts parses. That is a limitation of the 1.11.0 type worth designing around early.
One fact's journey
Look at the diagram from the point of view of one fact. It is born in T0 as words in a turn, may be noted in T1 during the session, is consolidated into T3 when the session closes, is promoted to T2 if it keeps being recalled, is demoted back to T3 if the profile fills, and ends in T4. At every point exactly one tier owns it.
Worked example: a returning support user
Consider a support agent for a software product. The user's plan tier and region are pinned in T2 (about 40 tokens) because every answer depends on them. In a long troubleshooting session the window passes 70% of budget at turn 30, and the oldest 20 turns become an 800-token summary in T1. The session closes after an hour idle, and a job consolidates it into T3 as an episode: the symptom, the fix that worked and the version.
Three weeks later the user returns with a similar error. The pinned block loads as usual. The recall query finds the episode, and because the user's plan is already pinned, Recall.render drops the copy of the plan from the recalled text. That saves tokens and, more importantly, avoids showing the model two plan values if the user has upgraded since. The user did upgrade, and the T2 write-through stamped a newer version, so the older plan text in the episode is marked as superseded. The episode is recalled in three consecutive sessions, so its fix is promoted to a pinned user:known_workaround key. At the next profile overflow, that key is the least recently used and is unpinned again.
Failure modes
- Two truths. A fact in T2 and an older copy in T3 both reach the prompt. Give each fact one owner and filter superseded copies.
- Profile bloat. T2 grows with every turn's "useful" detail. Use typed keys, a token cap and LRU demotion.
- Recall on every turn. Search runs for greetings and confirmations, which adds latency and noise. Gate T3 reads on the turn's need.
- Recall crowds out the conversation. Without a budget, ten long hits push recent turns out. Cap recalled tokens below the recent-turns slot.
- Nothing reaches T3. No code calls
addSessionToMemory, so cross-session recall is empty. Add a consolidation job and alert on zero writes. - Promotion storms. A popular item is promoted and demoted every session. Add hysteresis, such as promoting at 3 of 5 sessions and demoting at 0 of 10.
- Cold start after a rebuild. Re-indexing T3 from the archive resets recall counts, so promotion rules see every item as cold and demote the whole profile. Persist usage counters outside the index, and pause demotion while a rebuild runs.
Trade-offs
Every tier above T3 costs tokens on every call, so pinning trades money and window space for reliability. Every T3 read costs latency and risks irrelevant hits, so recall trades accuracy for breadth. Write-through keeps tiers consistent but doubles writes for pinned facts. Write-back is cheap but leaves a window in which T3 is stale. A vector store makes T3 recall better but adds an embedding step and an index to keep consistent, as described in semantic memory with vector stores. Most agents do well with a small pinned profile, write-back for conversation, write-through for pinned facts and gated recall.
What to do next
- Draw your agent's tiers and list, for each, the ADK mechanism, scope, owner and token budget.
- Measure a week of traces and set the per-slot budgets from real request sizes.
- Define the typed key set for
user:state and enforce a token cap in code. - Write a single before-model callback that assembles pinned, summary and recalled blocks within their budgets and de-duplicates against pinned keys.
- Add the consolidation job that calls
addSessionToMemorywhen sessions go idle, and alert when it writes nothing for a day. - Stamp versions on pinned facts and drop superseded copies at render time.
- Write promotion and demotion rules with hysteresis, and log every move.