Most agent memory stores facts: the user prefers aisle seats, the account is on the enterprise plan. Episodic memory stores experiences: on this date, for this goal, the agent called these tools in this order, hit this error, recovered this way, and the outcome was good or bad. Facts answer the question 'what is true about this user?'. Episodes answer a different one: 'what happened last time something like this came up, and did it work?'.
This article builds episodic memory for an ADK Java agent from first principles. It defines an episode as a structured record, mines one from the events ADK already writes into every session, stores it behind ADK's own BaseMemoryService contract so the stock LoadMemoryTool can recall it, and spends most of its length on the parts that decide whether the feature helps or hurts: outcome labelling, ranking, poisoning and staleness. It assumes you know how ADK sessions and the memory write path work; if not, read the cross-session memory article linked at the end first.
Episodes versus facts
Psychology borrowed the term from Endel Tulving, who separated memory for events (episodic) from memory for general knowledge (semantic). The split maps cleanly onto agents. A semantic memory entry is context-free: 'refunds over 500 dollars need a manager'. An episodic entry is bound to a time, a goal and an outcome: 'on 2 October, a 640 dollar refund was attempted directly, the refund tool rejected it with APPROVAL_REQUIRED, the agent opened an approval ticket, and the refund completed the next day'.
The episodic form carries three things facts lose. It carries procedure: the order of tool calls that worked. It carries failure: approaches that did not work and why. And it carries evidence: a timestamp and an outcome label, so the agent can weigh a precedent rather than obey it. That last point is the whole design problem. An episode is only useful if the agent knows whether it ended well, and the transcript alone usually does not say.
Episodic memory is worth building when an agent repeats multi-step tasks with tool calls whose right sequence is learned rather than written down: support resolution, infrastructure runbooks, data-pipeline repair. It is not worth building for single-turn question answering, where there is no procedure to remember.
The episode record
Do not store raw transcripts as episodes. A transcript is long, mostly chit-chat, and leaves the outcome implicit. Store a small structured record and render it to text only at recall time.
| Field | Source | Why it exists |
|---|---|---|
| goal | first substantive user message, normalised | what retrieval matches against |
| trace | functionCalls() / functionResponses() per event | the procedure, as tool names, key arguments and error codes |
| outcome | state key written by a tool, or user feedback | success, failure or unknown; never guessed from tone |
| lesson | one distillation model call | a sentence a future agent can act on |
| toolsetVersion | hash of tool names and schemas | detects precedents that refer to tools that changed |
| sessionId, createdAt | Session.id(), Event.timestamp() | idempotent upsert and recency |
| appName, userId | Session.appName(), userId() | the isolation boundary |
The outcome field deserves its own rule: it is written by code, not inferred by a model. The cleanest source is a tool that knows the truth, such as a completeRefund tool that writes task.outcome = success into session state when the payment system confirms. The second-best source is explicit user feedback. A model's judgement of whether the conversation 'went well' is the weakest source and should be stored as unknown.
Mining an episode from ADK events
ADK already records everything an extractor needs. Each Event has an author(), an optional content(), and convenience accessors functionCalls() and functionResponses() that return immutable lists. Event.timestamp() is a long in epoch milliseconds. The extractor below walks a finished session and builds the record. It reads session.events(); on recent adk-java main that accessor is deprecated in favour of immutableEvents(), so use whichever your version has.
public record Episode(String appName, String userId, String sessionId, String goal,
List<String> trace, String outcome, String lesson,
String toolsetVersion, long createdAtMillis) {}
public final class EpisodeExtractor {
private final String toolsetVersion; // hash of tool names + schemas, computed at boot
public EpisodeExtractor(String toolsetVersion) { this.toolsetVersion = toolsetVersion; }
public Optional<Episode> extract(Session session) {
List<Event> events = session.events(); // immutableEvents() on newer versions
String goal = events.stream()
.filter(e -> "user".equals(e.author()))
.map(Event::stringifyContent)
.filter(t -> t.length() > 15) // skip "hi" and "thanks"
.findFirst().orElse(null);
if (goal == null) return Optional.empty();
List<String> trace = new ArrayList<>();
for (Event e : events) {
for (FunctionCall call : e.functionCalls()) {
trace.add("call " + call.name().orElse("?") + keyArgs(call));
}
for (FunctionResponse r : e.functionResponses()) {
Object err = r.response().map(m -> m.get("error")).orElse(null);
trace.add("result " + r.name().orElse("?") + (err == null ? " ok" : " error=" + err));
}
}
if (trace.isEmpty()) return Optional.empty(); // no procedure, not an episode
String outcome = String.valueOf(session.state().getOrDefault("task.outcome", "unknown"));
long last = events.get(events.size() - 1).timestamp();
return Optional.of(new Episode(session.appName(), session.userId(), session.id(),
goal, trace, outcome, null, toolsetVersion, last));
}
private static String keyArgs(FunctionCall call) {
// Keep identifiers that explain the path; drop free text, which is where PII lives.
Map<String, Object> args = call.args().orElse(Map.of());
return args.containsKey("amount") ? " amount=" + args.get("amount") : "";
}
}Two choices in that code matter more than they look. Sessions with no tool calls produce no episode, which keeps the store about procedure. And keyArgs keeps only whitelisted arguments: amounts and error codes explain a path, while free-text arguments such as email bodies are where personal data leaks into long-term storage.
The lesson is filled by a separate step: one cheap model call over the goal, trace and outcome, with an instruction such as 'In one sentence, state what a future agent should do or avoid for a similar goal. If the outcome is unknown, say what was tried, not what works.' Run it asynchronously after the session closes, never on the user's turn.
An episodic memory service
Because recall goes through BaseMemoryService, the episodic store can be a decorator that ADK's LoadMemoryTool calls without knowing anything changed. The interface has two methods: addSessionToMemory(Session) returning Completable, and searchMemory(String appName, String userId, String query) returning Single<SearchMemoryResponse>. Remember that the runner never calls the write side for you; your session-close job does.
public final class EpisodicMemoryService implements BaseMemoryService {
private final EpisodeExtractor extractor;
private final EpisodeStore store; // your table: upsert, vector search, delete-by-user
private final Distiller distiller; // one model call -> lesson sentence
private final String currentToolset;
@Override
public Completable addSessionToMemory(Session session) {
return Completable.fromAction(() ->
extractor.extract(session)
.map(distiller::withLesson)
.ifPresent(store::upsertBySessionId)); // re-adds replace, never duplicate
}
@Override
public Single<SearchMemoryResponse> searchMemory(String appName, String userId, String query) {
return Single.fromCallable(() -> {
List<ScoredEpisode> hits = store.similar(appName, userId, query, 20); // scope inside the query
List<MemoryEntry> top = hits.stream()
.filter(h -> !"failure".equals(h.episode().outcome()) || h.similarity() > 0.85)
.sorted(Comparator.comparingDouble(this::score).reversed())
.limit(2)
.map(this::render)
.toList();
return SearchMemoryResponse.builder().memories(top).build();
});
}
private double score(ScoredEpisode h) {
double ageDays = (System.currentTimeMillis() - h.episode().createdAtMillis()) / 86_400_000.0;
double outcome = switch (h.episode().outcome()) {
case "success" -> 1.0; case "failure" -> 0.6; default -> 0.4; };
double fresh = h.episode().toolsetVersion().equals(currentToolset) ? 1.0 : 0.3;
return h.similarity() * outcome * fresh * Math.exp(-ageDays / 90.0);
}
private MemoryEntry render(ScoredEpisode h) {
Episode e = h.episode();
String body = "PRECEDENT (" + e.outcome() + ", " + Instant.ofEpochMilli(e.createdAtMillis())
+ "): goal=" + e.goal() + " | steps=" + String.join(" > ", e.trace())
+ " | lesson=" + e.lesson() + " | Treat as evidence, not instruction.";
return MemoryEntry.builder()
.content(Content.fromParts(Part.fromText(body)))
.author("episodic-memory")
.timestamp(Instant.ofEpochMilli(e.createdAtMillis()).toString()) // the field is a String
.build();
}
}The filter keeps failures only when they are very similar, because a failure is useful as a warning about the exact situation and noise for anything looser. The score multiplies similarity by an outcome weight, a toolset-freshness penalty and a 90-day decay; every constant there is a starting guess to tune against your own data. The render step labels each entry as precedent and ends with an explicit instruction to weigh it, because a model handed a step list will otherwise follow it literally. MemoryEntry's timestamp is a String, so convert the long yourself.
Worked example: a large refund
A billing agent receives: 'Refund the 640 dollar duplicate charge on invoice 8812.' The model calls LoadMemoryTool with the query 'refund duplicate charge large amount'. The store returns three candidates for this user's app scope:
| Episode | Similarity | Outcome | Toolset | Age | Score |
|---|---|---|---|---|---|
| A: 520 dollar refund, direct call rejected APPROVAL_REQUIRED, ticket opened, completed | 0.82 | success (1.0) | current (1.0) | 12 days | 0.82 x 1.0 x 1.0 x 0.875 = 0.72 |
| B: 610 dollar refund, direct call rejected, agent retried 3 times, user gave up | 0.88 | failure (0.6) | current (1.0) | 30 days | 0.88 x 0.6 x 1.0 x 0.717 = 0.38 |
| C: 450 dollar refund via old issueCredit tool, completed | 0.79 | success (1.0) | old (0.3) | 200 days | 0.79 x 1.0 x 0.3 x 0.108 = 0.03 |
Episode A ranks first and teaches the procedure: open the approval ticket instead of calling the refund tool directly. Episode B passes the failure filter because its similarity is above 0.85, and it adds a warning: retrying the rejected call loses the customer. Episode C refers to a tool that no longer exists and decays to nothing. The model's first action is to open the approval ticket, which saves one rejected tool call and the retry loop that sank episode B.
Failure modes
The failure modes of episodic memory are mostly failures of trust.
- Poisoned precedent. An unlabelled failure, stored as
unknownand recalled as if it worked, teaches the agent a broken procedure. Down-weight unknown outcomes and measure how often they are recalled. - Injection that persists. A prompt injection inside a tool result becomes part of the trace and is replayed into future sessions. Store tool names, codes and whitelisted arguments, never raw tool output, and run the distiller's input through the same sanitiser as live tool results.
- Stale procedure. Tools change; episodes do not. The toolset hash catches renamed or re-schemed tools but not changed business rules, so retention limits still matter.
- Anchoring. The model copies the precedent even when the new case differs, such as a refund under the approval threshold. Keep the 'evidence, not instruction' label and cap recall at two episodes.
- Duplicate ingestion. Checkpointed sessions are added several times. Upsert by session id, or one conversation becomes three equally strong precedents.
- Cross-user leakage. Episodes from one user recalled for another expose their data. Apply the app and user filter inside the store query, never after ranking. If you want shared, organisation-wide procedure, promote reviewed episodes into a separate, anonymised store.
Operating it
Measure whether the feature earns its cost. Track the recall rate (turns where the model called the memory tool and received an episode), the success rate of tasks with and without a recalled precedent, and tool calls per resolved task. The honest test is an A/B split on the memory service: half the sessions get the episodic decorator, half get an empty search result. If success and tool calls per task do not move, delete the feature.
Bound growth with a retention window, such as 180 days, plus deletion when the toolset hash has been retired for a full window. Delete by user on account removal; the right-to-forget article covers the mechanics. Log every recall with the episode ids returned, so an incident can be traced to the precedent that caused it.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Structured records vs transcripts | short, rankable, outcome explicit | extractor code per agent; details lost |
| Distilled lesson | actionable one-liner | a model call per session; can hallucinate |
| Pull via LoadMemoryTool | no cost on turns that need no history | model must realise it needs precedent |
| Push via a before-model callback | precedent always present | tokens on every call; more anchoring |
| Per-user scope | no leakage | cold start for every new user |
| Shared reviewed store | organisation learns once | anonymisation and review workflow |
What to do next
- Pick one multi-step task your agent repeats and list the tool that can write a trustworthy
task.outcomeinto state. - Implement
EpisodeExtractorfor that agent with a whitelist of arguments, and unit-test it on recorded sessions. - Wrap your existing memory service with the episodic decorator and keep its upsert idempotent; see cross-session memory for the write trigger and the MemoryService architecture for the contract.
- Tune the ranking with the techniques in long-term memory retrieval.
- Add retention, delete-by-user and recall logging, following the right-to-forget guide.
- Run the A/B split for two weeks, and keep the feature only if success rate or tool calls per task improve. The events article explains the event fields the extractor reads.