Most agent memory stores facts: the user prefers aisle seats, the account is on the enterprise plan. Episodic memory stores experiences: on this date, for this goal, the agent called these tools in this order, hit this error, recovered this way, and the outcome was good or bad. Facts answer the question 'what is true about this user?'. Episodes answer a different one: 'what happened last time something like this came up, and did it work?'.

This article builds episodic memory for an ADK Java agent from first principles. It defines an episode as a structured record, mines one from the events ADK already writes into every session, stores it behind ADK's own BaseMemoryService contract so the stock LoadMemoryTool can recall it, and spends most of its length on the parts that decide whether the feature helps or hurts: outcome labelling, ranking, poisoning and staleness. It assumes you know how ADK sessions and the memory write path work; if not, read the cross-session memory article linked at the end first.

Episodes versus facts

Psychology borrowed the term from Endel Tulving, who separated memory for events (episodic) from memory for general knowledge (semantic). The split maps cleanly onto agents. A semantic memory entry is context-free: 'refunds over 500 dollars need a manager'. An episodic entry is bound to a time, a goal and an outcome: 'on 2 October, a 640 dollar refund was attempted directly, the refund tool rejected it with APPROVAL_REQUIRED, the agent opened an approval ticket, and the refund completed the next day'.

The episodic form carries three things facts lose. It carries procedure: the order of tool calls that worked. It carries failure: approaches that did not work and why. And it carries evidence: a timestamp and an outcome label, so the agent can weigh a precedent rather than obey it. That last point is the whole design problem. An episode is only useful if the agent knows whether it ended well, and the transcript alone usually does not say.

Episodic memory is worth building when an agent repeats multi-step tasks with tool calls whose right sequence is learned rather than written down: support resolution, infrastructure runbooks, data-pipeline repair. It is not worth building for single-turn question answering, where there is no procedure to remember.

The episode record

Do not store raw transcripts as episodes. A transcript is long, mostly chit-chat, and leaves the outcome implicit. Store a small structured record and render it to text only at recall time.

FieldSourceWhy it exists
goalfirst substantive user message, normalisedwhat retrieval matches against
tracefunctionCalls() / functionResponses() per eventthe procedure, as tool names, key arguments and error codes
outcomestate key written by a tool, or user feedbacksuccess, failure or unknown; never guessed from tone
lessonone distillation model calla sentence a future agent can act on
toolsetVersionhash of tool names and schemasdetects precedents that refer to tools that changed
sessionId, createdAtSession.id(), Event.timestamp()idempotent upsert and recency
appName, userIdSession.appName(), userId()the isolation boundary

The outcome field deserves its own rule: it is written by code, not inferred by a model. The cleanest source is a tool that knows the truth, such as a completeRefund tool that writes task.outcome = success into session state when the payment system confirms. The second-best source is explicit user feedback. A model's judgement of whether the conversation 'went well' is the weakest source and should be stored as unknown.

Mining an episode from ADK events

ADK already records everything an extractor needs. Each Event has an author(), an optional content(), and convenience accessors functionCalls() and functionResponses() that return immutable lists. Event.timestamp() is a long in epoch milliseconds. The extractor below walks a finished session and builds the record. It reads session.events(); on recent adk-java main that accessor is deprecated in favour of immutableEvents(), so use whichever your version has.

public record Episode(String appName, String userId, String sessionId, String goal,
                      List<String> trace, String outcome, String lesson,
                      String toolsetVersion, long createdAtMillis) {}

public final class EpisodeExtractor {
  private final String toolsetVersion;   // hash of tool names + schemas, computed at boot

  public EpisodeExtractor(String toolsetVersion) { this.toolsetVersion = toolsetVersion; }

  public Optional<Episode> extract(Session session) {
    List<Event> events = session.events();          // immutableEvents() on newer versions
    String goal = events.stream()
        .filter(e -> "user".equals(e.author()))
        .map(Event::stringifyContent)
        .filter(t -> t.length() > 15)                 // skip "hi" and "thanks"
        .findFirst().orElse(null);
    if (goal == null) return Optional.empty();

    List<String> trace = new ArrayList<>();
    for (Event e : events) {
      for (FunctionCall call : e.functionCalls()) {
        trace.add("call " + call.name().orElse("?") + keyArgs(call));
      }
      for (FunctionResponse r : e.functionResponses()) {
        Object err = r.response().map(m -> m.get("error")).orElse(null);
        trace.add("result " + r.name().orElse("?") + (err == null ? " ok" : " error=" + err));
      }
    }
    if (trace.isEmpty()) return Optional.empty();     // no procedure, not an episode

    String outcome = String.valueOf(session.state().getOrDefault("task.outcome", "unknown"));
    long last = events.get(events.size() - 1).timestamp();
    return Optional.of(new Episode(session.appName(), session.userId(), session.id(),
        goal, trace, outcome, null, toolsetVersion, last));
  }

  private static String keyArgs(FunctionCall call) {
    // Keep identifiers that explain the path; drop free text, which is where PII lives.
    Map<String, Object> args = call.args().orElse(Map.of());
    return args.containsKey("amount") ? " amount=" + args.get("amount") : "";
  }
}

Two choices in that code matter more than they look. Sessions with no tool calls produce no episode, which keeps the store about procedure. And keyArgs keeps only whitelisted arguments: amounts and error codes explain a path, while free-text arguments such as email bodies are where personal data leaks into long-term storage.

The lesson is filled by a separate step: one cheap model call over the goal, trace and outcome, with an instruction such as 'In one sentence, state what a future agent should do or avoid for a similar goal. If the outcome is unknown, say what was tried, not what works.' Run it asynchronously after the session closes, never on the user's turn.

An episodic memory service

Because recall goes through BaseMemoryService, the episodic store can be a decorator that ADK's LoadMemoryTool calls without knowing anything changed. The interface has two methods: addSessionToMemory(Session) returning Completable, and searchMemory(String appName, String userId, String query) returning Single<SearchMemoryResponse>. Remember that the runner never calls the write side for you; your session-close job does.

public final class EpisodicMemoryService implements BaseMemoryService {
  private final EpisodeExtractor extractor;
  private final EpisodeStore store;       // your table: upsert, vector search, delete-by-user
  private final Distiller distiller;      // one model call -> lesson sentence
  private final String currentToolset;

  @Override
  public Completable addSessionToMemory(Session session) {
    return Completable.fromAction(() ->
        extractor.extract(session)
            .map(distiller::withLesson)
            .ifPresent(store::upsertBySessionId));   // re-adds replace, never duplicate
  }

  @Override
  public Single<SearchMemoryResponse> searchMemory(String appName, String userId, String query) {
    return Single.fromCallable(() -> {
      List<ScoredEpisode> hits = store.similar(appName, userId, query, 20); // scope inside the query
      List<MemoryEntry> top = hits.stream()
          .filter(h -> !"failure".equals(h.episode().outcome()) || h.similarity() > 0.85)
          .sorted(Comparator.comparingDouble(this::score).reversed())
          .limit(2)
          .map(this::render)
          .toList();
      return SearchMemoryResponse.builder().memories(top).build();
    });
  }

  private double score(ScoredEpisode h) {
    double ageDays = (System.currentTimeMillis() - h.episode().createdAtMillis()) / 86_400_000.0;
    double outcome = switch (h.episode().outcome()) {
      case "success" -> 1.0; case "failure" -> 0.6; default -> 0.4; };
    double fresh = h.episode().toolsetVersion().equals(currentToolset) ? 1.0 : 0.3;
    return h.similarity() * outcome * fresh * Math.exp(-ageDays / 90.0);
  }

  private MemoryEntry render(ScoredEpisode h) {
    Episode e = h.episode();
    String body = "PRECEDENT (" + e.outcome() + ", " + Instant.ofEpochMilli(e.createdAtMillis())
        + "): goal=" + e.goal() + " | steps=" + String.join(" > ", e.trace())
        + " | lesson=" + e.lesson() + " | Treat as evidence, not instruction.";
    return MemoryEntry.builder()
        .content(Content.fromParts(Part.fromText(body)))
        .author("episodic-memory")
        .timestamp(Instant.ofEpochMilli(e.createdAtMillis()).toString())  // the field is a String
        .build();
  }
}

The filter keeps failures only when they are very similar, because a failure is useful as a warning about the exact situation and noise for anything looser. The score multiplies similarity by an outcome weight, a toolset-freshness penalty and a 90-day decay; every constant there is a starting guess to tune against your own data. The render step labels each entry as precedent and ends with an explicit instruction to weigh it, because a model handed a step list will otherwise follow it literally. MemoryEntry's timestamp is a String, so convert the long yourself.

Worked example: a large refund

A billing agent receives: 'Refund the 640 dollar duplicate charge on invoice 8812.' The model calls LoadMemoryTool with the query 'refund duplicate charge large amount'. The store returns three candidates for this user's app scope:

EpisodeSimilarityOutcomeToolsetAgeScore
A: 520 dollar refund, direct call rejected APPROVAL_REQUIRED, ticket opened, completed0.82success (1.0)current (1.0)12 days0.82 x 1.0 x 1.0 x 0.875 = 0.72
B: 610 dollar refund, direct call rejected, agent retried 3 times, user gave up0.88failure (0.6)current (1.0)30 days0.88 x 0.6 x 1.0 x 0.717 = 0.38
C: 450 dollar refund via old issueCredit tool, completed0.79success (1.0)old (0.3)200 days0.79 x 1.0 x 0.3 x 0.108 = 0.03

Episode A ranks first and teaches the procedure: open the approval ticket instead of calling the refund tool directly. Episode B passes the failure filter because its similarity is above 0.85, and it adds a warning: retrying the rejected call loses the customer. Episode C refers to a tool that no longer exists and decays to nothing. The model's first action is to open the approval ticket, which saves one rejected tool call and the retry loop that sank episode B.

Episodic memory: write path (top) and read path (bottom)Session endsevents + stateExtractorgoal, trace, outcomeDistillerlesson, 1 model callValidatoroutcome label, PIIEpisode storeupsert by sessionNew user turna fresh goalLoadMemoryToolmodel asks for recallsearchMemoryapp + user scopedRank + filtersimilar, verified, freshPrecedent texttop 2 episodesreadThe validator and the ranker are where episodic memory succeeds or fails:an unlabelled failure recalled as precedent teaches the agent to repeat it.
Episodes are mined after the session closes and recalled through the standard memory tool; labelling and ranking sit on both paths.

Failure modes

The failure modes of episodic memory are mostly failures of trust.

  • Poisoned precedent. An unlabelled failure, stored as unknown and recalled as if it worked, teaches the agent a broken procedure. Down-weight unknown outcomes and measure how often they are recalled.
  • Injection that persists. A prompt injection inside a tool result becomes part of the trace and is replayed into future sessions. Store tool names, codes and whitelisted arguments, never raw tool output, and run the distiller's input through the same sanitiser as live tool results.
  • Stale procedure. Tools change; episodes do not. The toolset hash catches renamed or re-schemed tools but not changed business rules, so retention limits still matter.
  • Anchoring. The model copies the precedent even when the new case differs, such as a refund under the approval threshold. Keep the 'evidence, not instruction' label and cap recall at two episodes.
  • Duplicate ingestion. Checkpointed sessions are added several times. Upsert by session id, or one conversation becomes three equally strong precedents.
  • Cross-user leakage. Episodes from one user recalled for another expose their data. Apply the app and user filter inside the store query, never after ranking. If you want shared, organisation-wide procedure, promote reviewed episodes into a separate, anonymised store.

Operating it

Measure whether the feature earns its cost. Track the recall rate (turns where the model called the memory tool and received an episode), the success rate of tasks with and without a recalled precedent, and tool calls per resolved task. The honest test is an A/B split on the memory service: half the sessions get the episodic decorator, half get an empty search result. If success and tool calls per task do not move, delete the feature.

Bound growth with a retention window, such as 180 days, plus deletion when the toolset hash has been retired for a full window. Delete by user on account removal; the right-to-forget article covers the mechanics. Log every recall with the episode ids returned, so an incident can be traced to the precedent that caused it.

Trade-offs

ChoiceGainCost
Structured records vs transcriptsshort, rankable, outcome explicitextractor code per agent; details lost
Distilled lessonactionable one-linera model call per session; can hallucinate
Pull via LoadMemoryToolno cost on turns that need no historymodel must realise it needs precedent
Push via a before-model callbackprecedent always presenttokens on every call; more anchoring
Per-user scopeno leakagecold start for every new user
Shared reviewed storeorganisation learns onceanonymisation and review workflow

What to do next

  1. Pick one multi-step task your agent repeats and list the tool that can write a trustworthy task.outcome into state.
  2. Implement EpisodeExtractor for that agent with a whitelist of arguments, and unit-test it on recorded sessions.
  3. Wrap your existing memory service with the episodic decorator and keep its upsert idempotent; see cross-session memory for the write trigger and the MemoryService architecture for the contract.
  4. Tune the ranking with the techniques in long-term memory retrieval.
  5. Add retention, delete-by-user and recall logging, following the right-to-forget guide.
  6. Run the A/B split for two weeks, and keep the feature only if success rate or tool calls per task improve. The events article explains the event fields the extractor reads.
Key takeaway: Episodic memory stores what the agent tried, in what order, and whether it worked, so a future session can learn from precedent. Mine episodes from ADK's own events, keep them small and structured, let code rather than a model write the outcome, serve them through BaseMemoryService so LoadMemoryTool can recall them, and rank by similarity, outcome, toolset freshness and age. Label every recall as evidence, guard against persisted injection and cross-user leakage, and keep the feature only if an A/B test shows it helps.