Giving an agent memory sounds like a storage decision: pick a vector database, embed what the user says, retrieve the nearest neighbours next time. Teams that stop there find the same problems within weeks. The agent recalls a preference the user changed a month ago, surfaces a trivial remark ahead of a medical allergy, contradicts itself because two versions of a fact are both in the store, and cannot honour a request to forget something because copies of it live in summaries and caches.

Those are lifecycle problems, not storage problems. This article treats agent memory as a pipeline with stages that each need a design: what a memory record contains, how retrieval ranks candidates, how new information reconciles with old, how episodes consolidate into facts, how memories are forgotten and deleted, and how the whole thing stays isolated and resistant to poisoning. The kinds of memory are compared in agent memory layers compared, and the store layout and read and write paths for long-running agents in memory architectures for long-running agents; this page assumes those and focuses on the lifecycle.

Advertisement

Memory is a lifecycle, not a store

Every memory an agent keeps goes through the same stages. It is captured from a conversation turn or tool result, extracted into a short claim, reconciled against what is already known, stored, retrieved and scored when a later turn needs it, periodically consolidated with related memories, and eventually forgotten by expiry, decay or explicit deletion.

The stages split cleanly into two paths. The request path, extraction and retrieval, must be fast because a user is waiting. Everything else, reconciliation when it is expensive, consolidation and forgetting, can run in the background. Designing the split up front keeps the agent responsive and makes the slow, careful work possible.

Agent turnuser + tool eventsExtractcandidate memoriesReconcileADD / UPDATE / DELETE / NOOPMemory storerecords + vectorspartitioned by tenantand userRetrieve + scorerelevance, recency, importanceContext assemblytoken budgetConsolidation jobepisodes -> factsForgetting jobTTL, decay, deletionwritereadpromptbackground, off the request path
The memory lifecycle. Writes pass through extraction and reconciliation; reads are scored under a token budget; consolidation and forgetting run as background jobs against the same partitioned store.

The memory record

Most memory bugs trace back to a record that is just text plus a vector. A useful record carries enough metadata for every later stage to make a decision without calling a model.

from dataclasses import dataclass, field
from datetime import datetime

@dataclass
class Memory:
    id: str
    tenant_id: str                 # hard partition key, never optional
    user_id: str
    kind: str                      # "episode" | "fact" | "preference" | "procedure"
    text: str                      # the claim, in one sentence
    embedding: list[float]
    source: str                    # "user_said" | "tool_result" | "inferred" | "consolidated"
    source_refs: list[str]         # conversation / event ids it came from
    created_at: datetime
    last_used_at: datetime
    importance: float              # 0..1, set at write time
    confidence: float              # 0..1, lower for inferred claims
    supersedes: list[str] = field(default_factory=list)
    expires_at: datetime | None = None
    sensitivity: str = "normal"    # "normal" | "personal" | "restricted"
    deleted: bool = False

Three fields do most of the work. source and confidence separate what the user said from what the agent inferred, so an inference never overrides a direct statement. supersedes links a new version of a claim to the one it replaced, which is how you resolve contradictions without losing history. tenant_id is a hard partition key that every query must filter on, not an optional attribute.

Keep each record to one claim. A memory that says the user is vegetarian, lives in Leeds and prefers morning meetings cannot be updated or deleted cleanly when only one of those changes.

Advertisement

Retrieval scoring: relevance, recency, importance

Nearest-neighbour similarity alone is a poor ranking function. It ignores how old a memory is and how much it matters. The Generative Agents work by Park and colleagues popularised a score that combines three signals: relevance to the current situation, recency, and importance assigned when the memory was written. The version below follows that shape, with weights and decay rates that are our own illustrative choices.

import math

HALF_LIFE_DAYS = {"episode": 7, "preference": 90, "fact": None, "procedure": None}
W_REL, W_REC, W_IMP = 0.6, 0.2, 0.2    # illustrative weights, tune on your own traces

def recency(m, now):
    h = HALF_LIFE_DAYS[m.kind]
    if h is None:                        # stable facts do not fade with age
        return 1.0
    age = (now - m.last_used_at).total_seconds() / 86400
    return 0.5 ** (age / h)

def score(m, relevance, now):
    return W_REL * relevance + W_REC * recency(m, now) + W_IMP * m.importance

def retrieve(store, query_vec, user, now, k=8, min_rel=0.35):
    # the tenant and user filter is applied inside the store query, before similarity
    cands = store.search(query_vec, top=50,
                         where={"tenant_id": user.tenant_id, "user_id": user.id, "deleted": False})
    cands = [(m, rel) for m, rel in cands
             if rel >= min_rel and (m.expires_at is None or m.expires_at > now)]
    ranked = sorted(cands, key=lambda mr: score(mr[0], mr[1], now), reverse=True)[:k]
    for m, _ in ranked:
        store.touch(m.id, now)           # reading refreshes last_used_at
    return [m for m, _ in ranked]

The important design choice is that decay depends on the kind of memory. Episodes, such as what the user asked for yesterday, lose value quickly. Stable facts, such as an allergy or a legal name, should not decay at all; they stay until they are superseded or deleted.

A worked scoring example

A user asks a travel agent to book dinner. Three memories pass the relevance threshold:

MemoryKindRelevanceAgeRecencyImportanceScore
Liked the Thai place last monthepisode0.8230 days0.0510.30.562
Wants to keep this trip cheapepisode0.741 day0.9060.60.745
Severe peanut allergyfact0.61200 days1.00.950.756

With per-kind decay, the allergy ranks first even though its similarity to a dinner request is the lowest of the three. If the allergy were treated as an episode with a 7-day half-life, its recency would be close to zero and its score would fall to about 0.556, below the Thai restaurant, and with a small retrieval budget it could drop out of context entirely. That is the concrete failure a single global decay rate causes, and it is why importance and kind must be decided at write time.

Refreshing last_used_at on retrieval keeps useful episodes alive, but it can also create a rich-get-richer loop in which the same few memories keep winning. Cap how often one memory can be refreshed per session, or refresh only when the agent actually uses the memory in its answer.

Reconciliation: ADD, UPDATE, DELETE, NOOP

Appending every extracted claim produces contradictions. When the user says they have moved from Leeds to York, a store that already contains the Leeds address now holds both, and retrieval may return either. A common pattern, used by several open-source memory layers, is to retrieve the nearest existing memories for each candidate and ask a model to choose one operation: add a new record, update an existing one, delete one that is no longer true, or do nothing.

from dataclasses import replace

RECONCILE_PROMPT = """Existing memories (id: text):
{existing}

New candidate: {candidate}

Reply with JSON: {{"op": "ADD" | "UPDATE" | "DELETE" | "NOOP", "target_id": "...", "text": "..."}}
UPDATE when the candidate refines or corrects one existing memory; DELETE when it
states that an existing memory is no longer true; NOOP when it adds nothing."""

def reconcile(store, llm, cand, user):
    near = store.search(cand.embedding, top=5,
                        where={"tenant_id": user.tenant_id, "user_id": user.id, "deleted": False})
    decision = llm.json(RECONCILE_PROMPT.format(
        existing="\n".join(f"{m.id}: {m.text}" for m, _ in near), candidate=cand.text))
    ids = {m.id for m, _ in near}
    if decision["op"] in ("UPDATE", "DELETE") and decision.get("target_id") not in ids:
        decision = {"op": "ADD"}          # never let the model touch a record it was not shown
    if decision["op"] == "ADD":
        store.insert(cand)
    elif decision["op"] == "UPDATE":
        store.insert(replace(cand, text=decision["text"], supersedes=[decision["target_id"]]))
        store.retire(decision["target_id"])      # kept for audit, excluded from retrieval
    elif decision["op"] == "DELETE":
        store.retire(decision["target_id"])
    store.audit(user, cand, decision)            # every write is explainable later

Two guards make this safe. The model may only update or delete records it was shown, so a hallucinated id cannot damage unrelated memories. And updates retire rather than overwrite, keeping the old record for audit and linking it through supersedes. When the reconciliation call is too slow for the request path, write the candidate as a provisional episode immediately and reconcile it in the background.

Consolidation: from episodes to knowledge

Over weeks, an agent accumulates hundreds of episodes that say the same thing in different words: the user chose the aisle seat, asked for an aisle seat, complained about a middle seat. Retrieval then fills the context with near-duplicates. A consolidation job runs periodically per user, clusters related episodes, and writes one higher-level memory, the user strongly prefers aisle seats, with source='consolidated' and references to the episodes it summarises.

Consolidation needs rules. Only consolidate claims supported by several independent episodes, so one sarcastic remark does not become a fact. Carry the lowest confidence of the inputs, not the highest. Never consolidate across users or tenants. And keep the source references, because when a user deletes an episode, any consolidated memory built from it must be revisited. Consolidation is also a natural point to compact old episodes, which relates closely to context compaction within a single session.

Forgetting and deletion

An agent that never forgets gets worse over time, not better: retrieval slows, stale memories crowd out fresh ones and the privacy exposure grows. Forgetting has four separate mechanisms, and a production system needs all of them.

  • Expiry. Records with a natural lifetime, such as a one-off booking reference, carry an expires_at and are filtered out of retrieval immediately, then purged by the forgetting job.
  • Decay-based eviction. Episodes whose score has stayed below a floor for a long period are archived or dropped.
  • Capacity budgets. A per-user limit, for example a few thousand active records, forces eviction of the lowest-value memories and bounds cost.
  • Explicit deletion. When a user asks the agent to forget something, or deletes their account, the request must reach every copy.

Explicit deletion is the hard one because memories propagate. A single fact may exist as a record, an embedding in a vector index, a line in a consolidated summary, a cached prompt, an evaluation trace and a backup. Build the deletion path as a cascade driven by source_refs: delete the record and its vector, find every consolidated memory that cites it and regenerate or retire it, invalidate caches, and record the deletion in an audit log that stores ids, not content. Test it with a planted canary fact that must not be retrievable afterwards.

Isolation and memory poisoning

Memory turns a one-off prompt injection into a persistent one. If a web page or email the agent reads contains an instruction, and extraction writes it as a memory, every future session retrieves it. Treat memory as untrusted data, with the same care as tool output.

  • Write memories only from designated sources, and record which source produced each one. Content that arrived through tools gets lower confidence and is never stored as an instruction.
  • Put retrieved memories into the prompt as clearly delimited data, and tell the model that they are notes, not commands.
  • Enforce tenant and user filters inside the store query, never by filtering results in application code afterwards, where a bug returns another customer's data.
  • Tag sensitivity at write time and keep restricted categories, such as health or credentials, out of memory unless the product explicitly needs them.
  • Give users a way to see and edit what the agent remembers; it is the best poisoning detector you have.

Failure modes and how to measure memory

  • Stale recall. An outdated preference wins because reconciliation missed the change. Measure how often retrieved memories have been superseded.
  • Crowding. Near-duplicate episodes fill the budget. Measure the share of retrieved items with high mutual similarity.
  • Over-personalisation. The agent applies an old preference where the user clearly wants something different. Let explicit instructions in the current turn always override memory.
  • Latency creep. Retrieval slows as stores grow; see memory and vector store options for scaling choices.
  • Incomplete deletion. A forgotten fact reappears through a summary. Run the canary deletion test in CI.

Evaluate with scripted multi-session scenarios: plant facts, change some, delete others, then ask questions whose correct answers depend on the latest state. Score precision of retrieved memories, whether the answer used the right version, and whether deleted content ever appears. Run the suite whenever you change weights, prompts or the embedding model.

What to do next

  1. Define a memory record with tenant, source, confidence, kind, supersedes and expiry fields before storing anything else.
  2. Give each kind of memory its own decay rule, and make stable facts non-decaying.
  3. Implement reconciliation with ADD, UPDATE, DELETE and NOOP, restricted to records the model was shown, and retire rather than overwrite.
  4. Add a background consolidation job that requires several supporting episodes and keeps source references.
  5. Build the deletion cascade and a canary deletion test before launch, not after the first request.
  6. Move tenant filtering into the store query and delimit memories as data in the prompt.
  7. Write a multi-session evaluation suite and run it on every change to scoring or prompts.
Key takeaway: Agent memory succeeds or fails on its lifecycle, not its database. Store one claim per record with its source, confidence, kind and lineage. Rank retrieval by relevance, recency and importance, with decay rules that depend on the kind of memory so that a critical fact never fades like yesterday's chat. Reconcile new claims against old ones instead of appending, consolidate repeated episodes into knowledge in the background, and treat forgetting as a first-class feature with a deletion path that reaches every derived copy. Keep tenants apart inside the query, and treat everything in memory as untrusted data.