An agent without memory starts every conversation as a stranger. Adding memory looks easy: save some facts, search them later, paste the results into the prompt. The hard parts appear once the system is real. Who decides what gets saved, the model or your code? What stops one customer's facts reaching another's prompt? What happens when two sub-agents update the same fact at the same moment, or when the memory store is slow and the user is waiting?

This article treats memory as a system rather than a data structure. The memory lifecycle itself, scoring, reconciliation, consolidation and forgetting, is covered in agent memory architecture, and the stores and context-window assembly for long-running work in memory architectures for long-running agents. Here the focus is the layer around them: control, the service contract, scoping, latency, concurrency and operations, with a worked example and code you can adapt.

Advertisement

Two ways to put memory under control

There are two places the decision to read or write memory can live, and every design is some mix of them.

In agent-directed memory, the model is given tools such as recall and remember and decides when to call them. It can look things up only when the task needs them and can save exactly the fact it judges important. The cost is that recall depends on the model remembering to recall: an agent that does not think to search will confidently answer without context, and one that over-saves fills the store with noise.

In harness-directed memory, your code runs hooks around each turn. A read hook retrieves relevant memories and places them in the context window before the model runs; a write hook extracts candidate facts from the finished turn and sends them for storage. Behaviour is predictable and testable, and the model cannot forget to do it. The cost is that the harness guesses what is relevant from the user's message alone, and it spends retrieval latency and tokens on every turn whether or not they help.

PropertyAgent-directed (tools)Harness-directed (hooks)
Who decidesThe model, per stepYour code, every turn
Recall qualityTargeted when called; missing when notConsistent; can be irrelevant
LatencyPaid only when used, but mid-reasoningPaid every turn, before the model starts
TestabilityDepends on model behaviourDeterministic, unit-testable
RiskModel writes injected content as factExtraction writes noise at scale

A sound default is hybrid: a harness read hook injects a small, high-precision set (the user's profile and any memories with very high relevance), the model gets a recall tool for deeper lookups, and writes go through the harness after the turn so they can be validated, rather than being committed directly by a tool call mid-turn.

An agent memory system: two control paths into one scoped serviceUser turnauth: tenant, userHarnessbuilds context windowModelplans, calls toolsMemory toolsrecall, rememberRead hookinject before turnWrite hookextract after turnMemory servicescope from auth, versions, idempotencyscoped callsqueueProfile storekeyed factsEpisode logappend-onlyVector indexsemantic recallAudit + deletionwho wrote what
Both control paths end at one memory service. Scope comes from authentication, never from model arguments; writes are queued; every write is audited and deletable.

The memory service contract

Whichever path calls it, put memory behind one service with a narrow interface. Agents, hooks and admin tools then share the same scoping, validation and audit, and you can change storage without touching agent code. The contract needs fewer operations than people expect, but each one needs precise semantics.

from dataclasses import dataclass, field
from typing import Optional

@dataclass(frozen=True)
class Scope:                      # derived from the authenticated request, never from the model
    tenant: str
    user: str
    agent: Optional[str] = None   # None = shared by all agents serving this user

@dataclass
class Memory:
    id: str
    subject: str                  # what the fact is about, e.g. "contact_preference"
    content: str
    source: str                   # "user_said", "tool_result", "inferred"
    version: int = 1
    superseded_by: Optional[str] = None
    tags: list = field(default_factory=list)

class MemoryService:
    def search(self, scope, query, k=5, subjects=None, timeout_s=0.15) -> list: ...
    def get(self, scope, subject) -> Optional[Memory]: ...
    def put(self, scope, subject, content, source,
            expected_version=None, idempotency_key=None) -> Memory: ...
    def forget(self, scope, subject=None, memory_id=None) -> int: ...   # returns count removed

Four properties carry most of the weight. search takes an explicit timeout because the caller is usually on the user's critical path. put takes an expected_version so a writer can say "update this only if nobody changed it since I read it". put takes an idempotency key so a retried write does not create a duplicate. And forget works by subject as well as by id, because deletion requests arrive as "forget my address", not as row ids.

When you expose this to the model as tools, expose less. The model should see recall(query) and perhaps remember(subject, content), with scope filled in by the server and expected_version handled by the harness. Every argument the model controls is an argument an injected instruction can control too.

Advertisement

Scope is enforced at the boundary

The most damaging memory bug is a cross-tenant leak: one user's facts retrieved into another user's prompt. It usually happens because scope is a filter the caller passes, and some code path passes the wrong one or none. Make scope impossible to choose. The service derives it from the authenticated request, and the storage layer refuses any query without it.

def scope_from_request(req) -> Scope:
    claims = verify_token(req.headers["authorization"])   # raises if invalid
    return Scope(tenant=claims["tenant"], user=claims["sub"], agent=req.agent_name)

class ScopedStore:
    def __init__(self, backend):
        self.backend = backend
    def query(self, scope: Scope, **kw):
        if not scope or not scope.tenant or not scope.user:
            raise PermissionError("unscoped memory query")
        return self.backend.query(tenant=scope.tenant, user=scope.user,
                                  agent=scope.agent, **kw)

Use physical separation where the stakes justify it: a partition or namespace per tenant in the vector index, not just a metadata filter, so a missing filter returns nothing rather than everything. Decide explicitly whether memories are per agent or shared across the agents serving a user; a coding agent and a billing agent probably should not read each other's notes.

The read path under a latency budget

A harness read hook sits between the user pressing enter and the model's first token, so it gets a hard budget, typically on the order of 100-200 milliseconds. Fetch the profile and run the semantic search in parallel, give each a timeout, and treat a timeout as "no memory this turn" rather than an error. A slightly less personal answer is better than a stalled one.

import asyncio

async def build_memory_context(svc, scope, user_msg, budget_s=0.15, max_tokens=600):
    async def profile():
        return await svc.get_profile(scope)
    async def recall():
        return await svc.search(scope, user_msg, k=5)
    results = await asyncio.gather(
        asyncio.wait_for(profile(), budget_s),
        asyncio.wait_for(recall(), budget_s),
        return_exceptions=True)
    items = []
    for r in results:
        if isinstance(r, Exception):
            metrics.incr("memory.read.degraded")      # visible, not fatal
            continue
        items.extend(r if isinstance(r, list) else [r])
    return render_with_provenance(items, max_tokens)   # trims to the token budget

Cache the profile per session; it changes rarely and is needed on every turn. Do not cache semantic search results across turns, because the query changes. Render memories with their source and date so the model can weigh a fact the user stated last week differently from one it inferred last year.

The write path is asynchronous

Writes should not block the reply. The write hook places candidate memories on a queue and returns; a worker validates, deduplicates and reconciles them against existing memories, then commits. Process the queue in order per scope, so that "my address is A" followed by "actually it is B" cannot be committed in the wrong order.

Asynchronous writes create one visible problem: the user says something, then refers to it two turns later, before the worker has committed it. Solve this with read-your-writes within the session: the harness keeps the session's pending writes in a small buffer and merges them into reads until the worker confirms. Across sessions, a few seconds of lag is acceptable.

Concurrent writers and shared memory

Concurrency appears as soon as an orchestrator runs sub-agents in parallel, or a user has two sessions open. The classic failure is the lost update: two writers read the same fact, each changes it, and the second write silently erases the first. Optimistic concurrency fixes it: each memory has a version, writers send the version they read, and the store rejects a write whose version is stale.

import hashlib

def update_fact(svc, scope, subject, transform, source, retries=3):
    for _ in range(retries):
        current = svc.get(scope, subject)
        new_content = transform(current)          # may keep current if it outranks us
        digest = hashlib.sha256(new_content.encode()).hexdigest()[:16]
        try:
            return svc.put(scope, subject, new_content, source=source,
                           expected_version=current.version if current else 0,
                           idempotency_key=f"{subject}:{digest}")
        except VersionConflict:
            continue              # someone else wrote; re-read and re-apply
    raise RuntimeError(f"could not update {subject} after {retries} attempts")

For facts that accumulate rather than replace, such as a list of open issues, avoid read-modify-write entirely: append records and let consolidation merge them. For shared memory between agents, record which agent wrote each memory and give each agent a trust level. A research sub-agent's summary of a web page should not overwrite a fact the user stated directly; source ranks above recency when they conflict.

Worked example: two sub-agents and one preference

A support orchestrator handles "please stop calling me, email only, and also cancel the extra seat". It runs two sub-agents in parallel. The account agent reads contact_preference at version 3 (phone) and writes email at version 3. The billing agent, finishing the seat cancellation, reads the same memory at version 3 and writes a note that billing confirmations go to phone, also at version 3.

Without versions, whichever write lands second wins, and the user may get the phone call they asked to stop. With versions, the first write commits as version 4; the second is rejected with a conflict, re-reads version 4, sees email, and its transform keeps the user's stated preference because user_said outranks inferred. The audit log shows both attempts, so when someone asks why the billing agent's note never appeared, the answer is recorded rather than guessed.

Failure modes

FailureWhat happensDefence
Memory poisoningInjected text in a document is saved as a user fact and steers later sessionsWrites go through validation; tag source; never auto-save tool output as user_said
Cross-tenant leakAnother user's facts appear in a promptScope from auth, physical partitions, unscoped queries refused
Stale profile cacheAgent uses an address the user changedInvalidate the session cache on every committed write for that scope
Write stormExtraction saves every turn as several memoriesRate-limit per scope, deduplicate, require a subject
Embedding model changeOld and new vectors are not comparable; recall collapsesVersion vectors by model; re-embed and dual-read during migration
Incomplete deletionA forgotten fact survives in a cache, index or summaryDeletion fans out to every derived store, with a verification job

Running it in production

Instrument the read path with latency percentiles and a degraded-read rate, and the write path with queue depth, conflict rate and rejected writes. Track how often injected memories are actually used in answers: if the model ignores most of them, retrieval is noise and costs tokens for nothing. Build a small regression set of conversations that depend on remembered facts and run it whenever extraction prompts, retrieval settings or the embedding model change.

Treat deletion as a feature with a service level. A user's request to forget must reach the profile store, the episode log, every vector index, every summary derived from those memories and every cache, within a stated time, and a periodic job should confirm the subject is really gone. Choosing the storage itself is covered in memory and vector store options.

Trade-offs

DecisionOption AOption B
ControlHooks: predictable, always pays latencyTools: targeted, depends on the model
Write timingSynchronous: simple, slows repliesQueued: fast replies, needs read-your-writes
SharingPer-agent scopes: safe, duplicated factsShared user scope: consistent, needs trust levels
IsolationMetadata filter: cheap, one bug from a leakPartition per tenant: costlier, fails closed

What to do next

  1. Put every memory read and write behind one service and delete direct store access from agent code.
  2. Derive scope from authentication in that service, and add a test that an unscoped query raises.
  3. Choose your control mix explicitly: which reads are hooks, which are tools, and route all writes through validation.
  4. Give the read hook a hard timeout and a degraded-read metric, and cache the profile per session.
  5. Add versions and idempotency keys to put, and test two concurrent writers against the same subject.
  6. Write a deletion runbook covering every derived store, and schedule a job that verifies deletions.
Key takeaway: An agent memory system is mostly not about storage. Decide deliberately whether the model or the harness controls each read and write, put all access behind one service whose scope comes from authentication, keep reads inside a latency budget that degrades to no memory, queue writes with read-your-writes inside the session, protect shared facts with versions and source-based trust, and make deletion reach everything derived from a memory.