A chat assistant that remembers your name and preferred tone has a memory problem, but a small one. An agent migrating 400 files from one framework to another over three days has a much larger one. It will fill and reset its context window dozens of times, be interrupted and resumed, discover facts about the codebase that it must not forget, make decisions whose reasons matter later, and try approaches that failed and must not be tried again. Its memory is not a personalisation feature. It is what makes the task finishable at all.

Much of what is written about agent memory covers the conversational case: working, episodic and semantic memory, consolidation, forgetting and user privacy, described well in agent memory layers compared. This article covers task memory for long-horizon agents. It proposes four stores with different contracts, a record format with provenance and supersession, a write path that validates what the agent tries to remember, a read path that assembles each context window under a token budget, and tests that show whether the memory actually works.

Advertisement

Why the context window is not the memory

The transcript of a long run is a log, not a memory. It is too large to reload, it mixes important facts with noise, and it records beliefs the agent later corrected without marking them as wrong. Compressing it helps within one window, as covered in context compaction for long-running agents, but a summary of a summary drifts. After a few rounds, the specific facts the agent needs, such as the exact build command or why a module was left alone, are the first to blur.

What a long-running agent needs to carry across windows is a small set of things with very different shapes. The goal and constraints never change and must never be lost. The plan and its progress change constantly and must be exactly right. Decisions accumulate and must keep their reasons. Facts about the environment accumulate, and some become false. The raw history is rarely needed but occasionally essential. Putting all of these in one vector store, retrieved by similarity, treats them as the same kind of thing, and that is the most common design mistake.

Four stores with different contracts

StoreHoldsWrite ruleRead rule
Task ledgerGoal, constraints, plan items with status, open questionsStructured updates through tools; every status change loggedPinned: goal and constraints always; current plan slice always
Decision logChoices made, alternatives rejected, reasons, who decidedAppend-only; a reversal is a new entry that references the old oneBy scope: decisions touching the current work item
Fact storeClaims about the environment, with source and validityValidated: must cite a source; conflicting facts supersede, never overwriteFiltered by scope, then ranked by relevance and recency
Episodic archiveRaw transcripts, tool calls and outputsEverything, automatically, outside the model's controlSearch only, on demand; never loaded whole
Task memory for a long-running agent: four stores, one write path, one read pathAgent loopone context windowMemory toolspropose writesWrite validatorprovenance, conflictsTask ledgergoal, plan, todo, statusDecision logappend-only, with reasonsFact storetyped, sourced, supersededEpisodic archiveraw transcripts, tool outputalwaysContext assemblerpinned goal, ledger slice, facts, tailsearch onlynext windowGround-truth checkgit, APIs, filesreconcileEvaluation: recall probes, resume-from-kill tests, stale-fact rate, context tokens per stepmeasured per run, compared across memory designs
The agent proposes memory writes through tools; a validator checks them before they reach the ledger, decision log or fact store. The episodic archive records everything automatically. Each new context window is assembled from the stores under a budget, and ledger state is reconciled against ground truth.

The split exists because each store fails differently. A ledger that is slightly wrong sends the agent to redo finished work or skip unfinished work. A decision log without reasons invites the agent to relitigate settled questions in every new window. A fact store without provenance cannot tell a verified observation from something the model guessed. Keeping them separate lets each have the right integrity rule and the right retrieval rule.

Storage technology is a secondary choice. A relational database handles the ledger, decisions and facts comfortably, with a vector index added for fact retrieval; object storage holds the archive. The trade-offs between vector store options are covered in memory and vector store options.

Advertisement

Records with provenance and supersession

The fact store carries most of the risk, so its records need more structure than a line of text and an embedding. Each fact states what kind of claim it is, where it came from, when it was observed, what it applies to, and whether a newer fact replaced it.

from dataclasses import dataclass, field
from typing import Optional

@dataclass
class Fact:
    id: str
    text: str                    # "Build command is `make build-ci`, not `make build`"
    kind: str                    # observation | inference | instruction-from-user
    scope: list[str]             # ["repo:billing", "path:services/invoice/"]
    source: str                  # "tool:run_shell#call-8812" or "user:msg-41"
    source_hash: Optional[str]   # hash of the file or output the fact was read from
    observed_at: str
    confidence: float            # set by the validator from kind and source, not by the model
    supports: list[str] = field(default_factory=list)   # fact ids an inference rests on
    supersedes: Optional[str] = None
    superseded_by: Optional[str] = None

Three rules keep this honest. First, the kind separates what the agent saw from what it concluded. An observation cites a tool call whose output is in the archive; an inference cites the facts it was drawn from. Inferences are ranked below observations and never used as the source of another observation. Second, facts are never edited or deleted during a run. When a newer observation contradicts an old one, the new fact supersedes it, and the chain shows when and why belief changed. Third, source_hash lets the system detect staleness: if the file a fact was read from has changed since, the fact is flagged for re-verification before it is served.

The write path: what deserves to be remembered

If the agent can write anything to memory, it will write too much, and some of it will be wrong. Give it narrow tools instead of a free-form notes field: record_fact, record_decision, update_plan_item, add_open_question. Each tool call goes through a validator that the model does not control.

def validate_fact(proposed, archive, store):
    if proposed.kind == "observation":
        call = archive.get(proposed.source)
        if call is None:
            return reject("observation must cite a recorded tool call")
        if call.untrusted_content and looks_like_instruction(proposed.text):
            return reject("instructions from tool output are not facts")
    if proposed.kind == "inference" and not proposed.supports:
        return reject("inference must reference supporting fact ids")
    proposed.confidence = CONFIDENCE[(proposed.kind, source_type(proposed.source))]
    for old in store.same_subject(proposed, scope=proposed.scope):
        if contradicts(old, proposed):
            if proposed.observed_at <= old.observed_at:
                return reject(f"older than current belief {old.id}")
            proposed.supersedes = old.id
    if store.near_duplicate(proposed):
        return merge_reference(proposed)
    return accept(proposed)

The contradiction check is the expensive part. A cheap approach narrows candidates by scope and embedding similarity, then asks a small model whether the two statements can both be true. The rest is deterministic. Rejections go back to the agent as tool errors, which teaches it within the run what is worth recording.

The instruction check matters more than it looks. A web page or document the agent reads can contain text that, stored as a fact, becomes a durable instruction retrieved into every future window. Memory turns a one-off prompt injection into a persistent one. Tool output is data; it can support facts, but it should never become a stored directive.

The read path: assembling each window under a budget

Every new context window, and every step within one, is built by an assembler with a fixed layout and a token budget per section. A fixed layout matters because the model learns where to look, and because budgets per section stop one growing store from crowding out the others.

BUDGET = {"pinned": 1500, "plan": 2000, "decisions": 1500, "facts": 3000, "tail": 4000}

def assemble(task, item, stores, tok):
    parts = []
    parts.append(fit(render_goal_and_constraints(task), BUDGET["pinned"], tok, must_fit=True))
    parts.append(fit(render_plan(stores.ledger, focus=item, done_window=5), BUDGET["plan"], tok))
    parts.append(fit(render(stores.decisions.for_scope(item.scope)), BUDGET["decisions"], tok))
    facts = stores.facts.current(scope=item.scope)             # superseded facts excluded
    facts = [f for f in facts if not stale(f)] + reverify_queue(facts)
    ranked = rank(facts, query=item.description,
                  w_relevance=0.6, w_recency=0.25, w_confidence=0.15)
    parts.append(fit(render(ranked), BUDGET["facts"], tok))
    parts.append(fit(render_recent_steps(stores.archive, n=8), BUDGET["tail"], tok))
    return "\n\n".join(parts)

Scope filtering comes before similarity. Similarity alone retrieves facts about a different module that happen to use the same words; filtering by the work item's scope first makes retrieval precise, and similarity then orders what is left. The ranking weights echo the recency, importance and relevance score from the Generative Agents paper (Park et al., 2023), with confidence standing in for importance; tune them on your own recall tests rather than copying these numbers. The paging approach of MemGPT (Packer et al., 2023), where the agent itself moves information between context and external storage, is a complementary design; the assembler here keeps paging decisions in deterministic code and leaves the model only the write tools and a search tool over the archive.

Handoffs between context windows

The moment one context window ends and the next begins is where long runs usually lose the thread. Treat it as a checkpoint with three steps. First, the agent writes a short handoff note: what it was doing, what it expected next, and anything it had noticed but not yet recorded. Second, the system reconciles the ledger with ground truth. For a code migration that means checking which files actually changed in version control and which tests pass, and correcting plan items whose recorded status disagrees. Third, the next window is assembled from the stores, with the handoff note in the tail section.

The reconciliation step is what keeps memory from drifting away from reality. The agent's belief that it finished a file is a claim; the diff is evidence. When they disagree, evidence wins, and the disagreement itself is recorded as a fact, because it usually means a step failed silently.

Worked example: a three-day framework migration

An agent migrates a service of 400 source files from one web framework to another. The ledger holds one plan item per module, 38 in all, each with a status. On day one, running the tests with make build fails; the agent reads the CI configuration, finds make build-ci, and records an observation citing that tool call. On day two, a teammate changes the Makefile. The fact's source_hash no longer matches, so the assembler queues it for re-verification, the agent re-reads the file, and a new fact supersedes the old one.

Memory writeStoreWhy it pays off later
Decision: keep the legacy auth middleware; the new framework's version drops a header the mobile app needsDecision logIn window 17 the agent is about to replace it again; the scoped decision is retrieved and it stops
Fact: the invoice module's tests need a local database containerFact store, scope invoiceRetrieved only when working on invoice, not in every window
Open question: are the admin routes still used?LedgerSurfaces at the end of the run for a human instead of being silently skipped
Failed approach: automated rewrite of template tags breaks escapingDecision log, marked rejectedPrevents the same dead end in every new window

Without the decision log, the most expensive failure in long runs appears: the agent rediscovers a problem, forgets its earlier conclusion, and re-applies a change it previously reverted. Recorded reasons, retrieved by scope, are what break that loop.

Failure modes

  • Memory poisoning. Injected instructions stored as facts and replayed into every window. Reject instruction-shaped text from untrusted sources and keep the source on every record.
  • Self-reinforcing inference. A guess stored as a fact, later cited as evidence for another fact. Separate kinds and never let inferences stand as sources for observations.
  • Stale facts. True when recorded, false now. Hash sources and re-verify on change; prefer re-reading a file to trusting a memory of it.
  • Ledger drift. Recorded progress disagrees with the world. Reconcile against ground truth at every handoff.
  • Crowding. One store grows until it pushes out the plan. Per-section budgets and scope filters.
  • Retrieval misses. The fact exists but is not retrieved. Measure with recall probes, below, rather than guessing.

Evaluating memory

Memory designs are easy to argue about and cheap to test. Three tests cover most of what matters. Recall probes: at fixed points in a recorded run, ask questions whose answers were established earlier, such as the build command or why a module was skipped, and score whether the assembled context lets the model answer correctly. Resume tests: kill the agent at random points, restart it from memory alone, and compare the final result and the amount of repeated work with an uninterrupted run. Contradiction audits: sample served facts and check them against ground truth to measure the stale-fact rate.

Track context tokens spent on memory per step alongside these scores. A design that improves recall by loading everything is not an improvement, and the budgets in the assembler are the dial you are tuning.

What to do next

  1. Write down what your agent must carry across windows and sort it into ledger, decisions, facts and archive.
  2. Record every tool call and output in an archive the model cannot edit.
  3. Replace free-form notes with narrow memory tools and a validator that requires a cited source.
  4. Add kind, scope, source hash and supersession to fact records, and never overwrite a fact during a run.
  5. Build a context assembler with a fixed layout, per-section budgets and scope filtering before similarity.
  6. Reconcile the ledger against ground truth at every window handoff.
  7. Reject instruction-shaped text from untrusted tool output at the write path.
  8. Build recall probes and resume-from-kill tests, and use them to tune budgets and ranking weights.
Key takeaway: A long-running agent's memory is task state, not chat history. Split it into a ledger, a decision log, a sourced fact store and an archive, each with its own write and read rules. Validate every write, supersede rather than overwrite, assemble each window under per-section budgets with scope before similarity, reconcile beliefs against ground truth at handoffs, and test recall and resumption directly.