A chat assistant that remembers your name and preferred tone has a memory problem, but a small one. An agent migrating 400 files from one framework to another over three days has a much larger one. It will fill and reset its context window dozens of times, be interrupted and resumed, discover facts about the codebase that it must not forget, make decisions whose reasons matter later, and try approaches that failed and must not be tried again. Its memory is not a personalisation feature. It is what makes the task finishable at all.
Much of what is written about agent memory covers the conversational case: working, episodic and semantic memory, consolidation, forgetting and user privacy, described well in agent memory layers compared. This article covers task memory for long-horizon agents. It proposes four stores with different contracts, a record format with provenance and supersession, a write path that validates what the agent tries to remember, a read path that assembles each context window under a token budget, and tests that show whether the memory actually works.
Why the context window is not the memory
The transcript of a long run is a log, not a memory. It is too large to reload, it mixes important facts with noise, and it records beliefs the agent later corrected without marking them as wrong. Compressing it helps within one window, as covered in context compaction for long-running agents, but a summary of a summary drifts. After a few rounds, the specific facts the agent needs, such as the exact build command or why a module was left alone, are the first to blur.
What a long-running agent needs to carry across windows is a small set of things with very different shapes. The goal and constraints never change and must never be lost. The plan and its progress change constantly and must be exactly right. Decisions accumulate and must keep their reasons. Facts about the environment accumulate, and some become false. The raw history is rarely needed but occasionally essential. Putting all of these in one vector store, retrieved by similarity, treats them as the same kind of thing, and that is the most common design mistake.
Four stores with different contracts
| Store | Holds | Write rule | Read rule |
|---|---|---|---|
| Task ledger | Goal, constraints, plan items with status, open questions | Structured updates through tools; every status change logged | Pinned: goal and constraints always; current plan slice always |
| Decision log | Choices made, alternatives rejected, reasons, who decided | Append-only; a reversal is a new entry that references the old one | By scope: decisions touching the current work item |
| Fact store | Claims about the environment, with source and validity | Validated: must cite a source; conflicting facts supersede, never overwrite | Filtered by scope, then ranked by relevance and recency |
| Episodic archive | Raw transcripts, tool calls and outputs | Everything, automatically, outside the model's control | Search only, on demand; never loaded whole |
The split exists because each store fails differently. A ledger that is slightly wrong sends the agent to redo finished work or skip unfinished work. A decision log without reasons invites the agent to relitigate settled questions in every new window. A fact store without provenance cannot tell a verified observation from something the model guessed. Keeping them separate lets each have the right integrity rule and the right retrieval rule.
Storage technology is a secondary choice. A relational database handles the ledger, decisions and facts comfortably, with a vector index added for fact retrieval; object storage holds the archive. The trade-offs between vector store options are covered in memory and vector store options.
Records with provenance and supersession
The fact store carries most of the risk, so its records need more structure than a line of text and an embedding. Each fact states what kind of claim it is, where it came from, when it was observed, what it applies to, and whether a newer fact replaced it.
from dataclasses import dataclass, field
from typing import Optional
@dataclass
class Fact:
id: str
text: str # "Build command is `make build-ci`, not `make build`"
kind: str # observation | inference | instruction-from-user
scope: list[str] # ["repo:billing", "path:services/invoice/"]
source: str # "tool:run_shell#call-8812" or "user:msg-41"
source_hash: Optional[str] # hash of the file or output the fact was read from
observed_at: str
confidence: float # set by the validator from kind and source, not by the model
supports: list[str] = field(default_factory=list) # fact ids an inference rests on
supersedes: Optional[str] = None
superseded_by: Optional[str] = NoneThree rules keep this honest. First, the kind separates what the agent saw from what it concluded. An observation cites a tool call whose output is in the archive; an inference cites the facts it was drawn from. Inferences are ranked below observations and never used as the source of another observation. Second, facts are never edited or deleted during a run. When a newer observation contradicts an old one, the new fact supersedes it, and the chain shows when and why belief changed. Third, source_hash lets the system detect staleness: if the file a fact was read from has changed since, the fact is flagged for re-verification before it is served.
The write path: what deserves to be remembered
If the agent can write anything to memory, it will write too much, and some of it will be wrong. Give it narrow tools instead of a free-form notes field: record_fact, record_decision, update_plan_item, add_open_question. Each tool call goes through a validator that the model does not control.
def validate_fact(proposed, archive, store):
if proposed.kind == "observation":
call = archive.get(proposed.source)
if call is None:
return reject("observation must cite a recorded tool call")
if call.untrusted_content and looks_like_instruction(proposed.text):
return reject("instructions from tool output are not facts")
if proposed.kind == "inference" and not proposed.supports:
return reject("inference must reference supporting fact ids")
proposed.confidence = CONFIDENCE[(proposed.kind, source_type(proposed.source))]
for old in store.same_subject(proposed, scope=proposed.scope):
if contradicts(old, proposed):
if proposed.observed_at <= old.observed_at:
return reject(f"older than current belief {old.id}")
proposed.supersedes = old.id
if store.near_duplicate(proposed):
return merge_reference(proposed)
return accept(proposed)The contradiction check is the expensive part. A cheap approach narrows candidates by scope and embedding similarity, then asks a small model whether the two statements can both be true. The rest is deterministic. Rejections go back to the agent as tool errors, which teaches it within the run what is worth recording.
The instruction check matters more than it looks. A web page or document the agent reads can contain text that, stored as a fact, becomes a durable instruction retrieved into every future window. Memory turns a one-off prompt injection into a persistent one. Tool output is data; it can support facts, but it should never become a stored directive.
The read path: assembling each window under a budget
Every new context window, and every step within one, is built by an assembler with a fixed layout and a token budget per section. A fixed layout matters because the model learns where to look, and because budgets per section stop one growing store from crowding out the others.
BUDGET = {"pinned": 1500, "plan": 2000, "decisions": 1500, "facts": 3000, "tail": 4000}
def assemble(task, item, stores, tok):
parts = []
parts.append(fit(render_goal_and_constraints(task), BUDGET["pinned"], tok, must_fit=True))
parts.append(fit(render_plan(stores.ledger, focus=item, done_window=5), BUDGET["plan"], tok))
parts.append(fit(render(stores.decisions.for_scope(item.scope)), BUDGET["decisions"], tok))
facts = stores.facts.current(scope=item.scope) # superseded facts excluded
facts = [f for f in facts if not stale(f)] + reverify_queue(facts)
ranked = rank(facts, query=item.description,
w_relevance=0.6, w_recency=0.25, w_confidence=0.15)
parts.append(fit(render(ranked), BUDGET["facts"], tok))
parts.append(fit(render_recent_steps(stores.archive, n=8), BUDGET["tail"], tok))
return "\n\n".join(parts)Scope filtering comes before similarity. Similarity alone retrieves facts about a different module that happen to use the same words; filtering by the work item's scope first makes retrieval precise, and similarity then orders what is left. The ranking weights echo the recency, importance and relevance score from the Generative Agents paper (Park et al., 2023), with confidence standing in for importance; tune them on your own recall tests rather than copying these numbers. The paging approach of MemGPT (Packer et al., 2023), where the agent itself moves information between context and external storage, is a complementary design; the assembler here keeps paging decisions in deterministic code and leaves the model only the write tools and a search tool over the archive.
Handoffs between context windows
The moment one context window ends and the next begins is where long runs usually lose the thread. Treat it as a checkpoint with three steps. First, the agent writes a short handoff note: what it was doing, what it expected next, and anything it had noticed but not yet recorded. Second, the system reconciles the ledger with ground truth. For a code migration that means checking which files actually changed in version control and which tests pass, and correcting plan items whose recorded status disagrees. Third, the next window is assembled from the stores, with the handoff note in the tail section.
The reconciliation step is what keeps memory from drifting away from reality. The agent's belief that it finished a file is a claim; the diff is evidence. When they disagree, evidence wins, and the disagreement itself is recorded as a fact, because it usually means a step failed silently.
Worked example: a three-day framework migration
An agent migrates a service of 400 source files from one web framework to another. The ledger holds one plan item per module, 38 in all, each with a status. On day one, running the tests with make build fails; the agent reads the CI configuration, finds make build-ci, and records an observation citing that tool call. On day two, a teammate changes the Makefile. The fact's source_hash no longer matches, so the assembler queues it for re-verification, the agent re-reads the file, and a new fact supersedes the old one.
| Memory write | Store | Why it pays off later |
|---|---|---|
| Decision: keep the legacy auth middleware; the new framework's version drops a header the mobile app needs | Decision log | In window 17 the agent is about to replace it again; the scoped decision is retrieved and it stops |
| Fact: the invoice module's tests need a local database container | Fact store, scope invoice | Retrieved only when working on invoice, not in every window |
| Open question: are the admin routes still used? | Ledger | Surfaces at the end of the run for a human instead of being silently skipped |
| Failed approach: automated rewrite of template tags breaks escaping | Decision log, marked rejected | Prevents the same dead end in every new window |
Without the decision log, the most expensive failure in long runs appears: the agent rediscovers a problem, forgets its earlier conclusion, and re-applies a change it previously reverted. Recorded reasons, retrieved by scope, are what break that loop.
Failure modes
- Memory poisoning. Injected instructions stored as facts and replayed into every window. Reject instruction-shaped text from untrusted sources and keep the source on every record.
- Self-reinforcing inference. A guess stored as a fact, later cited as evidence for another fact. Separate kinds and never let inferences stand as sources for observations.
- Stale facts. True when recorded, false now. Hash sources and re-verify on change; prefer re-reading a file to trusting a memory of it.
- Ledger drift. Recorded progress disagrees with the world. Reconcile against ground truth at every handoff.
- Crowding. One store grows until it pushes out the plan. Per-section budgets and scope filters.
- Retrieval misses. The fact exists but is not retrieved. Measure with recall probes, below, rather than guessing.
Evaluating memory
Memory designs are easy to argue about and cheap to test. Three tests cover most of what matters. Recall probes: at fixed points in a recorded run, ask questions whose answers were established earlier, such as the build command or why a module was skipped, and score whether the assembled context lets the model answer correctly. Resume tests: kill the agent at random points, restart it from memory alone, and compare the final result and the amount of repeated work with an uninterrupted run. Contradiction audits: sample served facts and check them against ground truth to measure the stale-fact rate.
Track context tokens spent on memory per step alongside these scores. A design that improves recall by loading everything is not an improvement, and the budgets in the assembler are the dial you are tuning.
What to do next
- Write down what your agent must carry across windows and sort it into ledger, decisions, facts and archive.
- Record every tool call and output in an archive the model cannot edit.
- Replace free-form notes with narrow memory tools and a validator that requires a cited source.
- Add kind, scope, source hash and supersession to fact records, and never overwrite a fact during a run.
- Build a context assembler with a fixed layout, per-section budgets and scope filtering before similarity.
- Reconcile the ledger against ground truth at every window handoff.
- Reject instruction-shaped text from untrusted tool output at the write path.
- Build recall probes and resume-from-kill tests, and use them to tune budgets and ranking weights.