Every agent that talks to someone more than once needs memory, and most agents get it wrong in the same way: they treat memory as one thing. In practice there are two problems with opposite requirements. Short-term state must be exact, ordered and complete enough to resume a conversation after a crash. Long-term memory must be small, curated and safe to show to the model in a future conversation that has nothing to do with this one.
This article builds the memory pattern around that split. It shows where each kind of memory hooks into a single agent turn, gives a small reference implementation, works through the token cost of different window policies, and defines the rule for promoting a fact from a thread into long-term memory. The lifecycle of long-term records (scoring, reconciliation, forgetting) is covered in agent memory architecture, and the memory service contract in agent memory systems. Here the subject is the pattern that ties them to the loop.
Two kinds of state with different contracts
| Property | Short-term thread state | Long-term memory |
|---|---|---|
| Scope | One conversation or run | A user, team or tenant |
| Content | Every message, tool call and result | A few hundred curated facts |
| Correctness | Exact; nothing may be lost or reordered | Lossy by design; wrong entries must be removable |
| Read pattern | Load the thread, build a window | Search by relevance to the current message |
| Write pattern | Append on every event, synchronously | Promote selectively, usually asynchronously |
| If lost | The run cannot resume | The agent forgets a preference |
| Main risk | Growing past the context window | Stale, wrong or injected facts |
Mixing the two causes most memory bugs. Store the whole transcript as long-term memory and recall fills with task chatter from old conversations. Keep user preferences only in thread state and the agent forgets them the moment a new thread starts. Keeping them apart also lets each use the storage that suits it: an ordered log for threads, a searchable store for memories.
Where memory hooks into the loop
A turn has four memory touch points. The incoming message is appended to the thread log before anything else, so a crash after this point can be resumed. The window policy turns the log into the recent part of the context. Recall searches long-term memory, scoped to the user, with the current message as the query. After the model and tools run, their events are appended. Promotion runs later, outside the turn, so it never adds latency to a reply.
import json, sqlite3
class ThreadLog:
# Short-term memory: an append-only, ordered event log per thread.
def __init__(self, path="threads.db"):
self.db = sqlite3.connect(path)
self.db.execute('''CREATE TABLE IF NOT EXISTS events(
seq INTEGER PRIMARY KEY AUTOINCREMENT, thread_id TEXT, turn_id TEXT, kind TEXT, body TEXT,
UNIQUE (thread_id, turn_id, kind))''')
def append(self, thread_id, turn_id, kind, body):
# kind is unique within a turn ("user", "assistant", "tool:1"), so a retried turn is a no-op
with self.db:
self.db.execute("INSERT OR IGNORE INTO events (thread_id, turn_id, kind, body) VALUES (?, ?, ?, ?)",
(thread_id, turn_id, kind, json.dumps(body)))
def events(self, thread_id):
rows = self.db.execute("SELECT turn_id, kind, body FROM events WHERE thread_id = ? ORDER BY seq",
(thread_id,))
return [(t, k, json.loads(b)) for t, k, b in rows]
def build_window(events, budget, count_tokens):
summary, recent = "", []
for turn_id, kind, body in events:
if kind == "summary": # a summary event replaces everything before it
summary, recent = body["text"], []
else:
recent.append((kind, body["text"]))
window, used = [], count_tokens(summary)
for kind, text in reversed(recent):
used += count_tokens(text)
if used > budget:
break
window.append((kind, text))
return summary, window[::-1]
def agent_turn(thread_id, user_id, turn_id, text, llm, log, ltm, count_tokens, budget=6000):
log.append(thread_id, turn_id, "user", {"text": text})
done = [b["text"] for t, k, b in log.events(thread_id) if t == turn_id and k == "assistant"]
if done: # retried turn: return the logged reply
return done[0]
summary, window = build_window(log.events(thread_id), budget, count_tokens)
memories = ltm.search(scope=user_id, query=text, k=5) # scope enforced by the store
reply = llm(system=SYSTEM, memories=memories, summary=summary, window=window)
log.append(thread_id, turn_id, "assistant", {"text": reply})
return reply
Short-term state: an append-only log, not a mutable transcript
The thread log is the source of truth for a conversation, and the context sent to the model is derived from it on every turn. Two rules keep it trustworthy. First, never edit history in place. Summaries, corrections and compaction results are new events that supersede older ones, so you can always reconstruct what the model saw and why. Second, make appends idempotent. Clients retry, workers crash after calling the model but before replying, and queues deliver twice. A unique key on thread, turn and event kind means a replayed turn writes nothing new.
Tool results deserve care. They are often the largest events, and a resumed run that loses them will repeat side-effecting calls. Store the result, or a reference to it, in the log before acting on it. The general technique of making each step replayable is covered in durable agent workflows, and how to compress old events without losing the thread is covered in context compaction.
Window policies and their token arithmetic
The window policy decides how much of the log the model sees. To compare policies, take a 40-turn conversation in which each turn, user message plus reply plus tool output, averages 800 tokens, and add up the input tokens across the whole conversation.
| Policy | Input at turn 40 | Total input over 40 turns | Prefix stable between turns? |
|---|---|---|---|
| Full buffer | 32,000 | 800 x (1 + 2 + ... + 40) = 656,000 | Yes, append-only |
| Last 8 turns | 6,400 | 800 x (36 + 32 x 8) = 233,600 | No, start shifts every turn |
| Summary + block of up to 8 turns, new summary every 8 | 7,400 | 176,000 plus about 28,600 for 4 summary calls = about 204,600 | Yes, within each block |
| Token budget (6,000) | about 6,000 | similar to last 8 turns | No |
Raw totals favour the windowed policies, and prompt caching changes the picture. Providers that cache a repeated prompt prefix charge less for cached tokens and serve them faster. A full buffer is append-only, so almost all of each turn's input is a cached prefix. A sliding window drops the oldest turn every time, so the prefix changes and the cache misses on nearly everything. Check your provider's caching rules and prices before deciding, because they decide which row is actually cheapest.
The summary row keeps the best of both because its window moves in blocks: keep the summary and the turns since the last summary fixed, append new turns, and only every eighth turn write a new summary event and restart the block. Within a block the prefix is stable and caches well, and the input never exceeds the summary plus one block. The cost is summary quality: whatever the summary leaves out is gone from the model's view, though still in the log.
Long-term memory: keep the interface narrow
The loop needs only three operations from long-term memory: search within a scope, upsert a fact with its source, and delete by id. Everything else, such as embedding, scoring by relevance and recency, merging duplicates and expiring stale facts, belongs behind that interface. A narrow interface lets you start with a table and a keyword search and move to a vector store later without touching the agent.
Show recalled memories to the model as a separate, labelled block with dates and sources, never mixed into the conversation. The model should be able to tell "the user said this today" from "we recorded this in March". Instruct it to prefer the current conversation when the two conflict, and to say so rather than silently following an old memory.
The promotion rule: what crosses from thread to long-term
Promotion decides what an agent will remember for years, so give it an explicit rule. A fact is promoted only if all four conditions hold: the user stated or confirmed it, rather than the model inferring it; it will still be true next week; it is about the user, their preferences or their environment rather than this task's progress; and it is not a secret, credential, payment detail or information about a third party.
Run promotion when a thread goes idle, for example after thirty minutes without a message, because many threads never formally end. Use a model to propose candidates, then validate them in code: the cited evidence must be a real user turn in this thread, and anything matching secret patterns is dropped. An explicit request such as "remember that I prefer aisle seats" can be promoted immediately through a tool, since the user has confirmed it.
import json, re
PROMOTE = '''List facts about the user worth remembering in FUTURE conversations.
Include a fact only if the user stated or confirmed it, it will still be true next week,
and it concerns the user, their preferences or their environment, not this task's progress.
Exclude secrets, credentials, payment data and facts about other people.
Return JSON: [{"fact": "...", "evidence_turn": "<turn id>"}] or [].'''
SECRET = re.compile(r"(password|api[_ ]?key|token|\b\d{13,19}\b)", re.I)
def promote(thread_id, user_id, llm, log, ltm):
events = log.events(thread_id)
user_turns = {t for t, kind, _ in events if kind == "user"}
for c in json.loads(llm(PROMOTE + "\n\n" + render_transcript(events))):
if c.get("evidence_turn") not in user_turns or SECRET.search(c.get("fact", "")):
continue # unverifiable or sensitive: drop
ltm.upsert(scope=user_id, text=c["fact"], source=f"{thread_id}/{c['evidence_turn']}")
Worked example: a travel assistant over two sessions
On Monday a user asks a travel assistant to find flights to Lisbon for the 14th, mentions they are vegetarian, and books an itinerary. The thread log holds the messages, three large flight-search results and the booking confirmation. Half an hour later the thread goes idle and promotion runs. The model proposes three candidates: the user is vegetarian, the user is flying to Lisbon on the 14th, and the user prefers morning flights. Validation keeps the first, which the user stated. The second is task progress, and the booking system, not memory, is the source of truth for it. The third was the model's inference from one choice, so it fails the stated-or-confirmed test.
On Thursday the user opens a new thread and asks for dinner suggestions near their hotel. The window is empty, recall returns "vegetarian" with Monday's date and source, and the agent calls the booking tool to find the hotel address. The answer is correct, and nothing from Monday's flight search entered the context.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Duplicate messages after retries | Non-idempotent appends | Unique key on thread, turn and event kind |
| Resumed run repeats a payment call | Tool result not logged before acting | Log results first; replay from the log |
| Costs jump after adding a window | Sliding window defeats prompt caching | Move the window in blocks, or keep the full buffer while it fits |
| Agent cites old task details | Transcripts promoted wholesale | Apply the promotion rule; keep task state in its system of record |
| Agent follows a wrong memory | Inferred fact promoted, no source | Stated-or-confirmed rule; show dates and sources |
| One user's facts appear for another | Scope chosen by the model | Enforce scope in the store from the authenticated user |
Memory is also an attack surface: text in a tool result can try to plant instructions that later get promoted. The evidence check above, which accepts only user turns, closes the most direct path. For long-running agents that span many context windows, memory architectures for long-running agents extends the pattern with task-level stores.
What to do next
- Separate your agent's storage into a thread log and a long-term store, even if both start as tables in one database.
- Make thread appends idempotent with a turn id, and log tool results before acting on them.
- Measure tokens per turn and cache hit rate, then choose a window policy using the arithmetic above.
- Write your promotion rule down and enforce it in code, not only in the prompt.
- Show recalled memories as a dated, sourced block and tell the model to prefer the current conversation.
- Give users a way to see and delete what the agent remembers about them.