Every agent that talks to someone more than once needs memory, and most agents get it wrong in the same way: they treat memory as one thing. In practice there are two problems with opposite requirements. Short-term state must be exact, ordered and complete enough to resume a conversation after a crash. Long-term memory must be small, curated and safe to show to the model in a future conversation that has nothing to do with this one.

This article builds the memory pattern around that split. It shows where each kind of memory hooks into a single agent turn, gives a small reference implementation, works through the token cost of different window policies, and defines the rule for promoting a fact from a thread into long-term memory. The lifecycle of long-term records (scoring, reconciliation, forgetting) is covered in agent memory architecture, and the memory service contract in agent memory systems. Here the subject is the pattern that ties them to the loop.

Advertisement

Two kinds of state with different contracts

PropertyShort-term thread stateLong-term memory
ScopeOne conversation or runA user, team or tenant
ContentEvery message, tool call and resultA few hundred curated facts
CorrectnessExact; nothing may be lost or reorderedLossy by design; wrong entries must be removable
Read patternLoad the thread, build a windowSearch by relevance to the current message
Write patternAppend on every event, synchronouslyPromote selectively, usually asynchronously
If lostThe run cannot resumeThe agent forgets a preference
Main riskGrowing past the context windowStale, wrong or injected facts

Mixing the two causes most memory bugs. Store the whole transcript as long-term memory and recall fills with task chatter from old conversations. Keep user preferences only in thread state and the agent forgets them the moment a new thread starts. Keeping them apart also lets each use the storage that suits it: an ordered log for threads, a searchable store for memories.

Where memory hooks into the loop

One agent turn, and where the two kinds of memory hook inUser messageturn_id attachedThread log appendshort-term, exactWindow policysummary + recent turnsRecallsearch(user scope, query)Contextsystem + memories + windowModel and toolsreply, tool callsThread log appendassistant and tool eventsPromotion (async, on idle)extract, validate, upsertLong-term storeper user, curated, lossyYellow: thread state, owned by the run. Purple: long-term memory, owned by the user scope.
Figure 1. Thread state is written on every event and read through a window policy. Long-term memory is read by relevance at the start of a turn and written only through promotion.

A turn has four memory touch points. The incoming message is appended to the thread log before anything else, so a crash after this point can be resumed. The window policy turns the log into the recent part of the context. Recall searches long-term memory, scoped to the user, with the current message as the query. After the model and tools run, their events are appended. Promotion runs later, outside the turn, so it never adds latency to a reply.

import json, sqlite3

class ThreadLog:
    # Short-term memory: an append-only, ordered event log per thread.
    def __init__(self, path="threads.db"):
        self.db = sqlite3.connect(path)
        self.db.execute('''CREATE TABLE IF NOT EXISTS events(
            seq INTEGER PRIMARY KEY AUTOINCREMENT, thread_id TEXT, turn_id TEXT, kind TEXT, body TEXT,
            UNIQUE (thread_id, turn_id, kind))''')

    def append(self, thread_id, turn_id, kind, body):
        # kind is unique within a turn ("user", "assistant", "tool:1"), so a retried turn is a no-op
        with self.db:
            self.db.execute("INSERT OR IGNORE INTO events (thread_id, turn_id, kind, body) VALUES (?, ?, ?, ?)",
                            (thread_id, turn_id, kind, json.dumps(body)))

    def events(self, thread_id):
        rows = self.db.execute("SELECT turn_id, kind, body FROM events WHERE thread_id = ? ORDER BY seq",
                               (thread_id,))
        return [(t, k, json.loads(b)) for t, k, b in rows]


def build_window(events, budget, count_tokens):
    summary, recent = "", []
    for turn_id, kind, body in events:
        if kind == "summary":                 # a summary event replaces everything before it
            summary, recent = body["text"], []
        else:
            recent.append((kind, body["text"]))
    window, used = [], count_tokens(summary)
    for kind, text in reversed(recent):
        used += count_tokens(text)
        if used > budget:
            break
        window.append((kind, text))
    return summary, window[::-1]


def agent_turn(thread_id, user_id, turn_id, text, llm, log, ltm, count_tokens, budget=6000):
    log.append(thread_id, turn_id, "user", {"text": text})
    done = [b["text"] for t, k, b in log.events(thread_id) if t == turn_id and k == "assistant"]
    if done:                                                    # retried turn: return the logged reply
        return done[0]
    summary, window = build_window(log.events(thread_id), budget, count_tokens)
    memories = ltm.search(scope=user_id, query=text, k=5)      # scope enforced by the store
    reply = llm(system=SYSTEM, memories=memories, summary=summary, window=window)
    log.append(thread_id, turn_id, "assistant", {"text": reply})
    return reply
Advertisement

Short-term state: an append-only log, not a mutable transcript

The thread log is the source of truth for a conversation, and the context sent to the model is derived from it on every turn. Two rules keep it trustworthy. First, never edit history in place. Summaries, corrections and compaction results are new events that supersede older ones, so you can always reconstruct what the model saw and why. Second, make appends idempotent. Clients retry, workers crash after calling the model but before replying, and queues deliver twice. A unique key on thread, turn and event kind means a replayed turn writes nothing new.

Tool results deserve care. They are often the largest events, and a resumed run that loses them will repeat side-effecting calls. Store the result, or a reference to it, in the log before acting on it. The general technique of making each step replayable is covered in durable agent workflows, and how to compress old events without losing the thread is covered in context compaction.

Window policies and their token arithmetic

The window policy decides how much of the log the model sees. To compare policies, take a 40-turn conversation in which each turn, user message plus reply plus tool output, averages 800 tokens, and add up the input tokens across the whole conversation.

PolicyInput at turn 40Total input over 40 turnsPrefix stable between turns?
Full buffer32,000800 x (1 + 2 + ... + 40) = 656,000Yes, append-only
Last 8 turns6,400800 x (36 + 32 x 8) = 233,600No, start shifts every turn
Summary + block of up to 8 turns, new summary every 87,400176,000 plus about 28,600 for 4 summary calls = about 204,600Yes, within each block
Token budget (6,000)about 6,000similar to last 8 turnsNo

Raw totals favour the windowed policies, and prompt caching changes the picture. Providers that cache a repeated prompt prefix charge less for cached tokens and serve them faster. A full buffer is append-only, so almost all of each turn's input is a cached prefix. A sliding window drops the oldest turn every time, so the prefix changes and the cache misses on nearly everything. Check your provider's caching rules and prices before deciding, because they decide which row is actually cheapest.

The summary row keeps the best of both because its window moves in blocks: keep the summary and the turns since the last summary fixed, append new turns, and only every eighth turn write a new summary event and restart the block. Within a block the prefix is stable and caches well, and the input never exceeds the summary plus one block. The cost is summary quality: whatever the summary leaves out is gone from the model's view, though still in the log.

Long-term memory: keep the interface narrow

The loop needs only three operations from long-term memory: search within a scope, upsert a fact with its source, and delete by id. Everything else, such as embedding, scoring by relevance and recency, merging duplicates and expiring stale facts, belongs behind that interface. A narrow interface lets you start with a table and a keyword search and move to a vector store later without touching the agent.

Show recalled memories to the model as a separate, labelled block with dates and sources, never mixed into the conversation. The model should be able to tell "the user said this today" from "we recorded this in March". Instruct it to prefer the current conversation when the two conflict, and to say so rather than silently following an old memory.

The promotion rule: what crosses from thread to long-term

Promotion decides what an agent will remember for years, so give it an explicit rule. A fact is promoted only if all four conditions hold: the user stated or confirmed it, rather than the model inferring it; it will still be true next week; it is about the user, their preferences or their environment rather than this task's progress; and it is not a secret, credential, payment detail or information about a third party.

Run promotion when a thread goes idle, for example after thirty minutes without a message, because many threads never formally end. Use a model to propose candidates, then validate them in code: the cited evidence must be a real user turn in this thread, and anything matching secret patterns is dropped. An explicit request such as "remember that I prefer aisle seats" can be promoted immediately through a tool, since the user has confirmed it.

import json, re

PROMOTE = '''List facts about the user worth remembering in FUTURE conversations.
Include a fact only if the user stated or confirmed it, it will still be true next week,
and it concerns the user, their preferences or their environment, not this task's progress.
Exclude secrets, credentials, payment data and facts about other people.
Return JSON: [{"fact": "...", "evidence_turn": "<turn id>"}] or [].'''

SECRET = re.compile(r"(password|api[_ ]?key|token|\b\d{13,19}\b)", re.I)

def promote(thread_id, user_id, llm, log, ltm):
    events = log.events(thread_id)
    user_turns = {t for t, kind, _ in events if kind == "user"}
    for c in json.loads(llm(PROMOTE + "\n\n" + render_transcript(events))):
        if c.get("evidence_turn") not in user_turns or SECRET.search(c.get("fact", "")):
            continue                                    # unverifiable or sensitive: drop
        ltm.upsert(scope=user_id, text=c["fact"], source=f"{thread_id}/{c['evidence_turn']}")

Worked example: a travel assistant over two sessions

On Monday a user asks a travel assistant to find flights to Lisbon for the 14th, mentions they are vegetarian, and books an itinerary. The thread log holds the messages, three large flight-search results and the booking confirmation. Half an hour later the thread goes idle and promotion runs. The model proposes three candidates: the user is vegetarian, the user is flying to Lisbon on the 14th, and the user prefers morning flights. Validation keeps the first, which the user stated. The second is task progress, and the booking system, not memory, is the source of truth for it. The third was the model's inference from one choice, so it fails the stated-or-confirmed test.

On Thursday the user opens a new thread and asks for dinner suggestions near their hotel. The window is empty, recall returns "vegetarian" with Monday's date and source, and the agent calls the booking tool to find the hotel address. The answer is correct, and nothing from Monday's flight search entered the context.

Failure modes

SymptomCauseFix
Duplicate messages after retriesNon-idempotent appendsUnique key on thread, turn and event kind
Resumed run repeats a payment callTool result not logged before actingLog results first; replay from the log
Costs jump after adding a windowSliding window defeats prompt cachingMove the window in blocks, or keep the full buffer while it fits
Agent cites old task detailsTranscripts promoted wholesaleApply the promotion rule; keep task state in its system of record
Agent follows a wrong memoryInferred fact promoted, no sourceStated-or-confirmed rule; show dates and sources
One user's facts appear for anotherScope chosen by the modelEnforce scope in the store from the authenticated user

Memory is also an attack surface: text in a tool result can try to plant instructions that later get promoted. The evidence check above, which accepts only user turns, closes the most direct path. For long-running agents that span many context windows, memory architectures for long-running agents extends the pattern with task-level stores.

What to do next

  1. Separate your agent's storage into a thread log and a long-term store, even if both start as tables in one database.
  2. Make thread appends idempotent with a turn id, and log tool results before acting on them.
  3. Measure tokens per turn and cache hit rate, then choose a window policy using the arithmetic above.
  4. Write your promotion rule down and enforce it in code, not only in the prompt.
  5. Show recalled memories as a dated, sourced block and tell the model to prefer the current conversation.
  6. Give users a way to see and delete what the agent remembers about them.
Key takeaway: The memory pattern is two persistence problems, not one. Short-term thread state is an exact, append-only, idempotent log from which each turn's context window is derived. Long-term memory is a small, curated, user-scoped store that is searched at the start of a turn and written only by promotion. Choose a window policy with both token totals and prompt caching in mind; a summary plus a block-moving window often balances them. Promote only facts the user stated, that last, that concern the user and are not secrets, and validate that rule in code.