A language model has no memory between calls. Everything it knows about the user, the conversation, your documents and its tools must be in the tokens you send on each request, and the window is finite. Context management is the discipline of deciding, for every call, what goes in, in what order, and what stays out. It matters as much as model choice. A strong model given a stale policy and no order record will confidently give a wrong answer.
Most teams start with string concatenation and grow a pile of special cases. This article treats context as the output of a component, a context assembler, with typed inputs, a budget, explicit filtering, scoring and placement, and a manifest of every decision so that bad answers can be traced and replayed. It is provider-neutral, and the worked example's numbers come from running the code shown. Token budgeting, the overflow ladder and summarisation have their own page, context window management. This one is about the machinery around them.
The architecture
Why context needs a component
Three properties justify a dedicated component. First, one place to reason about. Context is assembled from four or five sources owned by different teams: the conversation log, a memory store, a retriever, tool results, and the system prompt and tool schemas. If each appends its own text, nobody owns the total, and the first long conversation overflows. Second, attribution. When an answer is wrong, the first question is "what did the model see?". Without a record, you cannot tell a retrieval miss from a model error. Third, testability. A pure function from (request, candidate items) to (placed prompt, manifest) can be unit-tested and replayed. A tangle of f-strings cannot.
The contract is simple. The assembler receives candidate items. Each one has an id, a kind, its text, its token count measured with the target model's own tokenizer, a value score for this request, and a pinned flag. It returns the placed sequence plus a manifest. It never calls the model and never fetches data itself, which keeps it deterministic. Measure tokens with the real tokenizer, because character-based estimates drift by language and content; see tokenization.
Filter before you score
Filter before you score, because filtering is about correctness and safety, not space. Four filters belong here. Tenant and permission: an item the caller may not see must never be a candidate, however relevant it is. A budget-based drop is not access control. Freshness and supersession: when policy v3 replaces v2, v2 must be removed, not merely ranked lower, because a lower-ranked contradiction that still fits will be read. Deduplication: the same chunk often arrives twice, once from the retriever and once quoted in an earlier turn. Hash the normalised text and keep one copy. Validity: tool results that errored, timed out or returned an empty body should become one short line, such as "order lookup failed", rather than a raw stack trace.
Of the four, supersession is the one most systems lack, and the worked example below shows what it costs. A pure relevance score cannot express "this is true but outdated".
Scoring and packing under a budget
After filtering, the problem is a knapsack: pick items whose total tokens fit the budget while maximising value. The budget is the window minus a reserved output allowance. Forgetting that reserve is the classic bug where a request fits but the answer is cut off. Pinned items (the system prompt, tool schemas, the current message, the last turn or two) go in unconditionally, and if they alone exceed the budget, the assembler fails loudly instead of truncating the system prompt. The rest are taken greedily by value per token. Greedy is not optimal for knapsack, but it is fast, explainable and close enough when items are small relative to the budget. Where one large item matters, consider splitting it into chunks rather than adding an exact solver.
import hashlib, json
from dataclasses import dataclass
@dataclass
class Item:
id: str
kind: str # system | tools | memory | retrieved | tool_result | turn | user
text: str
tokens: int # counted with the target model's tokenizer
value: float # usefulness for THIS request, 0..1
pinned: bool = False
seq: int = 0 # conversation position for turns and tool results
PREFIX = ("system", "tools", "memory") # stable kinds, fixed order
def digest(text):
return hashlib.sha256(text.encode()).hexdigest()
def pack(items, window, reserve_output):
budget = window - reserve_output
chosen = [i for i in items if i.pinned]
used = sum(i.tokens for i in chosen)
if used > budget:
raise ValueError(f"pinned context needs {used} tokens; budget is {budget}")
seen, dropped = {digest(i.text) for i in chosen}, []
for it in sorted((i for i in items if not i.pinned),
key=lambda i: i.value / i.tokens, reverse=True):
if digest(it.text) in seen:
dropped.append((it.id, "duplicate"))
elif used + it.tokens > budget:
dropped.append((it.id, "budget"))
else:
chosen.append(it); seen.add(digest(it.text)); used += it.tokens
return place(chosen), used, dropped
def place(chosen):
def of(kind):
return [i for i in chosen if i.kind == kind]
prefix = [i for k in PREFIX for i in sorted(of(k), key=lambda i: i.id) if i.pinned]
notes = [i for i in of("memory") if not i.pinned]
docs = sorted(of("retrieved"), key=lambda i: i.value, reverse=True)
docs = docs[0::2] + docs[1::2][::-1] # strongest at both edges, weakest in the middle
convo = sorted(of("turn") + of("tool_result"), key=lambda i: i.seq)
return prefix + notes + docs + convo + of("user")
def manifest(placed, used, dropped, window):
entries = [{"id": i.id, "tokens": i.tokens, "sha": digest(i.text)[:12]} for i in placed]
return {"window": window, "used": used, "items": entries, "dropped": dropped,
"context_hash": digest(json.dumps(entries))[:16]}The value score is where judgement lives. Start simple: retriever similarity for documents, a recency decay for turns, and a fixed high value for memory facts the user stated explicitly. Then let evaluation, not intuition, adjust the weights.
Placement: cache stability and position
Order matters for two independent reasons. The first is caching. Many serving stacks and hosted APIs reuse computation for a shared prompt prefix, so a byte-stable prefix makes repeated calls cheaper and faster. That is why place() puts only pinned system, tool and memory items first, in a fixed order, and moves optional memory notes after them. Including or excluding an optional note must not change the prefix bytes. Never put timestamps, request ids or per-user data ahead of shared content. Mechanisms and limits differ by provider; see prompt caching.
The second is attention position. Liu et al.'s "Lost in the Middle" study found that models used relevant information best when it sat at the beginning or end of a long context, and noticeably worse when it sat in the middle. The docs[0::2] + docs[1::2][::-1] line puts the strongest document first, the second strongest last, and the weakest in the middle. The current user message goes at the very end, so the question is the last thing the model reads. Treat the effect as a tendency to measure on your own model, not a law.
Worked example: the refund question
Here is a support assistant on a model with a hypothetical 16,000-token window, reserving 2,000 for the answer, which leaves a budget of 14,000. The user asks: "Can I still get a refund for order 4471?" The pinned items are the system prompt (900 tokens), tool schemas (1,600), the customer profile (300), the last two turns (600 and 450) and the question (250), for a total of 4,100. The candidates are:
| Item | Tokens | Value | Value per 1k tokens |
|---|---|---|---|
| mem:prefers-email | 120 | 0.50 | 4.17 |
| doc:order-history | 1,200 | 0.85 | 0.71 |
| turn:7 | 700 | 0.45 | 0.64 |
| doc:refund-policy (v3) | 1,800 | 0.92 | 0.51 |
| doc:refund-policy-copy | 1,800 | 0.90 | 0.50 |
| turn:5 | 900 | 0.30 | 0.33 |
| doc:policy-2024 (v2, superseded) | 2,000 | 0.35 | 0.18 |
| tool:order-lookup | 3,200 | 0.55 | 0.17 |
| doc:shipping-faq | 2,400 | 0.40 | 0.17 |
Run 1, without a supersession filter. The packer drops the policy copy as a duplicate, admits the 2024 policy and the shipping FAQ, and then has no room for the 3,200-token order lookup. It uses 13,220 tokens. The model now sees two contradictory refund policies and no live lookup of order 4471's current state. That is exactly the setup for a confident wrong answer, and nothing overflowed.
Run 2, with v2 removed by the freshness filter. Same code and same budget. The order lookup now fits, the shipping FAQ is dropped for budget, and usage is 12,020 tokens across 12 items. The placed order is: system, tools, profile, the email note, refund policy v3, order history, then turns 5, the order lookup, turns 7, 8 and 9, and finally the question. Note the lesson. The fix was not a bigger window or a better score but a filter, and only the manifest diff between the two runs makes the cause visible.
The manifest: provenance, replay and ablation
The manifest records what the model saw: item ids, token counts, content hashes, dropped items with reasons, and a hash of the whole context. Log it with the request id and the model response, and store the item texts by hash in a content-addressed store, so the exact prompt can be rebuilt without logging duplicate copies of every document. Apply the same retention and redaction rules as the conversation itself, because a manifest store is a copy of user data.
This record enables three things. Debugging: a complaint becomes "retrieval returned v2" or "the order lookup was dropped for budget" rather than "the model hallucinated". Replay: rebuild past contexts, change one thing (scoring weights, a filter, the model version) and diff the answers on a fixed set of real cases. Ablation: remove one item kind at a time to learn what actually earns its tokens. Teams are often surprised how little the third and fourth retrieved chunks contribute. For long-running agents, apply the same discipline to each step; see context engineering for agents.
The write path: what to remember
Reading context is half the job; the other half is deciding what to keep for later calls. Every turn produces material that could become a memory item: a stated preference, an order number, a resolved question. Write memory deliberately, not by saving whole transcripts. A useful record holds a short normalised statement ("prefers email contact"), its source turn id, a timestamp, a confidence, and an expiry or review date. Statements the user made explicitly deserve more trust than facts the model inferred, so store the two kinds separately and score them differently.
Updates need the same supersession logic as documents. When the user says they have moved, the old address must be marked replaced, not left to compete with the new one. Give users a way to see and delete what is remembered, and apply the same tenant keys and retention rules as the conversation log. Finally, keep memory writes out of the assembler. A separate step after the response, with its own tests, keeps the assembler a pure function and makes a bad memory write easy to find in the log.
Failure modes
- No output reserve. The prompt fits, and the answer is truncated mid-sentence.
- Estimated tokens. A characters-divided-by-four rule undercounts code and non-English text, so the real request overflows and the provider rejects or truncates it.
- Unstable prefix. A timestamp or an optional note at the top silently stops cache reuse, and latency and cost rise without any error.
- Cross-tenant leakage. Permissions are enforced after packing, or a shared cache is keyed without the tenant. Filter first and key everything by tenant.
- Stale truths. Superseded documents and old memory facts ("lives in Pune" after a move) outrank nothing but still fit. Version and expire them.
- Tool-result bloat. One verbose JSON payload evicts the conversation. Project tool output to the fields the task needs before it becomes a candidate.
- Silent drops. Something important is dropped for budget and nobody knows. Alert on drops of high-value items.
Trade-offs
Bigger windows versus selection. Long-context models reduce overflow but not the need to choose. More tokens cost more, add latency, and dilute attention. Selection usually beats stuffing, as context length explains. Greedy versus optimal packing. Exact knapsack is rarely worth its complexity; chunking large items helps more. Pinned versus scored. Each pin makes behaviour predictable and shrinks the room left for evidence, so keep the pin list short and reviewed. Cache stability versus freshness. A frozen prefix is cheap, but it cannot carry per-request facts, so keep volatile content late. Manifest cost versus blindness. Logging by hash is cheap; debugging without it is not.
What to do next
- Wrap every model call in one assembler function that takes typed candidate items and returns placed items plus a manifest.
- Count tokens with the target model's tokenizer and reserve an explicit output allowance.
- Add the four filters (tenant and ACL, supersession, dedupe, validity) before any scoring.
- Make the prefix byte-stable: pinned system, tools and memory first, volatile content last, the question at the end.
- Log manifests with content hashes and build a replay set of 50 to 100 real requests.
- Run one ablation per item kind and cut what does not move answer quality.
- Alert on budget drops of high-value items and on a falling cache-hit rate.