Most prompt-injection writing is about what an attacker says. Context stuffing is about how much they say. Every LLM call has a fixed input budget, and every application has some code that decides what goes into that budget when there is more material than room. An attacker who can make one input very large, very repetitive or very numerous can change what the model sees without writing a single clever instruction: the system prompt falls off the front, the policy reminder gets buried in the middle, honest retrieved documents are pushed out by near-copies, or the bill for one request grows a hundredfold.

This article treats the context window as an allocated resource, like memory, and shows how allocation bugs become security bugs: four families of volume attack, a truncation bug in a support agent, a defensive context assembler, and detection signals. Boundary attacks, where text pretends to be a different role, are covered in context smuggling.

The context window is an allocation problem

A model accepts at most some number of tokens per request: its context length minus whatever you reserve for the output. Your application fills that space from several sources: the system prompt, tool definitions, prior turns, retrieved documents, tool results and the new user message. Some of those sources are written by you and are trusted. Others are written, directly or indirectly, by people you do not control.

When the sources exceed the budget, something must give. Often one early line of code decides: keep the last N messages, or cut at a character limit. That line is a policy that decides which source loses when an untrusted source grows. If the answer is ever the system prompt or the safety instructions, you have a context stuffing vulnerability, whether or not anyone has found it.

Even when nothing is dropped, size changes behaviour. Long-context evaluations such as Liu et al. 2023, Lost in the Middle, found that models use information at the start and end of a long input more reliably than information in the middle. The exact curve depends on the model and has improved in newer systems, but the safe assumption is that an instruction separated from the question by tens of thousands of attacker-chosen tokens carries less weight than the same instruction placed next to it.

A context assembler with reserved regions: volume from one source cannot evict anotherUser messageuntrusted, any sizeRetrieved chunksuntrusted, top-kTool outputsuntrusted, any sizeConversation historymixed trustAssemblercount tokens per sourcequota, dedupe, truncateflag anomaliesSystem prompt + policyreserved, never truncatedHistory (summarised tail)quota 20%Retrieved evidencequota 30%, dedupedTool outputquota 20%, head+tail keptUser messagequota 25%, head+tail keptFinal instruction recapreserved, last positionTelemetrytokens per region, drops, repetition scoreModel callfixed max input tokens
A defensive assembler gives each source its own token quota, keeps trusted instructions in reserved regions at both ends, and records what it dropped. No single untrusted source can grow into another source's space.

Four families of volume attack

FamilyWhat the attacker controlsWhat goes wrong
DisplacementSize of a message, file or tool resultTruncation evicts the system prompt, tool rules or earlier evidence
DilutionAmount of filler around a requestTrusted instructions still present but far from the question and weakly followed
CrowdingMany near-duplicate documents or resultsTop-k retrieval or a result list fills with attacker content; honest sources drop out
AmplificationInput size per request or per turnCost and latency grow with tokens; quotas and other tenants suffer

Displacement is the most damaging because it is silent. A chat history stored as a list with the system message at index 0, trimmed with messages[-50:], loses that system message the moment the conversation passes fifty entries. A user who pastes one enormous block can force the same outcome in one turn if the trimming is by tokens from the oldest end.

Dilution needs no bug at all. Nothing is removed; the request simply surrounds the question with so much material that the policy text becomes a small, distant fraction of the input. The best-known variant is many-shot jailbreaking, described by Anil et al. at Anthropic in 2024: hundreds of fabricated dialogue examples inside one prompt, each showing an assistant complying, followed by the real request. They reported that attack success rose with the number of examples following a power law, the same shape seen in ordinary in-context learning, which is why it is hard to remove without hurting useful behaviour. The model-side picture is covered in multi-turn jailbreaks; here the point is that your assembler decides whether such a prompt can be that long in the first place.

Crowding moves the attack into data. If an attacker can add documents to a corpus you index, they can add fifty slight variations of one page. Similarity search returns the top k nearest chunks, and fifty near-copies of a page tuned to a likely query will take every slot. The model then reads one voice fifty times and the honest documents not at all. Oversized tool outputs crowd earlier tool results the same way. The retrieval side is covered in depth in prompt injection via RAG.

Amplification targets your budget rather than your model: prefill time and per-token billing grow with input size. The defence is the same as for any API, per-tenant token budgets and rate limits.

Worked example: the refund rule that fell off the front

Consider a customer-support agent with a 32,000-token input budget. Its system prompt is 1,200 tokens and includes the rule that refunds above 100 dollars require a human. Tool definitions take 2,000 tokens. The agent can call lookup_order and issue_refund. History is stored as a list of messages and trimmed like this:

def build_messages(history, user_msg, max_tokens=32_000):
    msgs = history + [{"role": "user", "content": user_msg}]
    # Drop oldest messages until it fits.
    while count_tokens(msgs) > max_tokens:
        msgs.pop(0)
    return msgs

The system prompt lives at history[0]. A user writes: here is my full order log, please check it and refund the 450-dollar duplicate charge and pastes 30,000 tokens of text. The loop pops the oldest message first, which is the system prompt, then the earliest turns, until the pasted log plus a few recent messages fit. The model now sees tool definitions, a giant log and a request for a refund. The refund rule is gone. Nothing in the logs shows that anything was dropped, because the code never recorded it. In testing everything worked, because nobody tested with a 30,000-token message.

The fix has three parts. First, the system prompt is not part of the trimmable history; it is a separate, reserved region. Second, the user message has its own quota: a 30,000-token paste is cut to its quota, keeping the head and tail with an explicit marker in the middle, so the model knows text was removed. Third, and most important, the refund rule is enforced in the issue_refund tool itself, which checks the amount and the caller's approval state in code. Prompt instructions are a convenience; a limit that matters must live where text cannot remove it.

A defensive context assembler

A defensive assembler is short. It computes each region in tokens with the same tokenizer the model uses, applies quotas, deduplicates retrieved chunks, and returns both the messages and a report of what it changed.

from dataclasses import dataclass, field

@dataclass
class Report:
    used: dict = field(default_factory=dict)
    truncated: list = field(default_factory=list)
    dropped_chunks: int = 0

QUOTAS = {"history": 0.20, "evidence": 0.30, "tools": 0.20, "user": 0.25}

def clip(text, limit, tok, name, report):
    ids = tok.encode(text)
    if len(ids) <= limit:
        return text
    head, tail = ids[: limit * 2 // 3], ids[-(limit // 3 - 16):]
    report.truncated.append((name, len(ids), limit))
    return tok.decode(head) + "\n[... content removed: input too long ...]\n" + tok.decode(tail)

def assemble(system, recap, history, chunks, tool_out, user, tok, budget):
    rep = Report()
    fixed = len(tok.encode(system)) + len(tok.encode(recap))
    free = budget - fixed
    if free <= 0:
        raise ValueError("system prompt and recap exceed the budget")
    lim = {k: int(free * v) for k, v in QUOTAS.items()}

    # near-duplicate removal (shingle Jaccard), then at most 2 chunks per source
    evidence = cap_per_source(dedupe(chunks, threshold=0.8), per_source=2)
    rep.dropped_chunks = len(chunks) - len(evidence)

    parts = {
        "history": clip(summarise_tail(history), lim["history"], tok, "history", rep),
        "evidence": clip("\n\n".join(evidence), lim["evidence"], tok, "evidence", rep),
        "tools": clip(tool_out, lim["tools"], tok, "tools", rep),
        "user": clip(user, lim["user"], tok, "user", rep),
    }
    rep.used = {k: len(tok.encode(v)) for k, v in parts.items()}
    msgs = [{"role": "system", "content": system},
            {"role": "user", "content": wrap_untrusted(parts)},
            {"role": "system", "content": recap}]
    return msgs, rep

Helpers such as dedupe and wrap_untrusted are left as stubs to fill in. Several details carry the weight. Quotas are fractions of what is left after the reserved regions, so a longer system prompt shrinks the untrusted space rather than the other way round. The recap is a short restatement of the rules that matter, placed last, which counters dilution by putting policy next to the question. Some provider APIs only accept a system message first; in that case put the recap at the end of the final user turn inside a clearly delimited block, as in spotlighting. Truncation keeps the head and tail, since logs and documents tend to put context at the start and the actionable part at the end, and it tells the model that content was removed instead of silently splicing. The per-source cap on evidence is the cheapest crowding defence there is: fifty copies of one page from one domain can occupy at most two slots.

Detection signals

Quotas limit damage; detection tells you someone is trying. Four signals are cheap to compute per request.

  • Length anomaly. A message in the top 0.1 percent of lengths for its route is worth a flag, not a block.
  • Repetition score. Compress the input with zlib and divide the compressed size by the raw size. Natural prose rarely compresses below about a third of its size; padding, repeated blocks and template-generated filler compress far more. Calibrate the threshold on your own traffic.
  • Embedded dialogue count. Count role markers such as lines starting with User:, Assistant:, Human: or chat-template special tokens inside a single user message. Real users occasionally paste one transcript; a message with two hundred alternating turns is the shape of a many-shot attempt.
  • Retrieval concentration. After retrieval, measure how many of the top k chunks come from one source, one author or one ingestion batch, and how similar they are to each other. A sudden rise for a given query family is a sign of corpus flooding.
import re, zlib

ROLE = re.compile(r"^\s*(user|assistant|human|ai|system)\s*:", re.I | re.M)

def stuffing_signals(text, p999_len):
    raw = text.encode("utf-8")
    ratio = len(zlib.compress(raw, 6)) / max(len(raw), 1)
    turns = len(ROLE.findall(text))
    return {
        "too_long": len(raw) > p999_len,
        "repetitive": len(raw) > 4000 and ratio < 0.15,
        "embedded_dialogue": turns >= 20,
    }

Route flagged requests to a stricter path: a smaller quota, a classifier pass, or a response that asks the user to attach the material as a file instead. Log the signals with the request ID so incidents can be reconstructed. The guardrail layers that can consume these signals are covered in RAG defence architecture.

How the defences fail

Defences against volume fail in predictable ways. Each of these has been seen in real systems.

  • Counting characters, not tokens. A character limit tuned on English lets through several times more tokens of text in some scripts, of emoji, or of unusual Unicode. Always count with the model's tokenizer, or with a conservative upper bound per byte.
  • Summaries as laundering. Replacing old turns with a model-written summary saves space, but an attacker who controls those turns now controls the summary, which you may label as trusted. Treat summaries of untrusted text as untrusted.
  • Quotas that starve real work. A legal team uploading a 200-page contract is not an attacker. If the quota cuts their document, answers get worse and they lose trust. Give long-document routes a different path, such as chunked map-reduce over the document, rather than raising the global quota.
  • Silent truncation elsewhere. A gateway or SDK may trim oversize inputs on its own. Check what your stack does on overflow and prefer an explicit error.
  • No regression test. The bug in the worked example survives because nobody tests with maximum-size inputs. Add a test that fills every untrusted region to twice its quota and asserts that the system prompt and recap are byte-identical in the final request.

Operations and trade-offs

In production, treat context like any other capacity. Export tokens per region, truncations, dropped chunks and stuffing signals per request, and alert on the user-region truncation rate per tenant. Enforce per-tenant input-token budgets at the gateway, before the model call, so amplification costs a rejected request rather than GPU time.

Evaluate robustness the same way you evaluate quality. Build a test set of policy-sensitive requests, then rerun it with filler inserted at several sizes and positions: before the request, after it, and spread through retrieved evidence. Measure whether the model still follows the policy, and whether your tools still refuse when it does not. The trade-offs look like this:

ControlStopsCosts
Reserved system region and recapDisplacement, much dilutionA few hundred tokens per request
Per-source quotas with head and tail clippingDisplacement, crowding by one sourceLong legitimate inputs get cut
Dedupe and per-domain caps on retrievalCorpus floodingFewer distinct chunks in rare topics
Stuffing signals and routingMany-shot and padding attemptsCalibration work, some false flags
Authorisation in tool codeConsequences of any of the aboveEngineering effort; the control that always holds
Per-tenant token budgetsAmplificationNeeds quota plumbing and customer messaging

What to do next

  1. Find the code that decides what goes into the context when inputs exceed the budget. Write down which source loses in each case.
  2. Move the system prompt and safety instructions out of trimmable history into a reserved region, and add a short recap of critical rules at the end of the input.
  3. Give each untrusted source a token quota, counted with the model tokenizer, and clip with a visible marker that keeps head and tail.
  4. Deduplicate retrieved chunks and cap chunks per source before they reach the prompt.
  5. Compute length, repetition and embedded-dialogue signals per request and log them with the request ID.
  6. Move every limit that matters, such as refund caps or data access, into tool code that checks it.
  7. Add a regression test that overfills every region and asserts the reserved regions arrive intact.
  8. Set per-tenant input-token budgets at the gateway and alert on truncation rates.
Key takeaway: Context stuffing works because applications allocate a fixed token budget with code nobody threat-modelled. Reserve space for trusted instructions, give every untrusted source a quota, deduplicate retrieval, log what you cut, and enforce the rules that matter in tool code, where no amount of text can remove them.