A model's context length is the number of tokens it can attend to in one forward pass: the system prompt, tool definitions, retrieved documents, the conversation so far and the answer it is writing all have to fit. Windows have grown from a few thousand tokens to hundreds of thousands and beyond, and it is tempting to treat that number as free space. It is not. Every token in the window costs compute before the first output token, memory for as long as the request lives, and bandwidth for every token generated. And a model that accepts a long input does not necessarily use all of it well.

This article is about using a context window, not extending one; training a model to a longer window is covered in the companion page on context extension. Here you will learn what counts against the window, how prompt length turns into prefill time and KV-cache memory with worked numbers for a 70B model, why advertised length and effective length differ, how to measure effective length on your own task, and how to budget tokens in an application. A contract question-answering service runs through as the worked example.

One window, many tenants

One request, one window: everything below shares the same token budgetSystem2KTools3KRetrieved documents40KHistory8KQuestion0.5KOutput reserve4KPrefill: all prompt tokens at oncecompute-bound, time to first tokenDecode: one token at a timereads the whole KV cache per stepKV cache: keys and values for every token, every layerabout 320 KiB per token for a 70B GQA model in BF16Recall is usually strongest near the start and end of the prompt and weaker in the middle.
System prompt, tools, documents, history, question and output all share one token budget; prefill cost and cache memory grow with everything in it.

What counts against the window

The window is measured in tokens of the model's own tokenizer, not words or characters. English prose often runs between 1.2 and 1.5 tokens per word with current tokenizers; code, tables, JSON and many other languages run higher. Never estimate from character counts in production; count with the tokenizer the model uses, because the same text can differ by tens of percent between model families.

Everything the model receives counts: the system prompt, every tool or function schema (these are easy to forget and can run to thousands of tokens), retrieved passages, prior turns including the model's own earlier answers, images or audio after encoding, and the chat template's special tokens. In most self-hosted serving engines the generated tokens share the same window, so a 32,768-token limit with a 31,000-token prompt leaves under 1,800 tokens to answer; vLLM's max_model_len, for instance, bounds prompt plus output. Hosted APIs usually also cap output separately. Either way, reserve the output budget before you fill the prompt.

Prefill: why long prompts are slow to start

Prefill is the forward pass over the whole prompt, done in parallel, before any output token appears. It has two parts. The matrix multiplications with the weights cost about 2 FLOPs per parameter per token, so they grow linearly with prompt length. Attention compares every token with every earlier token: per layer, each query-key pair costs about 2d FLOPs for the scores and 2d for the weighted sum over values, where d is the model width, and a causal prompt of n tokens has about n squared over 2 pairs. Attention therefore costs about 2 times layers times d times n squared.

Take the published Llama 3 70B shape: 80 layers, width 8192, about 70 billion parameters. The two terms are equal when n equals parameters divided by (layers times width), which is about 107,000 tokens. Below that the weights dominate; above it attention does. Worked numbers:

Prompt tokensWeight FLOPsAttention FLOPsAttention shareTotal
8,1921.15e158.8e137%1.2e15
32,7684.6e151.4e1523%6.0e15
131,0721.8e162.3e1655%4.1e16

Going from 8K to 128K multiplies the prompt by 16 but the prefill work by about 33. At a sustained 1 PFLOP/s, an assumed figure for a multi-GPU serving group rather than any specific product, the 128K prefill takes about 41 seconds and the 8K prefill about 1.2 seconds. Kernels such as FlashAttention cut the memory traffic of attention, not its arithmetic, so the quadratic term remains. The practical consequence is that time to first token for long prompts is dominated by prefill, and long prompts from one user delay short prompts from everyone else unless the engine chunks prefill.

Decode and the KV cache: why long prompts stay expensive

During decoding the model generates one token per step, and each step reads the keys and values of every previous token at every layer. The KV cache size per token is 2 times layers times KV heads times head dimension times bytes per value. The 70B model uses grouped-query attention with 8 KV heads of dimension 128, so in BF16 that is 2 x 80 x 8 x 128 x 2 bytes, 327,680 bytes or 320 KiB per token. A 131,072-token context holds 40 GiB of cache for one sequence.

Memory caps concurrency: a node with 640 GB of accelerator memory, holding about 141 GB of BF16 weights, has room for roughly a dozen 128K sequences and nothing else, against hundreds of 4K sequences. Bandwidth caps speed: each decode step reads the weights once for the whole batch but the cache once per sequence, so four 128K sequences read 160 GiB of cache per step, more than the weights. Long contexts therefore make every generated token slower, not just the first. The remedies are paged cache allocation, cache quantization, prefix caching for shared document prefixes, and simply sending fewer tokens.

Advertised versus effective length

The advertised context length is the longest input the model was trained or configured to accept. Effective context length is the longest input at which it still performs your task at acceptable quality. They differ, often a lot.

Two research results frame this. Lost in the Middle (Liu et al., 2023) moved the passage containing an answer through a long multi-document prompt and found accuracy highest when the passage was near the beginning or the end and lowest in the middle, a U-shaped curve. RULER (Hsieh et al., 2024) went beyond the single needle-in-a-haystack test with tasks such as multiple needles, variable tracking across the context and aggregation; models that scored near-perfectly on the simple needle test dropped sharply as length grew, and only about half of the models evaluated held up at 32K despite claiming that length or more. Exact scores change with every model release, so do not carry numbers from either paper into a design decision; carry the method.

The mechanism is intuitive. A single retrieval is easy because one distinctive passage matches the question. Real tasks need the model to find several relevant spans among many similar ones, combine them, and ignore distractors, and that gets harder as the haystack grows. Long-context training also uses fewer very long examples than short ones, so the model has less practice at the far end of its window.

Measuring effective length on your task

Measure effective length on your own documents and questions, not on a public benchmark. Take questions whose answers you know, place the answer-bearing passage at several depths inside realistic filler from your own corpus, and sweep total length.

import itertools, random

def build_prompt(question, evidence, filler_docs, total_tokens, depth, count_tokens):
    """Place evidence at relative depth (0.0 = start, 1.0 = end) inside filler of total_tokens."""
    docs, used = [], count_tokens(evidence)
    for d in filler_docs:                       # realistic distractors from the same corpus
        if used + count_tokens(d) > total_tokens:
            break
        docs.append(d)
        used += count_tokens(d)
    docs.insert(round(depth * len(docs)), evidence)
    return "\n\n".join(docs) + "\n\nQuestion: " + question

def sweep(cases, model, count_tokens, lengths=(4_000, 16_000, 32_000, 64_000, 128_000),
          depths=(0.0, 0.25, 0.5, 0.75, 1.0)):
    scores = {}
    for L, depth in itertools.product(lengths, depths):
        ok = 0
        for c in cases:
            filler = random.sample(c.filler_pool, k=len(c.filler_pool))
            answer = model(build_prompt(c.question, c.evidence, filler, L, depth, count_tokens))
            ok += c.grade(answer)
        scores[(L, depth)] = ok / len(cases)
    return scores   # the effective length is the largest L where every depth meets your bar

Use at least 50 questions per cell, include questions that need two or more passages, and grade with the same checker you use in production. The output is a grid of length against depth. Your effective length is the largest length at which every depth clears your quality bar, and it is usually well below the advertised figure.

Budgeting tokens in the application

Treat the window as a budget with priorities. Reserve output first, then fixed parts (system prompt, tools), then the current question, then fill the rest with retrieved context and history in priority order, trimming the lowest priority first. Put the most important material near the end, close to the question, or at the very start, not in the middle.

def assemble(parts, limit, output_reserve, count_tokens):
    """parts: list of (priority, name, text, droppable); lower priority number = more important."""
    budget = limit - output_reserve
    fixed = [p for p in parts if not p[3]]
    budget -= sum(count_tokens(p[2]) for p in fixed)
    if budget < 0:
        raise ValueError("fixed parts alone exceed the window")
    kept = list(fixed)
    for prio, name, text, _ in sorted(p for p in parts if p[3]):
        n = count_tokens(text)
        if n <= budget:
            kept.append((prio, name, text, True))
            budget -= n
    return kept      # order the kept parts for the prompt separately: key evidence last

For conversation history, keep the last few turns verbatim and replace older turns with a running summary. For documents, retrieve the passages that matter instead of sending whole files. Log the token count of every part on every request; when quality drops, that log tells you whether the window was crowded.

Worked example: questions over long contracts

A legal team wants answers about supplier contracts of 150 to 300 pages, roughly 90,000 to 180,000 tokens each, from a self-hosted 70B model. The first design sends the whole contract with every question. Prefill for a 150,000-token prompt is about 5e16 FLOPs and the cache is about 46 GiB per request. Suppose a sweep on 60 known questions shows accuracy of 92 percent at 16K, 85 percent at 64K and 71 percent at full length, with the worst results for clauses in the middle third of the document and for questions that combine a definition with a later clause.

The second design splits each contract by clause, retrieves the 12 most relevant clauses plus the definitions section (about 14,000 tokens), places definitions first and the retrieved clauses last before the question, and reserves 2,000 tokens for the answer. Suppose accuracy on the same 60 questions is 90 percent; prefill work falls by more than an order of magnitude, and the node serves many requests at once instead of a handful. For the remaining hard questions, those that need a whole-document view such as listing every termination right, the team uses a map-reduce pass over sections rather than one huge prompt. The accuracy figures are hypothetical; run the sweep on your own documents.

Failure modes

SymptomLikely causeWhat to do
Answers cut off mid-sentencePrompt left too little of the shared window for outputReserve output tokens before assembling the prompt
Correct passage was in the prompt but ignoredPlaced in the middle of a long promptMove key evidence near the question; send less
Time to first token jumps for everyoneA few very long prompts monopolise prefillChunked prefill, length-based routing, length limits per tenant
Out-of-memory or low concurrencyKV cache sized for the maximum lengthLower the engine's maximum length to what you need; paged and quantized cache
Quality drifts down over a long chatHistory crowds out instructions and evidenceSummarise older turns; re-send key instructions
Token counts disagree with the APIEstimating from characters or the wrong tokenizerCount with the model's tokenizer and template

Trade-offs: long context, retrieval or summarisation

Long context, retrieval and summarisation are not rivals. Long context is the right tool when the task needs a whole-document view and the document fits within your measured effective length, or when the same long prefix is reused across many questions and prefix caching makes repeated prefill cheap. Retrieval is right when only a small part of a large corpus matters per question. Summarisation and map-reduce are right when the answer depends on everything but no single pass can hold it reliably. Most production systems combine them, and the deciding inputs are your own effective-length grid and your cost per request.

Further reading on this site: extending a model's context window, tokenization, FlashAttention, PagedAttention and KV cache quantization.

What to do next

  1. Count tokens per part (system, tools, retrieval, history, question) on a day of real traffic using the model's tokenizer.
  2. Reserve an explicit output budget in code and fail loudly when fixed parts exceed the window.
  3. Run the depth-by-length sweep on 50 or more of your own questions and record the effective length.
  4. Set the serving engine's maximum length to what you need, not what the model allows.
  5. Put key evidence near the question, and summarise or retrieve instead of sending whole histories and files.
  6. Track time to first token and cache memory per request, broken down by prompt length.
  7. Re-run the sweep whenever you change model, prompt template or retrieval settings.
Key takeaway: Context length is a shared budget for system prompt, tools, documents, history, question and output. Prefill work grows with the square of prompt length once attention dominates, around 107,000 tokens for a 70B model, and the KV cache costs about 320 KiB per token in BF16, limiting concurrency and slowing every decode step. Models use the middle of long prompts less reliably than the ends, so measure effective length on your own task, reserve output tokens, put key evidence near the question, and prefer retrieval or summarisation when a smaller prompt does the job.