Context windows have grown from a few thousand tokens to hundreds of thousands, and some models accept a million or more. It is tempting to treat that as a solved problem: put every document in the prompt and ask the question. In practice, accuracy on long inputs depends heavily on how the input is laid out, where the question sits, and whether the model is asked to locate evidence before it answers. Two prompts with identical content in a different order can produce noticeably different results.

This article is about prompting once the material is already chosen: the layout of a long prompt, how to tag and index documents, the quote-first pattern, how to handle documents that disagree, and how to measure, on your own data, how far into the context your model reliably reads. Choosing and budgeting what goes into the window is a separate problem covered in context packing; this page assumes you have already decided, and are now deciding how to present it.

Advertisement

Why a bigger window is not a bigger memory

A context window is a limit on how many tokens the model can attend over, not a promise that it uses them all equally well. Attention weighting is learned, training data rarely puts the key material deep in very long inputs, and position encodings behave differently at lengths that were rare in training. The result is uneven recall across positions.

The best-known measurement is the 2023 paper Lost in the Middle by Liu and colleagues, which found that on multi-document question answering, accuracy was highest when the relevant document was near the start or the end of the context and dropped when it was in the middle. Newer models have reduced this effect, but not eliminated it, and the size of the effect varies by model and task. The RULER benchmark from NVIDIA (2024) made a related point: simple needle-in-a-haystack retrieval often looks near-perfect, while harder tasks such as multi-hop tracing and aggregation degrade well before the advertised window is full. Its authors call the length at which performance stays acceptable the effective context length.

Two practical consequences follow. First, a literal needle test, where the model finds one unusual sentence, tells you little about whether it can compare clauses across forty contracts. Second, the advertised limit is an upper bound on what is accepted, not the length at which you should plan to operate. You have to measure your own effective length, and the last part of this article shows how.

The layout that works

Most providers' guidance and most practitioners' measurements converge on the same layout. Put the long material first and the question last. Anthropic's long-context guidance, for example, recommends placing long documents near the top of the prompt, above the query and instructions, and reports that putting queries at the end improved response quality by up to 30% in its tests, especially with complex multi-document inputs. The intuition is simple: the model generates its answer immediately after reading the question, so the question and the instructions should be the freshest thing in its context, not something it read a hundred thousand tokens ago.

A short framing at the very top still helps: one or two sentences saying what the documents are and what task follows let the model read with a purpose. Many teams use a sandwich, a brief version of the question at the top and the precise version, with full instructions and output format, at the end.

Layout of a long-context prompt: stable material first, the question last, the output contract at the very end1. Role and short task framinga few sentences: who you are, what the documents are for2. Document indexid, title, source, date for every document, one line each3. Documentseach in its own tagged block with id, source and content4. Question and full instructionsrestated task, constraints, how to handle conflicts5. Output contractquotes first, then answer, cite ids, say when absentCacheable prefixsame bytes on every callVaries per requestquestion, user turnMost of the tokens are in block 3; most of the steering is in blocks 4 and 5, which the model reads lastKeeping blocks 1-3 byte-identical across calls is what lets a provider-side prompt cache reuse them
Blocks 1 to 3 are stable and large; blocks 4 and 5 are small and change per request. The model reads the steering last, and a prompt cache can reuse the stable prefix.

The same layout is also the one that caches well. Provider-side prompt caching reuses work for an identical prefix, so documents that do not change between requests belong before anything that does. If you put a per-user question at the top, every request has a different prefix and nothing can be reused. See prompt caching for how cache breakpoints and lifetimes work.

Advertisement

Tag every document and give the model an index

Pasted documents separated by blank lines give the model no reliable boundary and no way to refer to one precisely. Wrap each document in its own block with an id and metadata; XML-style tags are unambiguous and familiar to models (see delimiters and XML tags).

Put a one-line-per-document index before the documents. It gives the model a map, for example that three of 37 documents are amendments dated after the main agreement, before it reads any of them. Escape document text so a document containing your closing tag cannot break the structure.

from dataclasses import dataclass
from xml.sax.saxutils import escape

@dataclass
class Doc:
    doc_id: str      # stable, short: "D07"
    title: str
    source: str      # path, URL or system of record
    date: str        # ISO date, or "" if unknown
    text: str

def build_long_prompt(docs: list[Doc], framing: str, question: str, rules: list[str]) -> str:
    index = "\n".join(f"{d.doc_id} | {d.date or 'undated'} | {d.title} | {d.source}" for d in docs)
    blocks = []
    for d in docs:
        blocks.append(
            f'<document id="{d.doc_id}">\n'
            f"<source>{escape(d.source)}</source>\n<date>{d.date}</date>\n"
            f"<document_content>\n{escape(d.text)}\n</document_content>\n</document>"
        )
    rule_text = "\n".join(f"- {r}" for r in rules)
    return (
        f"{framing}\n\n<index>\n{index}\n</index>\n\n<documents>\n"
        + "\n".join(blocks)
        + f"\n</documents>\n\n<question>\n{question}\n</question>\n\n<instructions>\n{rule_text}\n</instructions>"
    )

Quote first, then answer

The single most effective instruction for long inputs is to make the model find its evidence before it reasons. Ask it to copy the exact passages that bear on the question, with the document id for each, into one block, and only then write the answer in a second block that may rely only on those quotes. This turns one hard task, answering from a huge context, into two easier ones: locating relevant text, then reasoning over a short excerpt that is now at the end of the context where recall is strongest.

Quote-first also makes answers checkable: because quotes should be verbatim, code can confirm each one appears in the cited document after normalising whitespace. A mismatch is a strong signal of fabrication, so reject or retry automatically. This deterministic guard complements grounding and citations.

import re

QUOTE_RULES = [
    "First, inside <quotes>, copy every passage relevant to the question verbatim, each as "
    "<quote doc=\"ID\">exact text</quote>. Copy, do not paraphrase.",
    "If no document contains the answer, write <quotes></quotes> and say so in <answer>.",
    "Then, inside <answer>, answer using only the quotes. Cite ids like [D07].",
    "If documents disagree, list each position with its id and date; do not pick one silently.",
]

def norm(s: str) -> str:
    return " ".join(s.split()).lower()

def verify_quotes(response: str, docs_by_id: dict[str, str]) -> list[str]:
    problems = []
    for doc_id, quote in re.findall(r'<quote doc="([^"]+)">(.*?)</quote>', response, re.S):
        body = docs_by_id.get(doc_id)
        if body is None:
            problems.append(f"unknown document {doc_id}")
        elif norm(quote) not in norm(body):
            problems.append(f"quote not found in {doc_id}: {quote[:60]!r}")
    return problems

Two details matter. Remember to unescape quotes before comparing if you escaped the document text. And allow an explicit empty result: if the prompt does not tell the model what to do when the answer is absent, it will tend to produce something plausible from the nearest related passage.

When documents disagree

Real sources conflict: an amendment overrides a clause, a newer policy replaces an older one. Left alone, a model often blends conflicting statements into one confident answer or picks whichever it read last. Give it a rule: recency (later dated documents win), authority (the signed contract outranks the summary email), or report-all (list every position with ids). Put dates in metadata so recency is computable, ask for a conflict list in the output, and include deliberate conflicts in your evaluation set.

Worked example: forty supplier contracts

In this illustrative scenario, a procurement team wants to know which supplier contracts allow termination for convenience with less than 90 days notice. There are 40 contracts and amendments totalling about 150,000 tokens, within the window of the model they use.

The first attempt pasted all forty files in folder order with the question at the top and returned eleven contracts. Review found two false positives, where a 60-day cure period was read as a notice period, and four misses, three of them amendments sitting in the middle of the prompt.

The second attempt changed only the presentation. Each file became a tagged document with id, counterparty, type and effective date; an index marked amendments and their parent agreements. The question moved to the end with a precise definition: notice for termination for convenience, not cure periods and not termination for cause. The output contract required quotes first, then a table of contract id, notice days and supporting quote, with a rule that a later amendment overrides its parent, and the verifier checked every quote.

This time all fourteen qualifying contracts were found, and one fabricated quote was caught by the verifier and corrected on retry. The gain came from layout, definitions, a conflict rule and a cheap check. Splitting the job into two calls, extracting every termination clause as rows and then answering from the rows, is often more accurate again and easier to audit, because the task is really extraction followed by filtering.

Measure your effective context length

Do not take a benchmark's word for your task. Build a positional sensitivity test from your own data: insert a known fact at several depths of a realistic corpus, run the same question at several total lengths, and record accuracy per cell. Use your real distractors, because difficulty depends on how similar irrelevant text is to relevant text.

Positional sensitivity test: same fact, same distractors, moved through the contextReal corpusyour own documentsInsert factat depth 0%..100%Ask questionfixed prompt templateGradeexact or judgedGrid of lengths x depthsaccuracy per cell, several seedsDecisionmax safe length, where to put key materialMeasure on your documents and your model: published benchmarks rarely match your distractors
The test moves one known fact through a realistic corpus and records accuracy for each combination of total length and depth.
import itertools, random, statistics

def build_haystack(corpus: list[str], target_tokens: int, count_tokens) -> list[str]:
    out, total = [], 0
    for chunk in corpus:
        n = count_tokens(chunk)
        if total + n > target_tokens:
            break
        out.append(chunk); total += n
    return out

def run_grid(corpus, fact, question, grade, call_model, count_tokens,
             lengths=(8_000, 32_000, 64_000, 128_000), depths=(0.0, 0.25, 0.5, 0.75, 1.0), seeds=3):
    results = {}
    for length, depth in itertools.product(lengths, depths):
        scores = []
        for seed in range(seeds):
            rng = random.Random(seed)
            chunks = build_haystack(rng.sample(corpus, len(corpus)), length, count_tokens)
            chunks.insert(int(depth * len(chunks)), fact)
            prompt = "\n\n".join(chunks) + "\n\n" + question
            scores.append(grade(call_model(prompt)))
        results[(length, depth)] = statistics.mean(scores)
    return results

The length at which middle-depth accuracy drops below your quality bar is your effective length for this task; operate below it. If accuracy is flat across depths but low overall, the problem is task difficulty, not position. For multi-hop questions, insert two or three linked facts at different depths, where degradation appears first.

Cost, latency and compression

Long prompts are paid for on every call, and time to first token grows with prompt length because the whole prompt is processed before the first output token. Caching the stable prefix cuts both for repeat requests over the same corpus.

When questions are narrow, ask whether the model needs the whole corpus. Retrieval narrows the input and compression removes low-value tokens; prompt compression covers the accuracy costs. Long context wins when the question needs global reading, such as comparing all contracts or finding contradictions anywhere. Retrieval wins when the answer is in a few places and the corpus dwarfs the window. Many systems retrieve generously, then read the retrieved set in full.

Failure modes

SymptomLikely causeFix
Misses facts in the middle of the inputPositional degradationIndex, quote-first, move key material, or split into extraction plus answer
Answers from the wrong documentNo clear boundaries or idsTag every document with id and metadata
Confident answer when nothing matchesNo instruction for the absent caseExplicit empty result in the output contract
Blends conflicting sourcesNo conflict ruleState recency or authority rule, ask for a conflict list
Ignores formatting or constraintsInstructions only at the topRestate full instructions after the documents
Fabricated quotesParaphrase pressure, weak groundingVerify quotes by substring, reject and retry
Follows instructions found inside a documentPrompt injection in source textMark documents as data, never as instructions, and filter untrusted input
Slow and expensiveWhole corpus sent every callCache the stable prefix, retrieve or compress

The injection row deserves emphasis: long prompts often carry text you did not write, and instruction-like text in it may be followed. Say that documents are reference material only, and treat that as a mitigation, not a guarantee.

What to do next

  1. Rewrite one long-context prompt so that documents come first, wrapped in tagged blocks with ids, sources and dates, and the question and full instructions come last.
  2. Add a one-line-per-document index above the documents.
  3. Add a quote-first output contract with an explicit empty result, and a verifier that checks every quote is a substring of the cited document.
  4. Write a conflict rule (recency, authority or report-all) and put dates in metadata so it can be applied.
  5. Build the positional sensitivity grid from your own corpus, find your effective length, and note it next to the prompt in your registry.
  6. Move the stable prefix ahead of anything per-request and turn on prompt caching; measure time to first token before and after.
  7. For narrow questions over large corpora, compare long context against retrieval plus a shorter prompt on the same evaluation set before choosing.
Key takeaway: A long context window lets you send more, not makes the model read everything equally well. Put stable documents first in tagged, indexed blocks, put the question and full instructions last, make the model quote its evidence before answering and verify those quotes in code, give it an explicit rule for conflicting sources, and measure your own effective context length with a positional test on your real data before you rely on the advertised limit.