Context windows have grown from a few thousand tokens to hundreds of thousands, and some models accept a million or more. It is tempting to treat that as a solved problem: put every document in the prompt and ask the question. In practice, accuracy on long inputs depends heavily on how the input is laid out, where the question sits, and whether the model is asked to locate evidence before it answers. Two prompts with identical content in a different order can produce noticeably different results.
This article is about prompting once the material is already chosen: the layout of a long prompt, how to tag and index documents, the quote-first pattern, how to handle documents that disagree, and how to measure, on your own data, how far into the context your model reliably reads. Choosing and budgeting what goes into the window is a separate problem covered in context packing; this page assumes you have already decided, and are now deciding how to present it.
Why a bigger window is not a bigger memory
A context window is a limit on how many tokens the model can attend over, not a promise that it uses them all equally well. Attention weighting is learned, training data rarely puts the key material deep in very long inputs, and position encodings behave differently at lengths that were rare in training. The result is uneven recall across positions.
The best-known measurement is the 2023 paper Lost in the Middle by Liu and colleagues, which found that on multi-document question answering, accuracy was highest when the relevant document was near the start or the end of the context and dropped when it was in the middle. Newer models have reduced this effect, but not eliminated it, and the size of the effect varies by model and task. The RULER benchmark from NVIDIA (2024) made a related point: simple needle-in-a-haystack retrieval often looks near-perfect, while harder tasks such as multi-hop tracing and aggregation degrade well before the advertised window is full. Its authors call the length at which performance stays acceptable the effective context length.
Two practical consequences follow. First, a literal needle test, where the model finds one unusual sentence, tells you little about whether it can compare clauses across forty contracts. Second, the advertised limit is an upper bound on what is accepted, not the length at which you should plan to operate. You have to measure your own effective length, and the last part of this article shows how.
The layout that works
Most providers' guidance and most practitioners' measurements converge on the same layout. Put the long material first and the question last. Anthropic's long-context guidance, for example, recommends placing long documents near the top of the prompt, above the query and instructions, and reports that putting queries at the end improved response quality by up to 30% in its tests, especially with complex multi-document inputs. The intuition is simple: the model generates its answer immediately after reading the question, so the question and the instructions should be the freshest thing in its context, not something it read a hundred thousand tokens ago.
A short framing at the very top still helps: one or two sentences saying what the documents are and what task follows let the model read with a purpose. Many teams use a sandwich, a brief version of the question at the top and the precise version, with full instructions and output format, at the end.
The same layout is also the one that caches well. Provider-side prompt caching reuses work for an identical prefix, so documents that do not change between requests belong before anything that does. If you put a per-user question at the top, every request has a different prefix and nothing can be reused. See prompt caching for how cache breakpoints and lifetimes work.
Tag every document and give the model an index
Pasted documents separated by blank lines give the model no reliable boundary and no way to refer to one precisely. Wrap each document in its own block with an id and metadata; XML-style tags are unambiguous and familiar to models (see delimiters and XML tags).
Put a one-line-per-document index before the documents. It gives the model a map, for example that three of 37 documents are amendments dated after the main agreement, before it reads any of them. Escape document text so a document containing your closing tag cannot break the structure.
from dataclasses import dataclass
from xml.sax.saxutils import escape
@dataclass
class Doc:
doc_id: str # stable, short: "D07"
title: str
source: str # path, URL or system of record
date: str # ISO date, or "" if unknown
text: str
def build_long_prompt(docs: list[Doc], framing: str, question: str, rules: list[str]) -> str:
index = "\n".join(f"{d.doc_id} | {d.date or 'undated'} | {d.title} | {d.source}" for d in docs)
blocks = []
for d in docs:
blocks.append(
f'<document id="{d.doc_id}">\n'
f"<source>{escape(d.source)}</source>\n<date>{d.date}</date>\n"
f"<document_content>\n{escape(d.text)}\n</document_content>\n</document>"
)
rule_text = "\n".join(f"- {r}" for r in rules)
return (
f"{framing}\n\n<index>\n{index}\n</index>\n\n<documents>\n"
+ "\n".join(blocks)
+ f"\n</documents>\n\n<question>\n{question}\n</question>\n\n<instructions>\n{rule_text}\n</instructions>"
)
Quote first, then answer
The single most effective instruction for long inputs is to make the model find its evidence before it reasons. Ask it to copy the exact passages that bear on the question, with the document id for each, into one block, and only then write the answer in a second block that may rely only on those quotes. This turns one hard task, answering from a huge context, into two easier ones: locating relevant text, then reasoning over a short excerpt that is now at the end of the context where recall is strongest.
Quote-first also makes answers checkable: because quotes should be verbatim, code can confirm each one appears in the cited document after normalising whitespace. A mismatch is a strong signal of fabrication, so reject or retry automatically. This deterministic guard complements grounding and citations.
import re
QUOTE_RULES = [
"First, inside <quotes>, copy every passage relevant to the question verbatim, each as "
"<quote doc=\"ID\">exact text</quote>. Copy, do not paraphrase.",
"If no document contains the answer, write <quotes></quotes> and say so in <answer>.",
"Then, inside <answer>, answer using only the quotes. Cite ids like [D07].",
"If documents disagree, list each position with its id and date; do not pick one silently.",
]
def norm(s: str) -> str:
return " ".join(s.split()).lower()
def verify_quotes(response: str, docs_by_id: dict[str, str]) -> list[str]:
problems = []
for doc_id, quote in re.findall(r'<quote doc="([^"]+)">(.*?)</quote>', response, re.S):
body = docs_by_id.get(doc_id)
if body is None:
problems.append(f"unknown document {doc_id}")
elif norm(quote) not in norm(body):
problems.append(f"quote not found in {doc_id}: {quote[:60]!r}")
return problemsTwo details matter. Remember to unescape quotes before comparing if you escaped the document text. And allow an explicit empty result: if the prompt does not tell the model what to do when the answer is absent, it will tend to produce something plausible from the nearest related passage.
When documents disagree
Real sources conflict: an amendment overrides a clause, a newer policy replaces an older one. Left alone, a model often blends conflicting statements into one confident answer or picks whichever it read last. Give it a rule: recency (later dated documents win), authority (the signed contract outranks the summary email), or report-all (list every position with ids). Put dates in metadata so recency is computable, ask for a conflict list in the output, and include deliberate conflicts in your evaluation set.
Worked example: forty supplier contracts
In this illustrative scenario, a procurement team wants to know which supplier contracts allow termination for convenience with less than 90 days notice. There are 40 contracts and amendments totalling about 150,000 tokens, within the window of the model they use.
The first attempt pasted all forty files in folder order with the question at the top and returned eleven contracts. Review found two false positives, where a 60-day cure period was read as a notice period, and four misses, three of them amendments sitting in the middle of the prompt.
The second attempt changed only the presentation. Each file became a tagged document with id, counterparty, type and effective date; an index marked amendments and their parent agreements. The question moved to the end with a precise definition: notice for termination for convenience, not cure periods and not termination for cause. The output contract required quotes first, then a table of contract id, notice days and supporting quote, with a rule that a later amendment overrides its parent, and the verifier checked every quote.
This time all fourteen qualifying contracts were found, and one fabricated quote was caught by the verifier and corrected on retry. The gain came from layout, definitions, a conflict rule and a cheap check. Splitting the job into two calls, extracting every termination clause as rows and then answering from the rows, is often more accurate again and easier to audit, because the task is really extraction followed by filtering.
Measure your effective context length
Do not take a benchmark's word for your task. Build a positional sensitivity test from your own data: insert a known fact at several depths of a realistic corpus, run the same question at several total lengths, and record accuracy per cell. Use your real distractors, because difficulty depends on how similar irrelevant text is to relevant text.
import itertools, random, statistics
def build_haystack(corpus: list[str], target_tokens: int, count_tokens) -> list[str]:
out, total = [], 0
for chunk in corpus:
n = count_tokens(chunk)
if total + n > target_tokens:
break
out.append(chunk); total += n
return out
def run_grid(corpus, fact, question, grade, call_model, count_tokens,
lengths=(8_000, 32_000, 64_000, 128_000), depths=(0.0, 0.25, 0.5, 0.75, 1.0), seeds=3):
results = {}
for length, depth in itertools.product(lengths, depths):
scores = []
for seed in range(seeds):
rng = random.Random(seed)
chunks = build_haystack(rng.sample(corpus, len(corpus)), length, count_tokens)
chunks.insert(int(depth * len(chunks)), fact)
prompt = "\n\n".join(chunks) + "\n\n" + question
scores.append(grade(call_model(prompt)))
results[(length, depth)] = statistics.mean(scores)
return resultsThe length at which middle-depth accuracy drops below your quality bar is your effective length for this task; operate below it. If accuracy is flat across depths but low overall, the problem is task difficulty, not position. For multi-hop questions, insert two or three linked facts at different depths, where degradation appears first.
Cost, latency and compression
Long prompts are paid for on every call, and time to first token grows with prompt length because the whole prompt is processed before the first output token. Caching the stable prefix cuts both for repeat requests over the same corpus.
When questions are narrow, ask whether the model needs the whole corpus. Retrieval narrows the input and compression removes low-value tokens; prompt compression covers the accuracy costs. Long context wins when the question needs global reading, such as comparing all contracts or finding contradictions anywhere. Retrieval wins when the answer is in a few places and the corpus dwarfs the window. Many systems retrieve generously, then read the retrieved set in full.
Failure modes
| Symptom | Likely cause | Fix |
|---|---|---|
| Misses facts in the middle of the input | Positional degradation | Index, quote-first, move key material, or split into extraction plus answer |
| Answers from the wrong document | No clear boundaries or ids | Tag every document with id and metadata |
| Confident answer when nothing matches | No instruction for the absent case | Explicit empty result in the output contract |
| Blends conflicting sources | No conflict rule | State recency or authority rule, ask for a conflict list |
| Ignores formatting or constraints | Instructions only at the top | Restate full instructions after the documents |
| Fabricated quotes | Paraphrase pressure, weak grounding | Verify quotes by substring, reject and retry |
| Follows instructions found inside a document | Prompt injection in source text | Mark documents as data, never as instructions, and filter untrusted input |
| Slow and expensive | Whole corpus sent every call | Cache the stable prefix, retrieve or compress |
The injection row deserves emphasis: long prompts often carry text you did not write, and instruction-like text in it may be followed. Say that documents are reference material only, and treat that as a mitigation, not a guarantee.
What to do next
- Rewrite one long-context prompt so that documents come first, wrapped in tagged blocks with ids, sources and dates, and the question and full instructions come last.
- Add a one-line-per-document index above the documents.
- Add a quote-first output contract with an explicit empty result, and a verifier that checks every quote is a substring of the cited document.
- Write a conflict rule (recency, authority or report-all) and put dates in metadata so it can be applied.
- Build the positional sensitivity grid from your own corpus, find your effective length, and note it next to the prompt in your registry.
- Move the stable prefix ahead of anything per-request and turn on prompt caching; measure time to first token before and after.
- For narrow questions over large corpora, compare long context against retrieval plus a shorter prompt on the same evaluation set before choosing.