A model card that says "128K context" tells you the length the model will accept without an error. It does not tell you the length at which the model still uses what you put in. Those two numbers are often far apart, and the only way to find the second one for your model, your serving stack and your kind of task is to measure it.
This article explains how to build that measurement: what skills a long-context evaluation has to separate, how public suites approach it, how to generate test cases that cannot be passed by accident, how to score and read them, what a sweep costs, and the many ways an evaluation harness quietly produces the wrong answer. It includes a small, working harness you can point at any model endpoint.
Five abilities, not one
Long-context ability is not one skill. A model can be excellent at finding one unusual sentence in a long document and poor at everything else. A useful evaluation separates at least five abilities:
- Retrieval: find a fact stated once and return it. This is the needle-in-a-haystack test, and it is the easiest.
- Retrieval under interference: find the right fact when several similar ones are present, such as many keys with values, of which only one is asked for.
- Tracking: follow a chain of references spread across the context, such as variable assignments where each refers to the previous one.
- Aggregation: answer a question whose answer depends on the whole context, such as counting or finding the most frequent items. No single passage holds the answer.
- Reasoning over spread evidence: answer a question that needs two or more facts located far apart, possibly phrased with no word overlap between question and evidence.
Each ability degrades at a different length. The useful output of an evaluation is therefore not one score but a profile: for each task family, accuracy as a function of length and of where the evidence sits. From that profile you derive an effective context length, the longest length at which the model still meets a quality bar you choose, and that number, not the advertised window, is what you design your retrieval chunking, prompt layout and product limits around.
What public suites taught us
Several public efforts shaped how this is measured, and each exposes a different weakness.
The needle-in-a-haystack test, popularised by Greg Kamradt in 2023, inserts one out-of-place sentence into long filler text at a chosen depth and asks for it back, sweeping length and depth into a heat map. Modern models mostly saturate it, so a green map says little about harder tasks.
Lost in the Middle (Liu et al., 2023) showed with multi-document question answering and key-value retrieval that accuracy is often highest when the relevant information is at the start or end of the context and lowest in the middle: the U-shaped curve. Depth is therefore a required axis, not an optional one.
RULER (NVIDIA, 2024) generalised the needle test into 13 synthetic tasks in four categories: retrieval variants with multiple keys and values, multi-hop tracing, aggregation, and question answering, at lengths from 4K to 128K. It defined effective length as the longest length at which a model's average score stays above a fixed bar, Llama-2-7B's score at 4K (85.6%). Of the models evaluated, all claimed 32K or more, but only about half stayed above that bar at 32K.
NoLiMa (Adobe Research, 2025) attacked the remaining shortcut: in most needle tests the question shares words with the needle, so the model can match strings instead of understanding. NoLiMa's needles and questions have minimal lexical overlap, so the model must make an associative step. Of 12 models claiming at least 128K, 10 fell below half of their short-context score by 32K.
Other suites move toward realistic tasks: LongBench draws on real long documents, and LongBench v2 adds expert-written questions; HELMET evaluates application-style tasks such as long-document QA, summarisation and citation; BABILong hides reasoning facts inside book text; NoCha asks for true or false claims about whole novels. The lesson across all of them is consistent: the harder and less literal the task, the earlier the decline.
Designing test cases that cannot be passed by accident
A home-grown harness measures your model behind your serving stack on your kind of content. Its design decisions are few but each matters.
Length is measured in the model's tokens. Count with the exact tokenizer and chat template the endpoint uses, and build each case to a token budget. Counting words or characters drifts by tens of percent between tokenizers, so your "64K" cell may really be 50K or 80K.
Depth is a fraction, placed at a sentence boundary. Insert the needle at a fraction of the haystack, 0% being the start and 100% the end, between sentences so it is not glued into a word.
Answers must not be guessable. Use random values, such as seven-digit codes bound to random keys, so the model cannot answer from what it learned in pre-training. A needle such as "the capital of France is Paris" measures memory, not context use.
Filler must be neutral and fresh. Widely reused filler, such as essays that appear in many public needle tests, may be memorised. Use text that is unrelated to the needle and, ideally, from your own corpus.
Distractors make it honest. Add several needles of the same form with different keys. A model that returns every number it saw must score zero, which strict scoring enforces.
Seeds make it repeatable. Generate every case from a fixed seed so a regression between two model versions is a difference in the model, not in the test.
A working length-by-depth harness
The harness below generates retrieval cases with distractors, sweeps a length by depth grid, scores strictly and computes effective length. model is any function from prompt to text; count_tokens should wrap the endpoint's real tokenizer. The code was run against fake models: one that always reads the answer (every cell 1.0), one blind to the middle half of its input (the U shape appears), and one that echoes every number (every cell 0.0, because strict scoring penalises distractors).
import random, re
from collections import defaultdict
ADJ = ["amber", "brisk", "cobalt", "dusty", "eager", "frosty", "gilded", "hollow"]
NOUN = ["falcon", "harbor", "lantern", "meadow", "orchard", "quarry", "summit", "willow"]
def make_case(rng, filler_sentences, count_tokens, length, depth, n_distractors=0):
keys = set()
while len(keys) < 1 + n_distractors:
keys.add(f"{rng.choice(ADJ)}-{rng.choice(NOUN)}-{rng.randint(100, 999)}")
keys = list(keys)
values = {k: str(rng.randint(1_000_000, 9_999_999)) for k in keys}
target = keys[0]
needle = f"The access code for {target} is {values[target]}."
others = [f"The access code for {k} is {values[k]}." for k in keys[1:]]
question = f"What is the access code for {target}? Answer with the number only."
budget = length - count_tokens(" ".join([needle, *others, question])) - 32
haystack, used = [], 0
while True:
s = filler_sentences[rng.randrange(len(filler_sentences))]
n = count_tokens(s)
if used + n > budget:
break
haystack.append(s)
used += n
for d in others: # distractors anywhere
haystack.insert(rng.randint(0, len(haystack)), d)
haystack.insert(round(depth * len(haystack)), needle) # target at the measured depth
return {"prompt": " ".join(haystack) + "\n\n" + question,
"answer": values[target], "wrong": {values[k] for k in keys[1:]}}
def score(case, output):
found = set(re.findall(r"\d{7}", output))
return int(case["answer"] in found and not (found & case["wrong"]))
def run_grid(model, filler, count_tokens, lengths, depths, reps=20, n_distractors=3, seed=0):
rng = random.Random(seed)
cells = defaultdict(list)
for length in lengths:
for depth in depths:
for _ in range(reps):
case = make_case(rng, filler, count_tokens, length, depth, n_distractors)
cells[(length, depth)].append(score(case, model(case["prompt"])))
return {k: sum(v) / len(v) for k, v in cells.items()}
def effective_length(grid, lengths, threshold):
best = 0
for length in sorted(lengths):
row = [acc for (l, _), acc in grid.items() if l == length]
if sum(row) / len(row) < threshold:
break # stop at the first failing length
best = length
return bestTwo choices in it are deliberate. Effective length stops at the first failing length rather than taking the longest passing one, because a model that fails at 32K and passes at 64K by luck has not earned 64K. And the threshold is an argument: RULER's 85.6% is one defensible choice, but the right bar is the accuracy your product needs.
Harder task families
Retrieval alone saturates quickly, so add generators for the harder families, each with synthetic content, a known answer and a token budget, reusing make_case's filler and placement code.
- Multi-value retrieval: bind several values to the same key at different depths and ask for all of them. Score recall over the set.
- Variable tracking: plant a chain such as
VAR_A = 48213, laterVAR_B = VAR_A, laterVAR_C = VAR_B, interleaved with unrelated chains, and ask which variables hold 48213. Chain length is a difficulty knob. - Aggregation: build the haystack from a word list with a controlled frequency distribution and ask for the most frequent words. There is no needle; the answer is spread over everything.
- Multi-hop QA: place two facts far apart, such as "Ilse works at the Varga lab" and "the Varga lab is in Tromso", and ask where Ilse works. Write the question so it shares no words with the second fact, which removes the string-matching shortcut NoLiMa identified.
Reserve model-graded scoring for open-ended tasks such as long-document summarisation, and calibrate the judge before trusting it, because judges have their own length and position biases.
Reading the results
Read the grid in three passes. Rows show length decay: where does mean accuracy first drop below your bar? Columns show positional bias: is the middle weaker than the edges, and does the sag grow with length?
Then size your uncertainty: a cell of n binary scores has a 95% interval of roughly 1.96 * sqrt(p * (1 - p) / n). At 20 cases and 70% accuracy that is about plus or minus 20 points, so neighbouring cells that differ by 10 points are noise. Spend extra repetitions on the cells that decide effective length, and compare model versions on the same seeded cases, which is far more sensitive than comparing two independent scores.
Finally, compare families. A typical profile holds retrieval at full length, loses tracking and aggregation much earlier, and loses low-overlap multi-hop earliest. Your effective length for a task is the one measured on the family that looks most like it.
What a sweep costs and how to run it
Prefill cost grows with every token, so count it before you run. A grid of six lengths from 4K to 128K sums to 252K tokens per depth and repetition; with 5 depths and 20 repetitions, one task family costs about 25 million input tokens. Five families are about 126 million. The 128K row dominates, because attention cost rises faster than linearly with length.
Three tactics cut it. First, run a coarse sweep (fewer repetitions, three depths) to find where the curve bends, then spend repetitions only near the bend. Second, reuse prefixes: cases that share filler before the needle can hit a prefix cache where the serving stack has one. Third, keep a small regression slice, such as 32K and 64K at three depths, that runs on every model or serving change, and the full grid only on releases.
Run the harness through the same serving path as production: same chat template, same maximum model length, same position-scaling configuration and same quantisation. Long-context quality is unusually sensitive to all four.
Failure modes and trade-offs
Most wrong long-context results come from the harness, not the model. Check for these before believing a number.
- Silent truncation. The serving stack clips prompts longer than its configured maximum, often from the left, so the needle disappears and the score looks like a model failure. Log the token count the server actually received for each case.
- Missing position scaling. A model extended with a RoPE scaling method needs the matching configuration at inference. Without it, accuracy collapses past the original training length.
- Leaky questions. A question that repeats the needle's wording measures string matching. Measure lexical overlap and keep it low for the families that claim to test understanding.
- Memorised content. Famous filler or well-known facts let the model answer without reading. Random values and private filler remove this.
- Sampling noise. Fix decoding settings, greedy where possible, across runs.
- Parser failures. Refusals, extra prose or formatted numbers score as wrong answers. Inspect a sample of failures by hand in every new run.
The trade-off running through all of this is realism against control. Synthetic tasks are cheap, exact and unmemorisable, but they overstate ability on real documents. Natural tasks predict product behaviour better but are costly to label and may leak into training data. Use synthetic grids to locate the cliff and real cases to confirm it.
What to do next
- Write down the bar: the accuracy your product needs on long inputs, per task type.
- Wire the harness above to your endpoint with the real tokenizer and chat template, and log received token counts.
- Run a coarse retrieval grid with distractors from 4K to the advertised window, then add tracking, aggregation and low-overlap multi-hop families.
- Compute effective length per family and set product limits, chunk sizes and retrieval depth from the lowest one that matches your use.
- Add a small seeded regression slice to CI for every model, quantisation or serving change.
- Build a few hundred natural cases from your own documents to confirm the synthetic results.
Keep learning: Context extension: training a model for long context, The math of RoPE context extension, Long-context prompting, Attention sinks and Calibrating LLM judges.