Long prompts cost money and time. A retrieval-augmented assistant that stuffs twenty passages into every request pays for all of them on every call, waits for all of them to be prefilled, and often answers worse because the passage that matters is buried in the middle. Prompt compression attacks the input side directly: it deletes the tokens the target model does not need and sends what is left. Done well, the model sees a shorter, denser prompt and answers about as well. Done badly, it deletes a negation, a number or a name and the answer is confidently wrong.

This page explains the LLMLingua family from Microsoft Research, the most widely used open-source compressors, from first principles: what signal they use to decide which tokens to drop, how each version improved on the last, how to call the library, and how to measure whether compression is helping your workload. How to divide a context window into budgets and keep compression compatible with caching at the architecture level is covered in prompt compression architecture. The library calls shown here were checked against the project README in October 2026.

Advertisement

The idea: tokens carry unequal information

A language model assigns every token a probability given the tokens before it. A token the model predicts easily, such as the second half of a common phrase or a function word, carries little information: if you delete it, a strong model can usually reconstruct the meaning. A token the model finds surprising, such as a rare name, a figure or a domain term, carries a lot. Information theory gives this a number, the self-information of a token, which is the negative log of its probability. Perplexity over a span is the same quantity averaged and exponentiated.

Every perplexity-based compressor follows from that observation. Run a small language model over the prompt, score each token or phrase by how surprising it is, and keep the surprising ones until you hit a budget. Selective Context (Li and colleagues, 2023) was an early version: it scored lexical units such as phrases and sentences by self-information under a small causal model and filtered out the least informative units. It worked, but it had two weaknesses that the LLMLingua line set out to fix. It treated every part of the prompt the same way, so an instruction was as likely to be cut as filler in a document. And it scored each unit independently, ignoring that once you delete one token, the surprise of the next one changes.

LLMLingua: budgets, iteration and alignment

LLMLingua (Jiang and colleagues, EMNLP 2023) added three pieces, and the paper reports up to 20x compression with small performance loss on its benchmarks.

The first is a budget controller. A prompt usually has an instruction, a set of demonstrations or documents, and a question. They do not deserve the same compression rate: the instruction and question are short and dense, while demonstrations are long and redundant. The controller assigns a higher keep-rate to the instruction and question and compresses the demonstrations harder. At a coarse level it ranks whole demonstrations by perplexity, keeps the most informative ones until the budget is spent and drops the rest, then passes what remains to the fine-grained stage.

The second is iterative token-level compression. Instead of scoring all tokens once, LLMLingua splits the text into segments and compresses them in order, scoring each segment conditioned on the already-compressed text before it. That way, the score of a token reflects what the target model will actually see, not what was in the original. Tokens whose perplexity exceeds a threshold derived from the budget are kept.

The third is distribution alignment. The small model doing the scoring and the large target model have different distributions, so a token that surprises the small model might be obvious to the large one or the reverse. The authors instruction-tuned the small model on text generated by the target model to bring the two closer.

Advertisement

LongLLMLingua: compress with the question in mind

Plain perplexity is question-blind. In retrieval-augmented generation that is the wrong objective: a passage can be full of surprising tokens and still irrelevant to the user's question. LongLLMLingua (Jiang and colleagues, ACL 2024) makes every stage question-aware, and its paper reports improving RAG performance by up to 21.4% while using about a quarter of the tokens on its benchmark.

At the coarse level it ranks documents by how much each one makes the question likely: roughly, the perplexity of the question conditioned on the document, so a document that explains the question scores well. At the fine level it uses contrastive perplexity, the change in a token's surprise when the question is present, which highlights tokens related to the question rather than tokens that are merely rare. It reorders documents so the most relevant appear first, countering the lost-in-the-middle effect where models attend poorly to the middle of long contexts. It assigns dynamic compression rates, keeping more of highly ranked documents. And it adds a subsequence-recovery step that maps spans in the model's answer back to the original uncompressed text, so names and figures the compressor mangled can be restored in the output.

LLMLingua-2: compression as token classification

Both earlier methods run a causal language model over the prompt, which costs real GPU time and still uses an indirect signal. LLMLingua-2 (Pan and colleagues, Findings of ACL 2024) reframes the problem. The authors asked GPT-4 to compress meeting transcripts from the MeetingBank dataset under strict rules: only delete words, never reorder, paraphrase or add. Aligning each compressed text back to its original gives a label for every word, keep or drop. They then trained a bidirectional encoder, XLM-RoBERTa large or multilingual BERT, as a token classifier on those labels.

At inference the classifier outputs a keep probability for every token in one forward pass, and the compressor keeps the highest-probability tokens until it reaches the requested rate. Because the encoder sees both directions of context it can judge a token by what follows, which a causal model cannot. The model is small and runs in one pass, and the README reports LLMLingua-2 is 3x to 6x faster than LLMLingua. Because it was trained on general text it is task-agnostic: it does not use the question, which makes it a good default for compressing long instructions, transcripts and tool outputs, and a weaker choice than LongLLMLingua when relevance to a specific question matters.

Using the library

The llmlingua package exposes one class, PromptCompressor, for all versions. Pass use_llmlingua2=True with one of the published LLMLingua-2 checkpoints to get the classifier; omit it to get the causal-model compressors. The compress_prompt method takes the text plus either a rate (fraction of tokens to keep) or a target_token count and returns a dictionary with compressed_prompt, origin_tokens, compressed_tokens, ratio and saving.

from llmlingua import PromptCompressor

# LLMLingua-2: a token classifier distilled from GPT-4 labels (bidirectional encoder, fast)
compressor = PromptCompressor(
    model_name="microsoft/llmlingua-2-xlm-roberta-large-meetingbank",
    use_llmlingua2=True,
)

result = compressor.compress_prompt(
    passages_text,              # the retrieved context only, never the system prompt
    rate=0.33,                  # keep roughly a third of the tokens
    force_tokens=["\n", "?"],   # tokens the classifier must always keep
)
print(result["origin_tokens"], "->", result["compressed_tokens"], result["ratio"])
prompt = SYSTEM_PREFIX + result["compressed_prompt"] + "\n\nQuestion: " + question

For question-aware compression of retrieved passages, pass a list of documents and the question, and turn on the LongLLMLingua options. The LongLLMLingua option values below are the ones the README uses in its own example; treat them as a starting point to tune.

# LongLLMLingua: question-aware, works on a list of documents (causal LM compressor)
long_compressor = PromptCompressor()   # library default causal model; a GPU is expected

result = long_compressor.compress_prompt(
    passages,                           # list[str], one entry per retrieved passage
    question=question,
    instruction="",
    rate=0.25,
    condition_in_question="after_condition",
    reorder_context="sort",             # most relevant passages first
    dynamic_context_compression_ratio=0.3,
    condition_compare=True,
    context_budget="+100",
    rank_method="longllmlingua",
)

Two practical notes. The causal-model compressors load a multi-billion-parameter model, so plan GPU memory and latency for it; PromptCompressor("microsoft/phi-2") is a smaller documented option. And the token counts in the result come from a GPT-3.5 tiktoken encoding, not your target model's tokenizer, so recount with the target tokenizer when you budget.

Worked example: a support assistant with twenty passages

Suppose a support assistant retrieves 20 passages of about 300 tokens each for every question: 6,000 tokens of context, plus a 1,200-token system prompt with tool definitions and a 50-token question. At 100,000 requests a day, the context alone is 600 million input tokens a day.

Apply LLMLingua-2 at rate=0.33 to the passages only. Context drops to roughly 2,000 tokens, saving about 4,000 tokens per request, or 400 million input tokens a day. Multiply by your provider's input price to get the saving, then subtract the compressor's cost: an encoder forward pass over 6,000 tokens per request, on hardware you run. Prefill latency on the target model falls roughly in proportion to the tokens removed, but the compressor adds its own latency in front of it, so measure end-to-end time, not just the model call.

Notice what was not compressed. The 1,200-token system prompt is identical across requests, so the provider's prefix cache already makes it cheap; compressing it would change its bytes and could break that cache for little gain. The question is short and every token in it matters. Compression belongs on the long, variable, redundant part of the prompt, and only there.

Where a compressor sits in a retrieval pipelineStable prefixsystem + tools + rulesRetrieved passagesvariable, longUser questionshort, never cutRank and reorderquestion-aware (coarse)Token pruningsmall LM or classifier (fine)Assembled promptprefix | compressed | questionTarget LLMprefix cache still hitsEval harness: answer accuracy and cost at each ratepick the rate from data, not from the paperkept tokensuntoucheduntouched
Compress only the variable retrieved context. The stable prefix stays byte-identical so the prefix cache keeps hitting, and the question passes through untouched.

Measure before you ship

Paper results come from paper benchmarks. Your documents, questions and target model are different, and the only number that matters is answer quality on your traffic at a given rate. Build a harness over a few hundred real questions with reference answers, run it uncompressed as a baseline, then sweep the rate and plot accuracy against tokens. Pick the most aggressive rate that stays within your tolerance of baseline, and record the compressor's own latency at p95.

import statistics, time

def evaluate(rate, cases, compressor, llm, grade):
    rows = []
    for case in cases:                       # 200+ real questions with reference answers
        t0 = time.perf_counter()
        out = compressor.compress_prompt(case.context, rate=rate, force_tokens=["\n", "?"])
        t_comp = time.perf_counter() - t0
        answer = llm(SYSTEM_PREFIX + out["compressed_prompt"] + "\n\nQuestion: " + case.question)
        rows.append((grade(answer, case.reference), out["compressed_tokens"], t_comp))
    acc = sum(r[0] for r in rows) / len(rows)
    return {"rate": rate, "accuracy": acc,
            "mean_tokens": statistics.mean(r[1] for r in rows),
            "p95_compress_s": sorted(r[2] for r in rows)[int(0.95 * len(rows)) - 1]}

baseline = evaluate(1.0, cases, identity_compressor, llm, grade)
for rate in (0.6, 0.45, 0.33, 0.25):
    print(evaluate(rate, cases, compressor, llm, grade))

Grade with whatever you already trust: exact match on extracted fields, a rubric-based model judge, or human review on a sample. Split results by question type. Compression often holds up on summary-style questions and fails first on questions whose answer is a specific number, date or identifier, because those are exactly the tokens a general-purpose scorer is least sure about.

Failure modes

  • Broken structure. Deleting tokens from JSON, code, SQL or tables produces text that no longer parses. Never compress structured payloads token by token; keep them whole or select whole rows and fields.
  • Lost negations and qualifiers. Short words such as not, only, except and before are cheap to predict and easy to drop, and dropping them inverts meaning. Inspect compressed samples, not just scores.
  • Mangled identifiers. Part numbers, account ids and URLs get split into many subword tokens and partly deleted. Protect them with force_tokens where possible, or mask them with placeholders before compression and restore them afterwards.
  • Cache misses. Compressing a stable prefix changes its bytes on every edit of the compressor or its settings, and can defeat prompt caching entirely, costing more than it saves.
  • Tokenizer mismatch. The library reports counts in a GPT-3.5 encoding, so 2,000 reported tokens may be 2,300 for your model and overflow a tight budget.
  • Silent drift. Upgrading the compressor checkpoint or library changes outputs. Pin versions and rerun the harness on every change.

Trade-offs

OptionSignalCostBest for
LLMLingua-2Distilled keep/drop classifierOne encoder pass, fastTask-agnostic compression of long text, transcripts, tool output
LongLLMLinguaQuestion-conditioned perplexityCausal LM passes, slowerRAG where relevance to a specific question decides what to keep
LLMLinguaIterative perplexity with budgetsCausal LM passesFew-shot prompts with many redundant demonstrations
Better retrievalRerank, fewer passagesA reranker callOften the first fix; sending 5 good passages beats compressing 20
SummarisationGenerative rewriteA full LLM callConversation history, where paraphrase is acceptable

Compression competes with simpler fixes, and they should usually come first. A reranker that cuts twenty passages to five removes three quarters of the context without touching any sentence. Prefix caching makes stable text cheap without changing it. Compression earns its place when the variable context is long, redundant and still necessary after retrieval is tuned. For the broader budget picture see context window management and cost optimization, and for how models use long inputs see long-context prompting.

What to do next

  1. Measure your prompt by segment and find the long, variable, redundant part; that is the only candidate for compression.
  2. Tune retrieval and turn on prefix caching first, then measure again.
  3. Build an evaluation set of a few hundred real questions with reference answers and run it uncompressed as a baseline.
  4. Try LLMLingua-2 on the variable context and sweep the rate; try LongLLMLingua if relevance to the question matters more than speed.
  5. Protect identifiers, numbers and structured payloads with forced tokens, placeholders or by excluding them.
  6. Recount tokens with the target model's tokenizer and record compressor latency at p95.
  7. Pin the compressor version and rerun the harness whenever it, the target model or the retriever changes.
Key takeaway: Prompt compression deletes the tokens a target model is least likely to need, scored by a small model. LLMLingua added per-segment budgets, iterative scoring and distribution alignment; LongLLMLingua made every stage question-aware and reorders passages; LLMLingua-2 replaced perplexity with a fast, task-agnostic token classifier distilled from GPT-4 labels. Apply it only to long, variable context, never to a cached prefix, structured data or the question, and choose the rate from an evaluation on your own traffic rather than from paper figures.