Long prompts cost money and time. A retrieval-augmented assistant that stuffs twenty passages into every request pays for all of them on every call, waits for all of them to be prefilled, and often answers worse because the passage that matters is buried in the middle. Prompt compression attacks the input side directly: it deletes the tokens the target model does not need and sends what is left. Done well, the model sees a shorter, denser prompt and answers about as well. Done badly, it deletes a negation, a number or a name and the answer is confidently wrong.
This page explains the LLMLingua family from Microsoft Research, the most widely used open-source compressors, from first principles: what signal they use to decide which tokens to drop, how each version improved on the last, how to call the library, and how to measure whether compression is helping your workload. How to divide a context window into budgets and keep compression compatible with caching at the architecture level is covered in prompt compression architecture. The library calls shown here were checked against the project README in October 2026.
The idea: tokens carry unequal information
A language model assigns every token a probability given the tokens before it. A token the model predicts easily, such as the second half of a common phrase or a function word, carries little information: if you delete it, a strong model can usually reconstruct the meaning. A token the model finds surprising, such as a rare name, a figure or a domain term, carries a lot. Information theory gives this a number, the self-information of a token, which is the negative log of its probability. Perplexity over a span is the same quantity averaged and exponentiated.
Every perplexity-based compressor follows from that observation. Run a small language model over the prompt, score each token or phrase by how surprising it is, and keep the surprising ones until you hit a budget. Selective Context (Li and colleagues, 2023) was an early version: it scored lexical units such as phrases and sentences by self-information under a small causal model and filtered out the least informative units. It worked, but it had two weaknesses that the LLMLingua line set out to fix. It treated every part of the prompt the same way, so an instruction was as likely to be cut as filler in a document. And it scored each unit independently, ignoring that once you delete one token, the surprise of the next one changes.
LLMLingua: budgets, iteration and alignment
LLMLingua (Jiang and colleagues, EMNLP 2023) added three pieces, and the paper reports up to 20x compression with small performance loss on its benchmarks.
The first is a budget controller. A prompt usually has an instruction, a set of demonstrations or documents, and a question. They do not deserve the same compression rate: the instruction and question are short and dense, while demonstrations are long and redundant. The controller assigns a higher keep-rate to the instruction and question and compresses the demonstrations harder. At a coarse level it ranks whole demonstrations by perplexity, keeps the most informative ones until the budget is spent and drops the rest, then passes what remains to the fine-grained stage.
The second is iterative token-level compression. Instead of scoring all tokens once, LLMLingua splits the text into segments and compresses them in order, scoring each segment conditioned on the already-compressed text before it. That way, the score of a token reflects what the target model will actually see, not what was in the original. Tokens whose perplexity exceeds a threshold derived from the budget are kept.
The third is distribution alignment. The small model doing the scoring and the large target model have different distributions, so a token that surprises the small model might be obvious to the large one or the reverse. The authors instruction-tuned the small model on text generated by the target model to bring the two closer.
LongLLMLingua: compress with the question in mind
Plain perplexity is question-blind. In retrieval-augmented generation that is the wrong objective: a passage can be full of surprising tokens and still irrelevant to the user's question. LongLLMLingua (Jiang and colleagues, ACL 2024) makes every stage question-aware, and its paper reports improving RAG performance by up to 21.4% while using about a quarter of the tokens on its benchmark.
At the coarse level it ranks documents by how much each one makes the question likely: roughly, the perplexity of the question conditioned on the document, so a document that explains the question scores well. At the fine level it uses contrastive perplexity, the change in a token's surprise when the question is present, which highlights tokens related to the question rather than tokens that are merely rare. It reorders documents so the most relevant appear first, countering the lost-in-the-middle effect where models attend poorly to the middle of long contexts. It assigns dynamic compression rates, keeping more of highly ranked documents. And it adds a subsequence-recovery step that maps spans in the model's answer back to the original uncompressed text, so names and figures the compressor mangled can be restored in the output.
LLMLingua-2: compression as token classification
Both earlier methods run a causal language model over the prompt, which costs real GPU time and still uses an indirect signal. LLMLingua-2 (Pan and colleagues, Findings of ACL 2024) reframes the problem. The authors asked GPT-4 to compress meeting transcripts from the MeetingBank dataset under strict rules: only delete words, never reorder, paraphrase or add. Aligning each compressed text back to its original gives a label for every word, keep or drop. They then trained a bidirectional encoder, XLM-RoBERTa large or multilingual BERT, as a token classifier on those labels.
At inference the classifier outputs a keep probability for every token in one forward pass, and the compressor keeps the highest-probability tokens until it reaches the requested rate. Because the encoder sees both directions of context it can judge a token by what follows, which a causal model cannot. The model is small and runs in one pass, and the README reports LLMLingua-2 is 3x to 6x faster than LLMLingua. Because it was trained on general text it is task-agnostic: it does not use the question, which makes it a good default for compressing long instructions, transcripts and tool outputs, and a weaker choice than LongLLMLingua when relevance to a specific question matters.
Using the library
The llmlingua package exposes one class, PromptCompressor, for all versions. Pass use_llmlingua2=True with one of the published LLMLingua-2 checkpoints to get the classifier; omit it to get the causal-model compressors. The compress_prompt method takes the text plus either a rate (fraction of tokens to keep) or a target_token count and returns a dictionary with compressed_prompt, origin_tokens, compressed_tokens, ratio and saving.
from llmlingua import PromptCompressor
# LLMLingua-2: a token classifier distilled from GPT-4 labels (bidirectional encoder, fast)
compressor = PromptCompressor(
model_name="microsoft/llmlingua-2-xlm-roberta-large-meetingbank",
use_llmlingua2=True,
)
result = compressor.compress_prompt(
passages_text, # the retrieved context only, never the system prompt
rate=0.33, # keep roughly a third of the tokens
force_tokens=["\n", "?"], # tokens the classifier must always keep
)
print(result["origin_tokens"], "->", result["compressed_tokens"], result["ratio"])
prompt = SYSTEM_PREFIX + result["compressed_prompt"] + "\n\nQuestion: " + questionFor question-aware compression of retrieved passages, pass a list of documents and the question, and turn on the LongLLMLingua options. The LongLLMLingua option values below are the ones the README uses in its own example; treat them as a starting point to tune.
# LongLLMLingua: question-aware, works on a list of documents (causal LM compressor)
long_compressor = PromptCompressor() # library default causal model; a GPU is expected
result = long_compressor.compress_prompt(
passages, # list[str], one entry per retrieved passage
question=question,
instruction="",
rate=0.25,
condition_in_question="after_condition",
reorder_context="sort", # most relevant passages first
dynamic_context_compression_ratio=0.3,
condition_compare=True,
context_budget="+100",
rank_method="longllmlingua",
)Two practical notes. The causal-model compressors load a multi-billion-parameter model, so plan GPU memory and latency for it; PromptCompressor("microsoft/phi-2") is a smaller documented option. And the token counts in the result come from a GPT-3.5 tiktoken encoding, not your target model's tokenizer, so recount with the target tokenizer when you budget.
Worked example: a support assistant with twenty passages
Suppose a support assistant retrieves 20 passages of about 300 tokens each for every question: 6,000 tokens of context, plus a 1,200-token system prompt with tool definitions and a 50-token question. At 100,000 requests a day, the context alone is 600 million input tokens a day.
Apply LLMLingua-2 at rate=0.33 to the passages only. Context drops to roughly 2,000 tokens, saving about 4,000 tokens per request, or 400 million input tokens a day. Multiply by your provider's input price to get the saving, then subtract the compressor's cost: an encoder forward pass over 6,000 tokens per request, on hardware you run. Prefill latency on the target model falls roughly in proportion to the tokens removed, but the compressor adds its own latency in front of it, so measure end-to-end time, not just the model call.
Notice what was not compressed. The 1,200-token system prompt is identical across requests, so the provider's prefix cache already makes it cheap; compressing it would change its bytes and could break that cache for little gain. The question is short and every token in it matters. Compression belongs on the long, variable, redundant part of the prompt, and only there.
Measure before you ship
Paper results come from paper benchmarks. Your documents, questions and target model are different, and the only number that matters is answer quality on your traffic at a given rate. Build a harness over a few hundred real questions with reference answers, run it uncompressed as a baseline, then sweep the rate and plot accuracy against tokens. Pick the most aggressive rate that stays within your tolerance of baseline, and record the compressor's own latency at p95.
import statistics, time
def evaluate(rate, cases, compressor, llm, grade):
rows = []
for case in cases: # 200+ real questions with reference answers
t0 = time.perf_counter()
out = compressor.compress_prompt(case.context, rate=rate, force_tokens=["\n", "?"])
t_comp = time.perf_counter() - t0
answer = llm(SYSTEM_PREFIX + out["compressed_prompt"] + "\n\nQuestion: " + case.question)
rows.append((grade(answer, case.reference), out["compressed_tokens"], t_comp))
acc = sum(r[0] for r in rows) / len(rows)
return {"rate": rate, "accuracy": acc,
"mean_tokens": statistics.mean(r[1] for r in rows),
"p95_compress_s": sorted(r[2] for r in rows)[int(0.95 * len(rows)) - 1]}
baseline = evaluate(1.0, cases, identity_compressor, llm, grade)
for rate in (0.6, 0.45, 0.33, 0.25):
print(evaluate(rate, cases, compressor, llm, grade))Grade with whatever you already trust: exact match on extracted fields, a rubric-based model judge, or human review on a sample. Split results by question type. Compression often holds up on summary-style questions and fails first on questions whose answer is a specific number, date or identifier, because those are exactly the tokens a general-purpose scorer is least sure about.
Failure modes
- Broken structure. Deleting tokens from JSON, code, SQL or tables produces text that no longer parses. Never compress structured payloads token by token; keep them whole or select whole rows and fields.
- Lost negations and qualifiers. Short words such as not, only, except and before are cheap to predict and easy to drop, and dropping them inverts meaning. Inspect compressed samples, not just scores.
- Mangled identifiers. Part numbers, account ids and URLs get split into many subword tokens and partly deleted. Protect them with
force_tokenswhere possible, or mask them with placeholders before compression and restore them afterwards. - Cache misses. Compressing a stable prefix changes its bytes on every edit of the compressor or its settings, and can defeat prompt caching entirely, costing more than it saves.
- Tokenizer mismatch. The library reports counts in a GPT-3.5 encoding, so 2,000 reported tokens may be 2,300 for your model and overflow a tight budget.
- Silent drift. Upgrading the compressor checkpoint or library changes outputs. Pin versions and rerun the harness on every change.
Trade-offs
| Option | Signal | Cost | Best for |
|---|---|---|---|
| LLMLingua-2 | Distilled keep/drop classifier | One encoder pass, fast | Task-agnostic compression of long text, transcripts, tool output |
| LongLLMLingua | Question-conditioned perplexity | Causal LM passes, slower | RAG where relevance to a specific question decides what to keep |
| LLMLingua | Iterative perplexity with budgets | Causal LM passes | Few-shot prompts with many redundant demonstrations |
| Better retrieval | Rerank, fewer passages | A reranker call | Often the first fix; sending 5 good passages beats compressing 20 |
| Summarisation | Generative rewrite | A full LLM call | Conversation history, where paraphrase is acceptable |
Compression competes with simpler fixes, and they should usually come first. A reranker that cuts twenty passages to five removes three quarters of the context without touching any sentence. Prefix caching makes stable text cheap without changing it. Compression earns its place when the variable context is long, redundant and still necessary after retrieval is tuned. For the broader budget picture see context window management and cost optimization, and for how models use long inputs see long-context prompting.
What to do next
- Measure your prompt by segment and find the long, variable, redundant part; that is the only candidate for compression.
- Tune retrieval and turn on prefix caching first, then measure again.
- Build an evaluation set of a few hundred real questions with reference answers and run it uncompressed as a baseline.
- Try LLMLingua-2 on the variable context and sweep the rate; try LongLLMLingua if relevance to the question matters more than speed.
- Protect identifiers, numbers and structured payloads with forced tokens, placeholders or by excluding them.
- Recount tokens with the target model's tokenizer and record compressor latency at p95.
- Pin the compressor version and rerun the harness whenever it, the target model or the retriever changes.