Run a language model with greedy decoding or a low temperature and sooner or later it writes the same sentence again, then again, until it hits the token limit. Repetition penalties are the inference-time patch: before sampling each token, lower the scores of tokens that have already appeared so the model is pushed toward something new. Nearly every serving stack exposes them, under similar names with quite different maths.
This article explains why loops happen, where penalties sit in the decoding pipeline, the exact formulas behind the multiplicative repetition penalty, the additive frequency and presence penalties, n-gram bans and DRY, and works through numbers that show why the most common one behaves strangely. It ends with the failure modes that penalties cause, which in practice are as common as the loops they fix, and a checklist for choosing settings. Parameter names and defaults are taken from the vLLM, llama.cpp and Hugging Face Transformers sources as read on 2026-10-03.
Why models repeat themselves
A decoder picks each token conditioned on everything before it, including its own previous output. If a phrase has appeared once, the context now contains evidence that this document repeats that phrase, and the model's probability of repeating it goes up. Xu and colleagues (NeurIPS 2022) measured this self-reinforcement effect: the more times a sentence has been repeated, the higher the probability of repeating it again, so a loop, once started, tends to tighten rather than break. Holtzman and colleagues' earlier study of neural text degeneration showed that maximisation-based decoding, greedy and beam search, falls into these loops far more than sampling does, because always taking the top token follows the reinforced path deterministically.
So loops are most common with greedy decoding, low temperature, small or heavily quantised models, long outputs, and prompts that are themselves repetitive, such as lists and templates. They are also a symptom of plain bugs: a wrong chat template or a missing end-of-sequence token leaves the model with no learned way to stop, and it fills the remaining budget with repeats. Check those causes before reaching for a penalty, as covered in sampling and decoding.
Where penalties sit in decoding
Penalties are logits processors. After the forward pass produces one score per vocabulary token, a chain of processors edits those scores before the sampler turns them into probabilities. Penalties need the token history; temperature divides all logits; top-k and top-p cut the tail; a structured-output mask, if used, sets disallowed tokens to minus infinity. Because the operations do not commute, applying a penalty before or after temperature gives different distributions. Libraries choose different orders, so a setting tuned on one stack does not transfer exactly to another.
The penalty family, formula by formula
Multiplicative repetition penalty. Introduced with the CTRL model (Keskar and colleagues, 2019), whose authors reported that a value around 1.2 with greedy decoding balanced truthful generation against repetition. For every token that already appears in the history, a positive logit is divided by the penalty and a negative logit multiplied by it, so the token always becomes less likely when the penalty is above 1. The sign split exists because dividing a negative logit would raise it. Values below 1 reward repetition. In Hugging Face Transformers the history for decoder-only models includes the prompt by default; vLLM's repetition_penalty likewise counts the prompt and the generated text. Each token is penalised once, however many times it occurred.
Frequency and presence penalties. These are additive. In the common formulation, a token's logit is reduced by its count so far times the frequency penalty, plus the presence penalty once if the count is above zero. OpenAI's API accepts both from -2.0 to 2.0, default 0, and describes frequency as scaling with how often a token has appeared and presence as applying once it has appeared at all. vLLM's presence_penalty and frequency_penalty follow the same idea and count only the generated text, not the prompt. Frequency grows with each reuse and so fights verbatim loops; presence is a flat nudge toward tokens not yet used.
N-gram bans. Transformers' no_repeat_ngram_size sets to minus infinity any token that would complete an n-gram already present. It guarantees no repeated n-gram of that length, which is exactly right for short summaries and exactly wrong for code, tables or names that must recur.
DRY (Don't Repeat Yourself). Available in llama.cpp, DRY looks for the token that would extend a sequence already seen earlier and penalises it in proportion to an exponential of the match length: short matches up to an allowed length cost nothing, long verbatim copies cost a lot. llama.cpp's defaults are multiplier 0.00 (disabled), base 1.75, allowed length 2, a 64-token look-back, and sequence breakers newline, colon, double quote and asterisk, which reset matching so ordinary formatting is not punished.
import numpy as np
def apply_penalties(logits, prompt_ids, output_ids,
repetition=1.0, frequency=0.0, presence=0.0,
penalise_prompt=True):
logits = logits.copy()
if repetition != 1.0:
seen = set(output_ids) | (set(prompt_ids) if penalise_prompt else set())
idx = np.fromiter(seen, dtype=np.int64)
x = logits[idx]
logits[idx] = np.where(x > 0, x / repetition, x * repetition)
if frequency or presence:
counts = np.bincount(np.asarray(output_ids, dtype=np.int64),
minlength=logits.shape[0])
logits -= counts * frequency + (counts > 0) * presence
return logitsllama.cpp exposes --repeat-penalty (default 1.00, meaning disabled) with --repeat-last-n (default 64), so only the last 64 tokens count, plus --presence-penalty and --frequency-penalty at 0.00. The window matters: a document-wide history penalises every common word eventually, while a short window targets loops.
Worked numbers: why the multiplicative penalty is not calibrated
Take four candidate tokens with logits 4.0, 3.0, 1.0 and -1.0. Softmax gives probabilities 0.702, 0.258, 0.035 and 0.005. Suppose token A, the 4.0 one, has already appeared three times.
A repetition penalty of 1.2 turns 4.0 into 3.33. A's probability falls from 0.702 to 0.547 and B rises to 0.392: A is still the favourite, just less dominant. A frequency penalty of 0.5 plus a presence penalty of 0.5 subtracts 0.5 times 3 plus 0.5, so 2.0, leaving A at 2.0. Its probability falls to 0.242 and B becomes the clear favourite at 0.657. Additive penalties scale with how often a token was used; the multiplicative one does not.
Now the strange part. Add 10 to every logit. Softmax is unchanged, 0.702 for A, because only differences between logits matter. But the multiplicative penalty now divides 14.0 by 1.2, giving 11.67, a cut of 2.33 instead of 0.67, and A's probability drops to 0.186. The same penalty value has a 3.5-times-larger effect purely because of where the logits sit. Different models, and the same model at different positions, centre their logits differently, so a multiplicative penalty is not a calibrated knob. Additive penalties shift a logit by a fixed amount and do not have this problem, which is one reason to prefer them when your stack offers both.
Tokens, prompts and copying
Penalties act on tokens, not words or ideas, and that has consequences explained in tokenization. Common tokens such as a leading-space "the", commas, newlines and indentation appear constantly in any long text. A document-wide multiplicative penalty that counts the prompt lowers all of them, and the model starts dropping articles, choosing odd punctuation or breaking indentation. Words split into several tokens are penalised piecemeal, so the model may avoid a correct spelling's first token and produce a misspelling or rare synonym.
Counting the prompt is especially harmful when the answer must copy from it. Retrieval-augmented question answering, extraction, translation and code editing all require reproducing names, identifiers and numbers that appear in the context. A repetition penalty over prompt tokens pushes the model away from exactly those, so it paraphrases a product name or invents a slightly wrong function name. Structured output suffers the same way: JSON repeats braces, quotes and key names by design, and even with a grammar mask from guided decoding keeping it valid, penalties bend the content toward unusual values and keys.
Worked example: fixing looping ticket summaries
Consider an illustrative case. A team runs an 8B instruction model at temperature 0 to summarise support tickets into five bullet points. About 3 percent of summaries end with the same bullet repeated until the 512-token limit. Someone sets repetition penalty 1.3, the loops disappear, and two weeks later customers report summaries that misspell product names and drop words.
The disciplined fix goes in order. First, check the cause: inspecting the looping outputs shows the model never emits its end-of-turn token, because the prompt was formatted without the model's chat template. Fixing the template cuts loops from 3 percent to 0.4 percent with no penalty at all. Second, cap length sensibly: five bullets need about 150 tokens, so max_tokens of 250 bounds any remaining loop. Third, add a gentle, output-only penalty: frequency 0.3, presence 0, which does not touch prompt tokens, so product names copied from the ticket are safe. Fourth, add a cheap guard after generation: if any line repeats verbatim, regenerate once with temperature 0.3. Loops fall below 0.1 percent and name accuracy recovers to its baseline. The repetition penalty is removed entirely.
def has_loop(text: str, min_len: int = 20) -> bool:
lines = [l.strip() for l in text.splitlines() if len(l.strip()) >= min_len]
return len(lines) != len(set(lines))
out = generate(prompt, temperature=0.0, max_tokens=250, frequency_penalty=0.3)
if has_loop(out):
out = generate(prompt, temperature=0.3, max_tokens=250, frequency_penalty=0.3)
Failure modes
Penalties fail in recognisable ways; knowing them lets you spot over-tuning quickly.
- Copy errors. Names, identifiers and numbers from the context come back altered. Cause: prompt tokens penalised. Fix: output-only penalties or none.
- Degraded fluency. Missing articles, odd punctuation, a thesaurus feel in long outputs. Cause: penalty too strong or history too long. Fix: smaller value, shorter window.
- Broken code and data. Indentation drifts, variable names change mid-function, JSON keys are renamed. Fix: disable penalties for code and structured tasks.
- Loops that change shape. The model escapes verbatim repetition by paraphrasing the same sentence. Token penalties cannot see meaning; fix the cause or detect semantic repeats at the application layer.
- Settings that do not transfer. A value tuned on one model or library behaves differently on another because of logit scale and processor order, as the numbers above show. Re-tune after any model or engine change.
- Masking a real bug. A penalty that suppresses loops caused by a missing stop token leaves the model generating past the natural end. Always look at why outputs did not stop.
Trade-offs and alternatives
Penalties are cheap, per-request and need no training, which is their whole appeal. Their cost is a bias against legitimate repetition that you cannot target. The alternatives each trade differently. Sampling at a moderate temperature instead of greedy breaks most loops while keeping copies exact. Stop sequences and a sensible maximum length bound the damage of any loop. Post-generation detection and retry fixes rare loops without biasing every output. Training-side fixes such as unlikelihood training (Welleck and colleagues) or better instruction data remove the tendency itself but require a training run. A good default for production is: no repetition penalty, no presence penalty, frequency penalty between 0 and 0.5 on open-ended prose only, moderate temperature where determinism is not required, and a loop detector. Raise penalties only with an evaluation that measures both the repetition rate and task accuracy. For long-output budgeting see context length.
What to do next
- Measure first: sample 500 production outputs and compute the rate of verbatim repeated lines and the share hitting the token limit.
- Inspect looping outputs for a wrong chat template or missing end-of-sequence token and fix those before tuning anything.
- Read your engine's documentation for which history each penalty counts (prompt, output, last-n window) and in what order processors run.
- Turn off penalties for extraction, RAG answers, code and structured output; start with frequency 0.2 to 0.5 for open-ended prose.
- Add stop sequences, a task-appropriate maximum length and a cheap loop detector with one retry.
- Build an evaluation that scores repetition and task accuracy together, including exact-match on copied names, and re-run it after any model or engine change.