Greedy decoding is the simplest way to turn a language model into text. At every step, take the single most likely next token, append it, and repeat. It has no randomness and no search, and nothing to tune except when to stop. It is the default in many libraries when sampling is off, it is what "temperature 0" usually means, and it is the reference behaviour that speculative decoding must reproduce.
This page explains greedy decoding from the logits up. It covers the loop with a KV cache, why the most likely token at each step does not give the most likely sequence, how greedy output degenerates into loops, and why temperature 0 is often not reproducible in production. It also shows how greedy verification makes speculative decoding exact, and when greedy is the right choice. It ends with code, a test harness and a checklist.
The argmax loop
A causal language model maps a prefix of tokens to a vector of logits, one score per vocabulary entry, for the next position. Softmax turns logits into probabilities, but greedy decoding never needs probabilities. Softmax is monotonic, so the index of the largest logit is the index of the largest probability. Greedy decoding is next = argmax(logits[-1]), repeated until an end-of-sequence token, a stop string or a token limit.
In practice the loop reuses the KV cache. The prompt is processed once (prefill), and each later step feeds only the newest token. Here is the loop with Hugging Face Transformers, written out so every step is visible:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(MODEL, torch_dtype=torch.bfloat16).eval()
ids = tok(prompt, return_tensors="pt").input_ids
past, generated = None, []
with torch.no_grad():
for _ in range(max_new_tokens):
step_input = ids if past is None else ids[:, -1:] # prefill once, then one token
out = model(input_ids=step_input, past_key_values=past, use_cache=True)
past = out.past_key_values
next_id = out.logits[:, -1, :].argmax(dim=-1, keepdim=True)
ids = torch.cat([ids, next_id], dim=-1)
generated.append(next_id.item())
if next_id.item() == tok.eos_token_id:
break
print(tok.decode(generated))The library equivalent is model.generate(**inputs, do_sample=False, num_beams=1, max_new_tokens=256). In vLLM, SamplingParams(temperature=0, max_tokens=256) selects greedy. Ties are broken by index: torch.argmax returns the first maximal index. Exact ties are rare in practice, but near-ties are not, and they matter below.
Locally best is not globally best
Greedy maximises each step's probability, not the probability of the whole output. In the diagram, step one offers A (0.50) or B (0.40). Greedy takes A. Its best continuation x has probability 0.40, so the pair scores 0.50 times 0.40, which is 0.20. The path through B continues with z at 0.90, giving 0.36. Greedy never looks at it. Beam search with width 2 keeps both A and B alive and finds the better pair.
Two things follow. First, greedy is a lower bound on what search can find under the model's own scoring. Second, finding the highest-probability sequence is often not what you want for open-ended text. The highest-probability continuation of a story tends to be bland and repetitive. That is why chat systems usually sample, and why beam search is mainly used where outputs are short and tightly constrained, such as translation or structured extraction.
Degeneration and repetition loops
Greedy decoding's characteristic failure is degeneration: it repeats a phrase, a sentence or a list item until it hits the token limit. Holtzman and colleagues documented this in "The Curious Case of Neural Text Degeneration" (2019). The mechanism is a feedback loop. Once a phrase appears in the context, the model assigns higher probability to repeating it, and greedy always takes the highest. A deterministic decoder has no way to escape a self-reinforcing cycle.
Mitigations, from least to most invasive:
- Stop conditions: an end-of-sequence token, stop strings such as a closing brace or a blank line, and a token limit sized to the task.
no_repeat_ngram_sizein Hugging Face forbids any n-gram from appearing twice. This is effective but blunt, because legitimate repeats such as code identifiers and table headers get blocked.repetition_penaltyrescales logits of tokens already present. Greedy then picks the argmax of the penalised logits, so output stays deterministic. Small values (around 1.1 to 1.2) are typical starting points to tune, not rules.- Detect and retry: watch the output for a repeated suffix and, when it triggers, cut back to the last clean point. Then fall back to low-temperature sampling or a different prompt.
- Switch strategy. Several reasoning-model cards explicitly advise against greedy decoding because long chains of thought loop under it. Check the card for the model you serve.
def repeated_suffix(token_ids, min_len=8, max_len=64):
"""Return the period of a suffix that repeats back-to-back, else 0."""
n = len(token_ids)
for L in range(min_len, min(max_len, n // 2) + 1):
if token_ids[n - L:] == token_ids[n - 2 * L:n - L]:
return L
return 0
Why temperature 0 is not reproducible
Teams pick temperature 0 to get reproducible outputs, then find that the same prompt gives different answers on different calls. Greedy is a deterministic function of the logits, but the logits are not a deterministic function of the prompt in a serving system.
Floating-point addition is not associative, so a matrix multiply or attention reduction that sums in a different order gives slightly different results. Inference kernels choose tiling and reduction strategies based on batch size and sequence layout. In a continuously batched server, your request shares a batch with whatever else is in flight, so the same prompt can go through differently shaped kernels on different calls. The differences are tiny, usually in the last bits of bfloat16. But when two candidate tokens are within that tiny margin, the argmax flips. From then on the context differs and every later token can diverge. Horace He and colleagues at Thinking Machines analysed this in "Defeating Nondeterminism in LLM Inference" (2025). They traced it to the lack of batch invariance in kernels, not to concurrency races, and showed batch-invariant kernels restore exact repeatability at some throughput cost.
What to do with this:
- Do not promise bit-identical outputs from a hosted API or a batched server at temperature 0 unless the provider documents it.
- For evaluations, record the full generated text, not just the score, so you can tell model changes from decode noise.
- Measure near-ties directly: log the gap between the top two logits at each step. Small gaps mark the positions where outputs can split.
- If you need exact repeatability, run at batch size 1 on a fixed software and hardware stack, or use an engine's documented batch-invariant mode, and test it with the harness below.
Logging the margin costs almost nothing inside the loop shown earlier. Take the top two logits at each step and record their difference next to the chosen token. In bfloat16, gaps below a few hundredths are the ones worth flagging. Treat that threshold as a starting point to calibrate against your own determinism runs, not as a constant.
top2 = out.logits[:, -1, :].float().topk(2, dim=-1).values
margin = (top2[:, 0] - top2[:, 1]).item()
margins.append(margin) # inspect the smallest margins after the runThe determinism report below calls the endpoint repeatedly while other traffic is running. One distinct output means the path is stable under that load. Several distinct outputs, together with small logged margins at the step where they first differ, confirm the floating-point explanation and not a bug in your own code.
from collections import Counter
def determinism_report(generate, prompt, runs=20):
"""generate(prompt) -> str. Call it under realistic concurrent load."""
outs = Counter(generate(prompt) for _ in range(runs))
top, count = outs.most_common(1)[0]
return {"distinct_outputs": len(outs), "modal_share": count / runs}
Greedy verification in speculative decoding
Speculative decoding speeds up generation by letting a small draft model propose k tokens. The large target model then scores all k positions in one forward pass. Under greedy decoding the acceptance rule is exact match. Keep the longest prefix of draft tokens where each equals the target's argmax at that position. At the first mismatch, emit the target's own argmax instead. If all k match, the same pass yields one extra token for free.
def speculative_greedy_step(target, draft, ctx, k):
proposal = draft.greedy(ctx, k) # k draft tokens
logits = target.forward(ctx + proposal) # one pass, k+1 positions scored
accepted = []
for i, tok in enumerate(proposal):
best = argmax(logits[len(ctx) - 1 + i])
if tok != best:
return accepted + [best] # target's choice at first mismatch
accepted.append(tok)
return accepted + [argmax(logits[len(ctx) + k - 1])] # bonus tokenEach emitted token is the target's argmax given the tokens before it, so the output matches the target's greedy decode, apart from the floating-point caveats above, since the verification pass is batched differently. The speed-up depends only on how often the draft agrees with the target. Greedy targets are the easiest case to reason about and to test: run the same prompts with and without speculation and diff the outputs.
When greedy is the right tool
| Situation | Greedy? | Why |
|---|---|---|
| Extraction, classification, short structured answers | Yes | One correct answer; variety is a bug |
| Regression tests and evals | Yes, with the caveats above | Lowest-variance baseline across model versions |
| Code completion of a few lines | Usually | Short horizon, little room to loop |
| Long-form writing, chat | No | Bland, repetitive text |
| Long chain-of-thought reasoning | Check the model card | Some models loop under greedy |
| Self-consistency or best-of-n | No | Needs diverse samples by design |
Greedy costs one forward step per token and no extra memory, the same as sampling, and less than beam search, which carries a KV cache per beam. When you need valid JSON or a fixed label set, combine greedy with constrained decoding: mask invalid tokens, then take the argmax of what remains.
For classification, a single greedy step is often all you need, and it is worth doing explicitly. Put the candidate labels in the prompt, run one forward pass, and compare the logits of each label's first token. If two labels share a first token, compare full-label log-probabilities instead. This gives the greedy answer plus a confidence margin for free, and it cannot produce an off-list label. That makes it more robust than generating free text and parsing it afterwards.
Worked example: looping JSON extraction
A team extracts invoice fields into JSON with an instruction-tuned model. It runs greedy at temperature 0 and gets about 2 percent of outputs that never close the JSON object. On inspection they repeat a line item until the 512-token limit.
Diagnosis: invoices with many near-identical line items put repeated text in the prompt. Greedy then copies the pattern one item too many and loops. Logging the top-two logit gap shows the gap shrinking to near zero at the point where the copy should stop.
Fix, in order: add a stop condition on the closing brace and cap max_new_tokens at about twice the longest valid output. Add the repeated-suffix detector, and on a hit cut back and retry once with a JSON grammar constraint. Keep greedy, because the task has one right answer. Track the invalid-output rate as a metric. Repetition penalties were tried and rejected, because they altered legitimately repeated SKU codes.
What to do next
- Write down why each endpoint uses greedy or sampling. Use greedy only where there is one right answer.
- Set explicit stop strings and a token limit for every greedy call.
- Add the repeated-suffix detector and count hits per endpoint.
- Run the determinism report under real load. If outputs split, log top-two logit gaps to find where.
- Store full outputs for evals, and never compare scores alone across serving changes.
- If you add speculative decoding, diff its greedy outputs against the non-speculative path on a fixed prompt set before you ship.
Related reading: sampling and decoding strategies, beam search, temperature, repetition penalty and speculative decoding.