Most people meet temperature as a slider in a chat playground: turn it down for facts, up for poems. That picture is not wrong, but it hides the fact that the same scalar appears in at least five different places in an LLM system, and the right value in each comes from a different procedure. In sampling you choose it per request. In best-of-n generation and reinforcement-learning rollouts you choose it per run, and the best value depends on how many samples you draw. In calibration you fit it on held-out data. In distillation you fix it on both teacher and student and must correct the gradient scale. In contrastive embedding training it is often a learned parameter that can blow up if left unclamped.
This article starts from the one equation all five share, works a numeric example by hand, and then walks each job with code and the mistakes that show up in practice. It deliberately does not repeat the sampler-ordering and provider-limits material already covered in Temperature, Top-p, Top-k, in depth or the entropy algebra in Sampling mathematics; read those for the decoding knobs in detail.
The equation and a worked example
A model produces a vector of logits z, one real number per vocabulary token or class. The softmax turns them into probabilities. Temperature inserts a single division first: p_i = exp(z_i / T) / sum_j exp(z_j / T). Three consequences follow directly from that formula and are worth stating precisely, because most confusion about temperature comes from forgetting one of them.
- The ranking never changes. Dividing every logit by the same positive number keeps their order, so the argmax, and therefore greedy decoding and top-1 accuracy, is identical at every T.
- Ratios are what move. The ratio of two probabilities is p_i / p_j = exp((z_i - z_j) / T). A logit gap of 2 means a 7.4x preference at T = 1, a 55x preference at T = 0.5 and a 2.7x preference at T = 2.
- The limits are degenerate. As T approaches 0 all mass collapses onto the argmax; as T grows the distribution approaches uniform. T = 0 itself is a division by zero, which is why APIs implement it as a special case that switches to greedy selection.
A worked example makes the size of the effect concrete. Take four logits [4, 3, 2, 0]. The numbers below were computed exactly; entropy is in nats, and the last column is exp(entropy), the effective number of choices the distribution behaves like.
| T | p(token 1..4) | entropy | effective choices |
|---|---|---|---|
| 0.5 | 0.867, 0.117, 0.016, 0.0003 | 0.444 | 1.56 |
| 1.0 | 0.657, 0.242, 0.089, 0.012 | 0.888 | 2.43 |
| 2.0 | 0.474, 0.288, 0.174, 0.064 | 1.193 | 3.30 |
Halving T took the long-tail token from 1.2 percent to 0.03 percent, a fortyfold drop, while the top token gained only 21 points. Doubling T multiplied the tail token by more than five. That asymmetry is the whole story of temperature at decode time: it mostly controls how often the tail gets picked, and the tail is where both creativity and nonsense live.
Temperature at decode time
At inference, temperature is applied independently at every generated position. That matters because a sequence probability is a product of per-step probabilities. A small per-step change compounds: if a 400 token answer has a one-in-a-hundred chance per step of picking a tail token at T = 1, the chance of at least one such pick in the whole answer is about 98 percent. Lowering T does not make one token much better; it makes long answers much less likely to contain a single derailing token.
The minimal sampler is a few lines. Real serving engines add truncation (top-k, top-p, min-p), penalties and batching, and the order in which those are applied relative to temperature changes the result; the detailed pipeline is in LLM sampling and decoding architecture.
import numpy as np
def sample(logits: np.ndarray, temperature: float, rng: np.random.Generator) -> int:
if temperature == 0.0: # special-cased: greedy
return int(np.argmax(logits))
z = logits / temperature
z = z - z.max() # stability: exp never overflows
probs = np.exp(z)
probs /= probs.sum()
return int(rng.choice(len(probs), p=probs))Two practical points are easy to miss. First, setting temperature to 0 does not guarantee identical output across calls: batching, floating-point reduction order and server-side changes can flip near-ties in the argmax, so treat T = 0 as low variance, not as determinism. Second, some providers fix or ignore temperature for particular models, notably reasoning-oriented ones; check the documentation of the model you call rather than assuming the parameter is honoured.
Many samples: pass@k, voting and RL rollouts
When you draw many samples and keep the best, the right temperature is not the one that maximises the quality of a single sample. The Codex paper (Chen et al., 2021) measured this on code generation: the temperature that maximised pass@1 was about 0.2, while for pass@100 it was about 0.8. With one attempt you want the most likely program; with a hundred attempts and a test suite that picks the winner, you want the attempts to be different from each other, and diversity comes from the tail.
If you evaluate this way, use the unbiased pass@k estimator from the same paper rather than literally drawing k samples. Generate n samples per problem, count c correct, and compute the probability that a random subset of k contains at least one correct sample:
import numpy as np
def pass_at_k(n: int, c: int, k: int) -> float:
"""Unbiased estimator, numerically stable product form (Chen et al. 2021)."""
if n - c < k:
return 1.0
return 1.0 - float(np.prod(1.0 - k / np.arange(n - c + 1, n + 1)))
# n = 200 samples, 10 correct
print(pass_at_k(200, 10, 1)) # 0.05
print(pass_at_k(200, 10, 10)) # about 0.409The same reasoning applies to self-consistency (sample several reasoning chains and vote) and to reinforcement-learning rollouts. Group-based policy-gradient methods compare several completions of the same prompt and normalise rewards within the group. If temperature is so low that every completion in a group is the same, every reward is the same, the normalised advantage is zero and the step produces no learning signal for that prompt. Rollout temperature is therefore a training hyperparameter, not a serving preference, and is usually kept near 1.
There is a second, quieter RL trap. If rollouts are sampled from softmax(z / T) but the training step computes log-probabilities from softmax(z), the policy whose gradient you take is not the policy you sampled from. Either sample at T = 1, or apply the same division in the training forward pass.
Fitted temperature: calibration
Temperature scaling, introduced for neural classifiers by Guo et al. (2017), uses the same division for a completely different purpose: making a model's confidence mean what it says. Modern networks tend to be overconfident; among predictions made with 90 percent confidence, fewer than 90 percent are right. Because dividing logits by T does not change the argmax, you can fit one scalar on a held-out set to minimise negative log-likelihood and improve calibration without touching accuracy.
For LLMs this applies whenever you turn logits into a decision with a threshold: a classifier built on the probability of the tokens Yes and No, a multiple-choice grader that reads the letter probabilities, or a router that escalates low-confidence answers to a human. Here T is fitted, not chosen, and it lives in your post-processing, not in the sampling call. The broader calibration toolbox, including Platt and isotonic methods and reliability diagrams, is covered in Model calibration architecture.
import torch, torch.nn.functional as F
def fit_temperature(logits: torch.Tensor, labels: torch.Tensor) -> float:
"""logits: [N, C] from a held-out set the model never trained on; labels: [N]."""
log_t = torch.zeros(1, requires_grad=True) # optimise log T so T stays positive
opt = torch.optim.LBFGS([log_t], lr=0.1, max_iter=200)
def closure():
opt.zero_grad()
loss = F.cross_entropy(logits / log_t.exp(), labels)
loss.backward()
return loss
opt.step(closure)
return float(log_t.exp())Measure expected calibration error (ECE), the confidence-weighted gap between stated confidence and observed accuracy across bins, before and after. A fitted T above 1 means the model was overconfident, below 1 underconfident. The fit is one parameter, so a few hundred labelled examples are usually enough, but it only holds for the distribution it was fitted on: a new prompt template, a new model version or a new user population can move it, so refit whenever any of those change.
Fixed temperature: distillation
In knowledge distillation (Hinton, Vinyals and Dean, 2015) a student learns from a teacher's full probability distribution rather than only from hard labels. At T = 1 a confident teacher puts nearly all mass on one token, and the information in the relative sizes of the small probabilities, which wrong answers are nearly right, is invisible. Raising T on both teacher and student softens both distributions so the loss can see that structure.
The detail people drop is the gradient scale. The gradient of the softened loss with respect to the logits shrinks roughly as 1 / T squared, so the paper multiplies the soft loss by T squared to keep its contribution comparable as T changes. Without that factor, raising T quietly turns the soft loss down and you end up tuning two coupled knobs at once.
import torch.nn.functional as F
def kd_loss(student_logits, teacher_logits, labels, T: float = 2.0, alpha: float = 0.5):
soft = F.kl_div(
F.log_softmax(student_logits / T, dim=-1),
F.log_softmax(teacher_logits / T, dim=-1),
reduction="batchmean",
log_target=True,
) * (T * T) # restore gradient scale
hard = F.cross_entropy(student_logits, labels) # T = 1 for the label term
return alpha * soft + (1 - alpha) * hardDistillation temperature is fixed for a run and must be identical for teacher and student; at evaluation and serving time the student runs at T = 1 with whatever sampling temperature the application chooses. The end-to-end recipe, including data curation and the choice between logit and sequence-level distillation, is in Model distillation architecture.
Learned temperature: contrastive embeddings
Embedding models used for retrieval in RAG systems are commonly trained with a contrastive loss: for each query, push its similarity to the matching document up and to the other documents in the batch down. Similarities are cosine values in [-1, 1], which is far too narrow a range for a softmax to express a confident choice, so they are divided by a temperature tau, typically well below 1. CLIP (Radford et al., 2021) made tau learnable, initialised it to the equivalent of 0.07 and clipped it so that logits are never scaled by more than 100.
import math, torch, torch.nn as nn, torch.nn.functional as F
class ContrastiveHead(nn.Module):
def __init__(self):
super().__init__()
self.logit_scale = nn.Parameter(torch.tensor(math.log(1 / 0.07))) # scale = 1 / tau
def forward(self, q: torch.Tensor, d: torch.Tensor) -> torch.Tensor:
q, d = F.normalize(q, dim=-1), F.normalize(d, dim=-1)
scale = self.logit_scale.exp().clamp(max=100.0)
logits = scale * q @ d.T # [B, B], positives on the diagonal
target = torch.arange(q.size(0), device=q.device)
return (F.cross_entropy(logits, target) + F.cross_entropy(logits.T, target)) / 2Small tau focuses the gradient on the hardest negatives, which sharpens retrieval but amplifies false negatives. Without the clamp, a learnable scale can grow until training becomes unstable, so log it.
Failure modes
- Comparing evaluations run at different temperatures. A model evaluated at T = 0.7 against one at T = 0 is a comparison of decoding settings, not of models. Record T, top-p and the seed with every result.
- Treating T = 0 as a cache key. Responses can differ at T = 0; cache on the response you got, not on the assumption that you would get it again.
- Zero-advantage RL groups. If the fraction of groups where all rewards are equal climbs, rollout temperature is too low or the task is too easy or too hard for the current policy.
- Calibration fitted on the wrong distribution. A T fitted on last quarter's traffic will be confidently wrong after a prompt change. Refit and re-measure ECE on every release.
- Distillation without the T squared factor, or with different temperatures on teacher and student.
- An unclamped learnable contrastive scale, visible as a scale that climbs steadily while retrieval metrics stall.
Choosing a sampling temperature
For the sampling job, treat temperature as a hyperparameter with a measurement behind it. Build a small eval set of real requests with a scoring function, which can be exact match, a unit test, a rubric or a judge model. Sweep T over 0, 0.3, 0.7 and 1.0 with everything else fixed, run several samples per prompt, and record both mean quality and the spread between samples. Pick the lowest temperature at which quality is still at its plateau if users want consistency, or the highest at which quality has not yet dropped if you rerank or vote over several samples.
Reasonable starting points, to be confirmed by that sweep, are low temperature for extraction, classification and single-shot code, and near 1 for brainstorming, best-of-n and RL rollouts.
What to do next
- Find every place your stack applies a temperature: sampling calls, eval harnesses, RL rollout config, calibration post-processing, distillation and embedding training. Write down who sets each value and why.
- Add T, top-p, top-k and the seed to every logged evaluation result, and refuse to compare runs whose settings differ.
- Run a four-point temperature sweep on one production task and keep the plot next to the config.
- If you threshold on model probabilities, fit a calibration temperature on held-out data with the code above, measure ECE before and after, and schedule a refit on each model or prompt release.
- If you train with distillation, check for the T squared factor and that teacher and student share T.
- If you train with RL, monitor the share of zero-advantage groups and verify that trainer and inference log-probabilities agree on the same tokens.
- If you train embeddings, log the contrastive scale and confirm it is clamped.