Self-attention compares every token with every other token, but nothing in the comparison says where those tokens are. Shuffle the input and, without positional information, each token's output is the same, only shuffled. "Dog bites man" and "man bites dog" would look alike. Positional encoding is how a transformer learns order, and the choice shapes how far a model can read, how its KV cache behaves and how hard it is to extend its context later.
This article builds the idea from first principles, compares the main designs with code, works through why rotary embeddings fail beyond their training length, explains the context-extension tricks built on that analysis, and ends with a decision table. Implementation detail for rotary embeddings, such as layouts and configuration files, is in RoPE architecture, in depth.
Why attention is blind to order
In one head, token i attends to token j with weight proportional to exp(qi · kj / √d), where q and k are linear projections of the token vectors. Permute the tokens and the set of dot products is permuted the same way, so the output is permuted too: attention is permutation-equivariant. A causal mask breaks the symmetry partly, because token 5 sees five predecessors and token 50 sees fifty, which is why some decoder-only models learn order with no encoding at all. Encoders, with no mask, need explicit position.
There are three places to add it, shown in the diagram: add a position vector to the token embedding before the first layer, rotate queries and keys inside every attention layer, or add a bias to the attention logits based on distance. The first gives absolute position; the other two give relative position, which is usually what language needs, since the relation between a verb and its subject does not depend on whether they appear at token 10 or token 10,000.
Absolute encodings: sinusoidal and learned
The original transformer added a fixed sinusoidal vector to each embedding. Dimension pair i oscillates with wavelength 2π · 100002i/d, so low dimensions change fast and high dimensions slowly, like the hands of a clock. A shift by k positions is a fixed linear map of the vector, which in principle lets attention learn relative offsets.
import torch
def sinusoidal(n_pos, d_model, base=10000.0):
pos = torch.arange(n_pos, dtype=torch.float32)[:, None] # (n, 1)
i = torch.arange(0, d_model, 2, dtype=torch.float32)[None, :] # (1, d/2)
angle = pos / base ** (i / d_model)
pe = torch.zeros(n_pos, d_model)
pe[:, 0::2] = torch.sin(angle)
pe[:, 1::2] = torch.cos(angle)
return pe # added to embeddingsThe derivation, including why a shift is a rotation, is worked through in positional encoding from first principles. Learned absolute embeddings, as in BERT and GPT-2, replace the formula with a trainable table of one vector per position. They work well within range but have a hard ceiling: GPT-2's table has 1,024 rows, and position 1,025 has no vector at all. Neither form extrapolates well in practice, and because position is mixed into the token vector once, deep layers must carry it forward through every residual update.
Relative biases: Shaw and T5
Shaw and colleagues (2018) added learned vectors for the clipped distance j − i into the attention computation, so attention depends on offset rather than absolute index. T5 simplified this to a learned scalar bias per head, added to the logit, indexed by a bucket of the distance. Small distances get their own buckets; larger ones share logarithmically spaced buckets up to a maximum distance, beyond which everything shares the last bucket.
import math, torch
def t5_bucket(distance, num_buckets=32, max_distance=128):
# causal case: distance = query_pos - key_pos >= 0
exact = num_buckets // 2
if distance < exact:
return distance
scaled = math.log(distance / exact) / math.log(max_distance / exact)
return min(exact + int(scaled * (num_buckets - exact)), num_buckets - 1)
def alibi_slopes(n_heads): # n_heads a power of two
return torch.tensor([2 ** (-8 * (h + 1) / n_heads) for h in range(n_heads)])
def alibi_bias(n, n_heads):
dist = torch.arange(n)[None, :] - torch.arange(n)[:, None] # j - i, <= 0 below diagonal
return alibi_slopes(n_heads)[:, None, None] * dist[None] # (h, n, n), add to logitsBucketed biases generalise to longer inputs because unseen distances fall into the last bucket. The cost is an extra n × n bias per head, which fused attention kernels must support, and the coarse buckets lose precise long-range positions.
Rotary embeddings in brief
RoPE (Su et al., 2021) splits each query and key into pairs of dimensions and rotates pair i by the angle m · θi, where m is the token's position and θi = base−2i/d. The dot product of two rotated vectors depends only on the difference of their angles, so qm · kn depends on m − n: absolute rotation produces relative attention, with no extra parameters and no bias matrix. Values are not rotated.
import torch
def rope_frequencies(head_dim, base=10000.0, scale=1.0, ntk_alpha=1.0):
# scale > 1: position interpolation (divide positions by scale)
# ntk_alpha > 1: NTK-aware scaling (raise the base instead)
base = base * ntk_alpha ** (head_dim / (head_dim - 2))
inv_freq = 1.0 / base ** (torch.arange(0, head_dim, 2).float() / head_dim)
return inv_freq / scale
def rotate(x, positions, inv_freq): # x: (..., n, head_dim), interleaved pair layout
angle = positions[:, None].float() * inv_freq[None, :] # (n, head_dim/2)
cos, sin = angle.cos(), angle.sin()
x1, x2 = x[..., ::2], x[..., 1::2]
out = torch.stack((x1 * cos - x2 * sin, x1 * sin + x2 * cos), dim=-1)
return out.flatten(-2) # applied to q and k, never to vBecause the rotation is applied to keys before they are cached, a KV cache stores already-rotated keys and decoding needs only the new token's position. RoPE is the default in most open decoder models today; Llama 2 used base 10,000 and Llama 3 raised it to 500,000 to support longer contexts.
ALiBi and NoPE
ALiBi (Press et al., 2021) drops learned position entirely and subtracts a penalty proportional to distance from each logit: head h uses slope mh, and for 8 heads the slopes are 1/2, 1/4, ... down to 1/256. Steep heads attend locally, shallow heads see far. The original paper trained on 1,024 tokens and evaluated longer with little loss in perplexity, the property it was designed for. The catch is that the linear penalty pushes attention toward recent tokens, which can hurt retrieval of a fact far back in a long document; see ALiBi attention architecture.
NoPE means no positional encoding at all in a causal decoder. Kazemnejad and colleagues (2023) showed such models learn position from the causal mask and can generalise to longer lengths in some tasks. It is an interesting baseline and is used in some layers of hybrid designs, but it is not a safe default for a general model without your own evaluation.
Worked example: why RoPE breaks past its training length
Take head dimension 128 and base 10,000. Pair 0 rotates by one radian per token, a wavelength of about 6.3 tokens. The last pair has θ = 10000−126/128, a wavelength of 2π × 10000126/128, about 54,000 tokens. A model trained on 4,096 tokens has seen every phase of the fast pairs many times, but its slowest pairs have turned only a fraction of a circle: about 4,096 / 54,000, under 8 percent of one rotation.
Run that model at 16,384 tokens and the slow pairs reach angles it never saw in training. Attention scores built from those dimensions become out of distribution, and perplexity typically climbs sharply beyond the training length. The fast pairs are fine because they wrapped around many times in training. That asymmetry is what every context-extension method works with.
Extending context: interpolation, NTK-aware scaling and YaRN
Position interpolation (Chen et al., 2023) divides positions by the extension factor, so 16,384 tokens map into the trained range 0 to 4,096. Slow pairs stay in distribution, but fast pairs are compressed fourfold and nearby tokens become harder to tell apart, so a short fine-tune at the new length is needed. In the code above this is scale=4.
NTK-aware scaling raises the base instead. High-frequency pairs barely change and keep local resolution, while low-frequency pairs stretch to cover the longer range. It often works without fine-tuning for moderate extension. In the code this is ntk_alpha.
YaRN (Peng et al., 2023) treats pairs separately: those whose wavelength is short relative to the training length are left alone, those whose wavelength exceeds it are interpolated, and a ramp blends between them. It also scales attention logits by a temperature, because longer sequences flatten the softmax. Llama 3.1 ships its own frequency-dependent scheme in the same spirit. Whatever method a checkpoint uses is part of the model: an inference engine that ignores the checkpoint's scaling settings produces fluent text that degrades at long range without any error. Measure the result with long-context retrieval tests rather than perplexity alone.
Systems consequences
- Kernels. RoPE leaves the attention kernel untouched because it rotates q and k before the call. Bias methods need kernel support for a bias or slopes; check that your fused kernel supports the method before choosing it. See FlashAttention.
- Sequence packing. When several documents are packed into one training row, reset position ids at each document boundary and mask across documents, or later documents train at positions they will never see alone.
- Padding. With left padding in batched generation, compute positions from the attention mask, not from the raw index, or every padded sequence is shifted.
- KV cache. Cached keys are rotated with the positions they were written at. Evicting or shifting cache entries, as streaming attention schemes do, must keep positions consistent with what the model saw in training.
Failure modes
| Symptom | Likely cause | Check |
|---|---|---|
| Garbage after a fixed length | Learned absolute table exhausted, or RoPE beyond training length | Model card context length; scaling config |
| Fine on short prompts, worse on long ones, no error | Inference ignores the checkpoint's rope scaling | Compare engine config with the checkpoint's |
| Batched outputs differ from single requests | Position ids counted over padding | Derive positions from the attention mask |
| Long-document recall poor despite low perplexity | Recency bias of ALiBi or over-compressed interpolation | Needle and multi-hop retrieval tests at target length |
| Training loss spikes on packed data | Positions not reset between documents | Inspect position ids of one packed row |
Choosing an encoding
| Method | Strengths | Weaknesses | Choose when |
|---|---|---|---|
| Learned absolute | Simple, strong in range | Hard length ceiling | Encoders with fixed short inputs |
| Sinusoidal | No parameters | Weak extrapolation | Teaching, small experiments |
| T5 buckets | Relative, graceful at length | Bias matrix, coarse far positions | Encoder-decoder models |
| RoPE | Relative, cache friendly, kernel neutral, extensible | Needs scaling to go past training length | Default for decoder LLMs |
| ALiBi | Extrapolates without tuning | Recency bias, kernel support needed | Long inputs where locality dominates |
| NoPE | Nothing to configure | Less predictable; needs evaluation | Research, hybrid layers |
What to do next
- For any checkpoint you serve, record its positional method, base, training length and scaling settings, and check that your inference engine reads all of them.
- Run the sinusoidal and RoPE code above and plot the wavelength of each pair for your model's head dimension and base.
- Test long-context behaviour at your target length with retrieval and multi-hop tasks, not perplexity alone.
- If you train with packing, verify position ids reset at document boundaries; if you batch with left padding, derive positions from the mask.
- Before extending context, try NTK-aware scaling or YaRN at inference, then decide whether a short fine-tune at the new length is worth it.
- Read RoPE architecture for implementation detail and attention sinks for how cache eviction interacts with position.