Rotary position embedding, RoPE, encodes a token's position by rotating pairs of channels in its query and key vectors by angles proportional to that position. The dot product between a rotated query at position t and a rotated key at position s then depends on t minus s, so attention sees relative distance without any extra parameters. The mathematics, including the frequency ladder and why it yields relative position, is derived in Rotary Position Embedding, the maths, and context extension with NTK scaling and YaRN is covered in YaRN and NTK scaling.
This article is about the architecture around the formula: where RoPE sits in an attention layer, the two incompatible memory layouts and how checkpoint converters reconcile them, partial rotary, how position IDs are built under padding and packing, why the KV cache stores rotated keys, what fused kernels change, how multi-head latent attention has to decouple it, and how to read a model's rope_scaling config. These are where RoPE bugs live: a wrong layout or offset gives a model that emits fluent text and is quietly worse.
Where rotation happens in the block
RoPE rotates queries and keys after projection, before attention scores. Values are never rotated: position should affect which tokens are attended to, not the content that is mixed. Nothing is added at the input embedding; every layer re-injects position into its own attention.
Rotating both sides is what makes it relative: the score depends on the difference of the two angles. In code it is two elementwise multiply-adds per tensor, negligible next to the projections.
A reference implementation
Most implementations have three parts: a cos/sin cache computed once, a helper that swaps and negates halves, and an apply function that indexes the cache by position IDs:
import torch
def rope_cache(head_dim, max_pos, base=10000.0, rotary_dim=None, device="cpu"):
"""cos/sin tables of shape [max_pos, rotary_dim], computed in float32."""
rd = rotary_dim or head_dim
inv_freq = 1.0 / (base ** (torch.arange(0, rd, 2, device=device, dtype=torch.float32) / rd))
pos = torch.arange(max_pos, device=device, dtype=torch.float32)
angles = torch.outer(pos, inv_freq) # [max_pos, rd/2]
emb = torch.cat([angles, angles], dim=-1) # half-split layout
return emb.cos(), emb.sin()
def rotate_half(x):
x1, x2 = x.chunk(2, dim=-1)
return torch.cat([-x2, x1], dim=-1)
def apply_rope(q, k, cos, sin, position_ids):
"""q, k: [batch, heads, seq, head_dim]; position_ids: [batch, seq]."""
rd = cos.shape[-1]
cos = cos[position_ids].unsqueeze(1).to(q.dtype) # [batch, 1, seq, rd]
sin = sin[position_ids].unsqueeze(1).to(q.dtype)
def rot(x):
x_rot, x_pass = x[..., :rd], x[..., rd:] # partial rotary if rd < head_dim
return torch.cat([x_rot * cos + rotate_half(x_rot) * sin, x_pass], dim=-1)
return rot(q), rot(k)Two details are deliberate. The angles are computed in float32 and only cast to the activation dtype at the end. In bfloat16, with 8 bits of significand precision, integers above 256 are no longer all representable: positions 4,001 and 4,007 both round to 4,000, and the rotation angles of neighbouring tokens become identical. Second, the cache is indexed by explicit position_ids rather than by the token's index in the tensor, which is what makes padding, packing and caching work.
Two layouts, one maths: interleaved versus half-split
RoPE rotates pairs of channels, and there are two conventions for which channels form a pair. The interleaved layout, used by the original RoFormer and GPT-J code and by Meta's original Llama release, pairs adjacent channels (0, 1), (2, 3) and so on. The half-split layout, used by GPT-NeoX and by the Hugging Face Llama implementation, pairs channel i with channel i plus d/2, which is what rotate_half expresses.
They are the same operation applied to a permuted vector. Because the permutation is the same for queries and keys, dot products are unchanged, so a model trained with one layout can be run with the other if and only if the rows of W_q and W_k are permuted to match.
import torch
def rope_interleaved(x, cos_half, sin_half):
"""GPT-J / original Meta layout: pairs are (x0, x1), (x2, x3), ..."""
x_even, x_odd = x[..., 0::2], x[..., 1::2]
out_even = x_even * cos_half - x_odd * sin_half
out_odd = x_even * sin_half + x_odd * cos_half
return torch.stack([out_even, out_odd], dim=-1).flatten(-2)
def rope_half_split(x, cos_half, sin_half):
"""GPT-NeoX / Hugging Face Llama layout: pairs are (x_i, x_{i + d/2})."""
x1, x2 = x.chunk(2, dim=-1)
return torch.cat([x1 * cos_half - x2 * sin_half, x1 * sin_half + x2 * cos_half], dim=-1)
def to_half_split(x):
"""Reorder the last dim from interleaved to half-split: evens first, then odds."""
return torch.cat([x[..., 0::2], x[..., 1::2]], dim=-1)
d, t = 8, 5
inv = 1.0 / (10000 ** (torch.arange(0, d, 2).float() / d))
cos_h, sin_h = torch.cos(t * inv), torch.sin(t * inv)
q = torch.randn(d)
a = to_half_split(rope_interleaved(q, cos_h, sin_h))
b = rope_half_split(to_half_split(q), cos_h, sin_h)
assert torch.allclose(a, b, atol=1e-6) # same maths once the channels are reorderedThis is exactly what the Llama conversion script in transformers does to Meta's weights. Its permute function reorders the output rows of the query and key projections, per head, from interleaved pairs to the half-split order:
# From transformers' convert_llama_weights_to_hf.py: reorders the OUTPUT rows of
# W_q and W_k, per head, from interleaved pairs to the half-split layout.
def permute(w, n_heads, dim1, dim2):
return w.view(n_heads, dim1 // n_heads // 2, 2, dim2).transpose(1, 2).reshape(dim1, dim2)This is the most common RoPE bug in ports and custom kernels: skipping the permute, applying it twice, or using a kernel with the other layout scrambles position while content still flows, so outputs look plausible and perplexity is badly off. Grouped-query attention adds a trap: the key projection has fewer heads than the query projection, so the permute must use the number of key-value heads for W_k.
Partial rotary and the choice of base
Nothing requires every channel to be rotated. GPT-NeoX rotates a fraction of each head's channels, set by its rotary_pct configuration of 0.25, and leaves the rest position-free. The idea is that some channels can then match content regardless of distance. In the reference code, rotary_dim below head_dim handles it: only the first slice is rotated.
The base, often rope_theta in configs, sets the frequency ladder. The original value of 10,000 gives the slowest channel a wavelength of roughly 54,000 positions for a 128-dimensional head, which is comfortably above a 4K or 8K training context but well short of 128K. Newer long-context models raise it; Llama 3 uses 500,000. Changing the base after training changes every angle, so never edit it casually when serving.
Position IDs under padding and packing
The rotation a token receives comes entirely from the position ID you pass, so the code that builds position IDs is part of the model.
Batched generation usually pads on the left so that every sequence's last real token is at the same index. If position IDs are simply 0 to length minus 1 across the padded tensor, a short prompt's first real token starts at position 2 or 20 instead of 0, and the model sees a different context from the one it was tested on. The fix derives positions from the attention mask.
Sequence packing for training puts several documents in one row to avoid padding. Positions must restart at 0 for each document, and attention must also be blocked across document boundaries, either with a block-diagonal mask or with variable-length attention kernels that take cumulative sequence lengths. Restarting positions without the mask, or the reverse, leaks one document into another.
import torch
# Left-padded batch for generation: 0 marks padding.
attention_mask = torch.tensor([[0, 0, 1, 1, 1],
[1, 1, 1, 1, 1]])
position_ids = attention_mask.long().cumsum(-1) - 1
position_ids.masked_fill_(attention_mask == 0, 1) # any valid index; these are masked anyway
# tensor([[1, 1, 0, 1, 2],
# [0, 1, 2, 3, 4]])
# Packed training sequence: three documents of lengths 3, 2, 4 in one row.
lengths = [3, 2, 4]
packed_position_ids = torch.cat([torch.arange(n) for n in lengths])
# tensor([0, 1, 2, 0, 1, 0, 1, 2, 3]) -- restart per document, and the attention
# mask (or varlen kernel boundaries) must also stop documents seeing each other.
RoPE and the KV cache
During decoding, each new token's key is rotated by its own position and then appended to the cache. Cached keys are never re-rotated, because the relative property means that the score between the new query at t and a cached key at s already depends only on t minus s. This is what makes RoPE cache-friendly: a decode step costs one rotation of one query and one key. The cache design itself is described in KV cache architecture.
Storing rotated keys has consequences. Anything that changes positions after the fact, such as dynamic scaling that alters frequencies as the sequence grows, makes old cached keys inconsistent with new ones; implementations either accept that mismatch or recompute. Streaming schemes that evict the middle of the cache must choose whether to keep original positions or to assign positions within the cache and re-rotate, a question attention sinks works through.
Fused kernels
In eager PyTorch, RoPE is several small elementwise kernels per layer, each reading and writing the whole query and key tensor. Production stacks fuse rotation into one kernel, into the kernel that writes keys into the paged cache, or into the attention prologue. The flash-attn package ships a rotary kernel with an explicit interleaved flag, which is a reminder that every fused implementation embeds one of the two layouts. Verify any swapped kernel against the reference at small and large positions. Attention kernels themselves are covered in FlashAttention.
Decoupled RoPE in multi-head latent attention
Multi-head latent attention compresses keys and values into a small latent vector per token and, at inference, absorbs the key up-projection into the query projection so that attention runs directly against the cached latent. Rotation breaks that absorption: a position-dependent rotation between the up-projection and the dot product means the matrices can no longer be pre-multiplied. DeepSeek-V2's solution is to decouple position. Each head's query and key get a small extra slice that carries RoPE, with a key slice shared across heads and cached alongside the latent, while the larger compressed part carries no position at all. The score is the sum of a content term and a rotary term; see multi-head latent attention.
Reading a rope_scaling config
Hugging Face configs describe RoPE with a base and an optional scaling block, and the rope_type key selects the formula. Recent transformers releases accept values including linear, dynamic, yarn, longrope and llama3. Newer transformers releases fold these into a rope_parameters dict, but checkpoint configs such as Llama 3.1's still read:
{
"head_dim": 128,
"max_position_embeddings": 131072,
"rope_theta": 500000.0,
"rope_scaling": {
"rope_type": "llama3",
"factor": 8.0,
"low_freq_factor": 1.0,
"high_freq_factor": 4.0,
"original_max_position_embeddings": 8192
}
}Read it as follows: the model was pretrained with 8,192 positions, and at load time its inverse frequencies are rescaled by band. High-frequency channels, whose wavelengths are short relative to the original context, are left alone; low-frequency channels are divided by the factor of 8; the band in between is interpolated smoothly using the two frequency factors. The rescaled frequencies go into the cos/sin cache once. An engine that ignores this block serves a 128K model with wrong long-range positions, while short-prompt tests look fine.
Worked example: diagnosing a ported checkpoint
A team ports a Llama-style model to a custom inference engine. Short prompts produce fluent text, but perplexity on a validation set is 40 percent worse than the reference implementation, and answers to questions about early parts of long documents are poor. They compare the reference and their engine layer by layer on the same input. Embeddings and the first projection match exactly; queries after rotation differ at every position except 0. At position 0 the rotation is the identity, so any layout mismatch is invisible there, a useful clue. Their kernel uses the interleaved layout while the checkpoint was converted to half-split. Applying the inverse permute to W_q and W_k, with the key-value head count for W_k, brings every layer within floating-point tolerance and perplexity back in line. They add a regression test at positions 0, 1, 1,000 and 100,000.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Fluent output, worse perplexity | Interleaved and half-split layouts mixed | Permute W_q and W_k or switch the kernel layout |
| Bug only at long context | Positions or inverse frequencies computed in bf16 or fp16 | Compute angles in float32, cast at the end |
| Batched output differs from single prompts | Position IDs ignore left padding | Derive positions from the attention mask |
| Packed training loss is odd | Positions not reset per document, or no cross-document mask | Reset positions and block-diagonal mask together |
| Long prompts degrade after a config edit | rope_theta or rope_scaling changed or ignored | Serve exactly the trained values; test beyond the original context |
| Port fails only with grouped-query attention | Permute used the query head count for W_k | Use the key-value head count |
What to do next
- Find your model's layout, base, rotary dimension and rope_scaling block, and write them down next to the checkpoint.
- Keep a reference implementation and compare rotated queries and keys at positions 0, 1, a mid value and a value beyond the pretraining context.
- Compute RoPE angles in float32 in every code path, including fused kernels.
- Build position IDs from the attention mask for left-padded batches and reset them per document when packing.
- Confirm that your serving engine applies the rope_scaling block and that cached keys are rotated exactly once.
- When porting weights, apply the layout permute with the correct head counts and test perplexity against the source.
- Read the maths and extension articles before changing the base or scaling for longer context.