Self-attention compares every query with every key through a dot product, and a dot product has no idea where either vector came from. Without extra information, a transformer sees a sentence as a bag of tokens. Rotary Position Embeddings, introduced by Su and colleagues in the 2021 RoFormer paper, fix this by rotating each query and key by an angle proportional to its position. RoPE is now the default in most open decoder-only models, including the Llama, Mistral, Qwen and DeepSeek families.

Plenty of pages derive the rotation. This one follows RoPE into the inference stack, where it decides what a KV cache stores, whether cached keys can be moved, how keys can be quantized and which bugs show up only in production. It starts from first principles, checks the two key properties numerically with code you can run, and ends with a checklist for anyone serving or porting a RoPE model. For the wider family of position schemes, see positional embeddings. For stretching a model past its training length, see context extension.

Why rotate: position from first principles

Position information can enter a transformer in three places. It can be added to the token embedding (learned or sinusoidal absolute positions). It can be added to the attention score as a bias (T5 relative buckets, ALiBi). Or it can be applied to the query and key vectors themselves, which is what RoPE does. The goal is for the score between a query at position m and a key at position n to depend on the content of both vectors and on the distance m minus n, not on m and n separately. A model that has learned what 'three tokens back' means at position 10 should apply the same rule at position 1,000.

Rotation in a plane achieves exactly this. Rotate vector q by angle m times theta and vector k by angle n times theta. Their dot product equals the dot product of the unrotated vectors after rotating one of them by (m minus n) times theta, because rotations compose and preserve length. Absolute positions go in, and only the relative offset survives in the score. RoPE needs no learned parameters and adds nothing to the residual stream. Values are never rotated, so what a token contributes to the output is unaffected. Only how strongly it is attended to changes.

The rotation and the frequency ladder

A head of dimension d is split into d/2 pairs, and each pair is rotated at its own frequency. Pair i uses theta_i = base ** (-2i / d). The first pairs spin quickly, through a full turn every few tokens, and resolve fine local order. The last pairs turn so slowly that their angle barely changes across thousands of tokens, which lets them carry coarse long-range position. The base was 10,000 in the original paper and in Llama 2. Llama 3 raised it to 500,000 (rope_theta in the model config), which slows every pair and leaves more room for long contexts.

import numpy as np

def rope_tables(positions, head_dim, base=10000.0):
    inv_freq = base ** (-np.arange(0, head_dim, 2) / head_dim)   # (d/2,)
    angles = np.outer(np.asarray(positions, dtype=np.float64), inv_freq)
    return np.cos(angles), np.sin(angles)                         # (n, d/2)

def apply_rope(x, cos, sin):
    """Half-split layout: pair i is (x[i], x[i + d/2])."""
    h = x.shape[-1] // 2
    x1, x2 = x[..., :h], x[..., h:]
    return np.concatenate([x1 * cos - x2 * sin, x1 * sin + x2 * cos], axis=-1)

There are two pairing conventions. The original paper rotates adjacent elements (x0 with x1, x2 with x3), called the interleaved layout. Hugging Face Llama code and many others rotate element i with element i + d/2, called the half-split layout, as shown above. The two are equivalent up to a fixed permutation of the projection weights, but they are not interchangeable at runtime. Weights trained for one layout and run with the other produce a model that still emits fluent-looking text and is quietly wrong. RoPE in the transformer block covers the layouts, partial rotary and fused kernels in implementation detail.

Worked example: checking the properties numerically

Two claims carry most of the weight in what follows: the score depends only on relative position, and a key already rotated for position n can be moved to n + delta by rotating it by delta again. Both can be checked in a few lines with a 64-dimensional head and random vectors:

rng = np.random.default_rng(0)
d = 64
q, k = rng.standard_normal(d), rng.standard_normal(d)

def score(m, n):
    cq, sq = rope_tables([m], d)
    ck, sk = rope_tables([n], d)
    return float((apply_rope(q, cq, sq) * apply_rope(k, ck, sk)).sum())

print("score(10, 7)     =", round(score(10, 7), 6))
print("score(1010, 1007)=", round(score(1010, 1007), 6))

n, delta = 7, 500
cached = apply_rope(k, *rope_tables([n], d))
shifted = apply_rope(cached, *rope_tables([delta], d))
fresh = apply_rope(k, *rope_tables([n + delta], d))
print("max |shifted - fresh| =", float(np.abs(shifted - fresh).max()))

qs = rng.standard_normal((2000, d))
for dist in (0, 16, 256, 4096):
    s = np.sum(apply_rope(qs, *rope_tables([dist], d)) *
               apply_rope(qs, *rope_tables([0], d)), axis=-1)
    print(f"dist {dist:5d}: self-similarity {s.mean():.2f}")

Output:

score(10, 7)     = -10.512923
score(1010, 1007)= -10.512923
max |shifted - fresh| = 6.994405055138486e-14
dist     0: self-similarity 64.18
dist    16: self-similarity 38.80
dist   256: self-similarity 22.62
dist  4096: self-similarity -0.12

The first two lines confirm translation invariance: the same pair of vectors scores identically three tokens apart, whether at position 10 or at position 1,010. The third line shows that re-rotation matches recomputation to float64 rounding. The last block shows the behaviour the RoFormer paper calls long-term decay. A query matched against an identical key scores about d (64) at distance zero. The score falls as distance grows, because the fast pairs stop agreeing first and the slow pairs follow, until at 4,096 tokens the alignment has averaged out. A trained model learns to work with this. It is a tendency, not a hard window, and nothing stops a trained head from attending sharply to a distant token.

Where RoPE sits in the serving stack

One decode step with RoPE and a KV cachehidden state x_tposition tW_qW_kW_vrotate q by trotate k by tKV cacheK rotated, V plainv unrotatedscores = q_t . K / sqrt(d)depends on t - s for each cached key ssoftmax, weighted sum of VDesign choices this forcescache post-rotation keys: no work per stepor cache pre-rotation keys: rotate on readshift cached keys by rotating by deltaquantize keys before or after rotation
RoPE is applied after the Q and K projections and before the cache write. Values are never rotated. Each choice in the red box trades compute against flexibility.

In standard serving, keys are rotated once at their own position and written to the cache already rotated. Each decode step rotates only the new query and the new key, and the dot product with every cached key automatically reflects the distance. This is why RoPE costs almost nothing at inference: the cos and sin tables are precomputed or computed inside a fused kernel, and the cache never needs revisiting while positions keep increasing.

The cache layout itself, with pages, block tables and prefix sharing, is described in paged attention. One RoPE consequence of prefix sharing is worth stating: because cached keys carry their absolute rotation, a cached block is only valid at the position it was computed for. Prefix caches that key each block on the entire token sequence before it get this right automatically, since an identical prefix implies identical positions. A cache that tries to reuse the same document chunk at different offsets in different prompts cannot just copy the keys. It must re-rotate them, and even then the keys still encode the attention context they were computed in, which is a separate correctness problem.

Shifting cached keys: re-rotation and attention sinks

Re-rotation is what makes cache surgery possible. When a conversation outgrows the context window, a runtime can evict old tokens and slide the rest down. The evicted region leaves a gap in positions, and closing it means every remaining key must move from position n to n minus g. Recomputing those keys would need the original hidden states. Rotating each cached key by minus g gives the same answer, as the experiment above showed, at the cost of one elementwise pass over the cache.

StreamingLLM (Xiao and colleagues, 2023) builds on a related observation. Models put large attention weight on the first few tokens, which act as attention sinks, so a cache that keeps those few sink tokens plus a recent window stays stable over very long streams, while a plain sliding window collapses once the first tokens are evicted. For RoPE models the paper assigns positions by place in the cache, not by place in the original text. To do that it stores keys before rotation and applies the rotation at each decoding step. That is the other branch of the design choice: pre-rotation caching costs a rotation per cached key per step (cheap when fused into the attention kernel) and makes shifting free.

Two cautions apply. Repeated in-place shifts of post-rotation keys in fp16 or bf16 accumulate rounding error, so shift rarely, or keep a pre-rotation copy if you shift often. And neither technique extends what the model can attend to. Evicted tokens are gone. These are memory-management tools, not context extension.

Quantizing keys around the rotation

Key caches dominate long-context memory, so they get quantized. RoPE complicates this. Before rotation, a few key channels carry consistently large magnitudes, and these outlier channels are stable across tokens, which is what per-channel quantization exploits. Rotation mixes each channel with its pair at an angle that changes with position, so the same outlier energy is spread across the pair differently at every position. The KVQuant paper (Hooper and colleagues, 2024) responds by quantizing keys per channel and before RoPE, then applying the rotation after dequantization when attention is computed.

# pre-RoPE key quantization, conceptually
k_raw = x @ W_k                          # no rotation yet
k_q, scale = quantize_per_channel(k_raw) # outlier channels stay in fixed columns
cache.write(k_q, scale, position=t)

# at attention time, inside the kernel
k = dequantize(k_q, scale)
k = apply_rope(k, cos[positions], sin[positions])
scores = q_rot @ k.T

The trade-off is the same as for pre-rotation caching: rotation moves from the write path to the read path, so it must be fused into the attention kernel to stay cheap. When evaluating a KV-cache quantization scheme for a RoPE model, ask which side of the rotation it quantizes on, and measure long-context retrieval, not only perplexity on short text.

Bugs that only show up in production

  • Layout mismatch. Interleaved weights run with a half-split kernel, or the reverse. Short prompts look almost fine. Test against reference logits from the original implementation, not against your own expectations.
  • Wrong position ids under padding. With left padding, the first real token must still get position 0, or whatever the reference does. Position ids derived from the padded tensor shift every token. With packed sequences, positions must restart at each document boundary.
  • Low-precision angles. Computing position * inv_freq in bf16 rounds large positions, since bf16 represents integers exactly only up to 256, and angles at long range drift. Compute angles in fp32 and cast the results.
  • Config drift. rope_theta or rope_scaling lost while converting a checkpoint. The model works at short lengths and falls apart beyond the original training length. Extension schemes such as NTK-aware scaling and YaRN are covered in YaRN and NTK scaling.
  • Rotating V or the output. Only Q and K are rotated. Rotating V changes the outputs themselves, not only the attention pattern.
  • Speculative and chunked decoding. Draft tokens, verification and prefill chunks each need the correct absolute positions. An off-by-one here shows up as a slightly lower acceptance rate, not as a crash.

Trade-offs and variants

SchemeStrengthsWeaknesses
RoPErelative scores, no parameters, cache-friendly, strong defaultneeds scaling to go past the training length; layout and precision bugs
Learned absolutesimplehard limit at the table size; no relative structure
ALiBilinear distance bias, extrapolates gracefullyfixed decay shape; less common in current open models
No positional encoding (NoPE)nothing to configurerelies on the causal mask for order; less studied at scale

Variants keep the core idea. Some models rotate only part of each head (partial rotary). Qwen2-VL splits the rotary dimensions across time, height and width for images and video (M-RoPE). DeepSeek-V2 keeps a small separate rotary component alongside its compressed latent attention, because rotation would otherwise prevent absorbing the key projection. All of them rest on the same identity: rotating both vectors leaves only their difference in the score.

What to do next

  1. Run the numpy snippets above and change the base from 10,000 to 500,000 to see how the self-similarity curve stretches.
  2. For a model you serve, read rope_theta, rope_scaling and the maximum position from its config, and confirm your runtime reads the same values.
  3. Write a golden test: logits for a fixed 4,000-token prompt from the reference implementation, compared against your runtime at a tight tolerance.
  4. Check how position ids are built under left padding and packing in your batching code.
  5. Confirm angles are computed in fp32 in your kernels or framework version.
  6. If you evict or shift cache entries, decide between post-rotation keys with re-rotation and pre-rotation keys rotated on read, and measure the cost of each.
  7. If you quantize the KV cache, find out whether keys are quantized before or after RoPE, and benchmark long-context retrieval before rolling it out.
Key takeaway: RoPE rotates each query and key pair by position times a per-pair frequency, so attention scores depend only on relative distance and nothing is learned or added to the residual stream. Keys are usually cached already rotated, and because rotations compose, a cached key can be moved by rotating it by the offset. That makes cache shifting cheap, and pre-rotation caching and pre-RoPE quantization move the rotation to the read path instead. Most RoPE failures in production are layout mismatches, wrong position ids, low-precision angles or lost rope_theta and rope_scaling settings, and golden-logit tests catch all of them.