Rotary position embeddings are the default way modern decoder-only language models encode token order. Llama, Qwen, Mistral, Gemma and most open models rotate their query and key vectors by an angle proportional to position before attention. The idea is compact enough to state in two lines, which makes it easy to implement and easy to get subtly wrong.

This article assumes you have seen the rotation once and goes further. It treats the rotary channels as a bank of clocks and puts numbers on their periods, summarises what interpretability work has found trained models doing with the fast and slow clocks, gives a property-test harness that catches most implementation bugs, works through a precision bug that only appears in long sequences, and explains how the base is chosen for pretraining and how RoPE extends to images and video. The derivation with a numeric example is in RoPE: relative position via rotation, and memory layouts, kernels and the KV cache are in RoPE architecture.

Advertisement

The rotation in one paragraph

Split each query and key head vector of dimension d into d/2 pairs of channels. Pair i gets a frequency θi = base−2i/d, with base 10,000 in the original RoFormer paper. The pair at token position m is rotated by the angle mθi. Because rotations compose, the dot product between a query at position m and a key at position n depends on the positions only through m − n. Absolute positions go in and relative positions come out, with no parameters and no change to the values. The rest of this article is about the consequences of that ladder of frequencies.

The frequency ladder as a bank of clocks

Each pair is a clock hand that turns θi radians per token, so it completes a turn every 2π/θi = 2π·base2i/d tokens. Pair 0 turns once per 6.3 tokens regardless of base; it distinguishes neighbours sharply but repeats constantly. For a 128-dimensional head, the slowest pair has a wavelength of about 54,410 tokens at base 10,000 and about 2,559,196 tokens at base 500,000. Between them the wavelengths grow geometrically.

The useful question is how many pairs never complete even one turn inside the context the model sees. Those pairs cannot alias, since no two positions in context share an angle, but each sees only part of its circle during training. The figure and table below are computed by this batch script from the formula above.

Wavelength of each rotary channel pair (head_dim 128), log scale1e01e11e21e31e41e51e61e76pair 0, base 10,0006pair 0, base 500,00063pair 16, base 10,000167pair 16, base 500,000628pair 32, base 10,0004,443pair 32, base 500,0006,283pair 48, base 10,000118,143pair 48, base 500,00019,869pair 56, base 10,000609,226pair 56, base 500,00054,410pair 63, base 10,0002,559,196pair 63, base 500,0008K context128K contextBars to the right of a context line never complete one turn inside that context.Blue: base 10,000. Violet: base 500,000. Values computed by this script.
Wavelength per channel pair for two bases. The fastest pairs repeat every few tokens; the slowest barely move within an 8K context.
ContextPairs slower than context, base 10,000Pairs slower than context, base 500,000
8,192 tokens14 of 6429 of 64
131,072 tokens0 of 6415 of 64

Raising the base shifts the ladder toward slower clocks, so more pairs act almost position-free within a context. The pairs that never complete a turn in training are the ones that break past the training length, because they meet unseen angles; extending RoPE to longer contexts covers the fixes.

Advertisement

What trained models do with the fast and slow clocks

The RoFormer paper motivated RoPE partly by a long-term decay property: an upper bound on the attention contribution that shrinks as relative distance grows. A bound is not a behaviour. For particular query and key vectors, the rotated dot product can stay large at long distances, and a model can learn such vectors.

Barbero and colleagues studied what a trained model actually does in Round and Round We Go! What makes Rotary Positional Encodings useful?, published at ICLR 2025, by examining Gemma 7B. They report that the model uses the highest frequencies to build robust positional attention patterns, heads that attend to the same or the previous token for instance, and that it puts much of its query and key mass in the lowest frequencies, which they argue carry semantic information because those channels barely rotate across the context. They argue that distance decay is unlikely to be the core reason RoPE works, and propose a variant, p-RoPE, that removes a fraction of the lowest frequencies so those channels become fully position-free, reporting improved performance in their experiments.

GPT-NeoX already rotated only a fraction of each head's channels, so the idea is not new. Treat these as findings about particular models, not laws. The practical point holds: channel bands do different jobs, so scaling or truncating one band changes behaviour unevenly.

A property-test harness for any implementation

RoPE bugs rarely crash: a wrong layout, sign or position offset yields a model that generates plausible text while being slightly worse. Test properties any correct implementation must satisfy. The tests below take a function rope(x, positions) that rotates a tensor of shape (batch, heads, seq, head_dim).

import torch

def make_rope(head_dim, base=10_000.0):
    inv_freq = base ** (-torch.arange(0, head_dim, 2, dtype=torch.float32) / head_dim)
    def rope(x, pos):                                   # half-split layout
        ang = pos.to(torch.float32)[:, None] * inv_freq[None, :]      # angles in fp32
        cos, sin = ang.cos(), ang.sin()
        x1, x2 = x.float().chunk(2, dim=-1)
        out = torch.cat([x1 * cos - x2 * sin, x1 * sin + x2 * cos], dim=-1)
        return out.to(x.dtype)
    return rope

# Tolerances: (vector, score). On scores of magnitude ~40, fp32 angle rounding at position
# 120,000 moves them by ~0.05 and bf16 rounding by ~0.15; a real position bug moves them by tens.
TOL = {torch.float32: (1e-4, 0.1), torch.bfloat16: (5e-2, 0.5)}

def check_rope(rope, head_dim=128, seq=64, offset=1000, dtype=torch.float32):
    vec_tol, score_tol = TOL[dtype]
    torch.manual_seed(0)
    q = torch.randn(1, 1, seq, head_dim).to(dtype)
    k = torch.randn(1, 1, seq, head_dim).to(dtype)
    pos = torch.arange(seq)
    f = lambda t: t.float()

    # 1. Position 0 is the identity.
    z = torch.zeros(1, dtype=torch.long)
    assert torch.allclose(f(rope(q[..., :1, :], z)), f(q[..., :1, :]), atol=vec_tol)

    # 2. Rotation preserves vector norms, also far out.
    assert torch.allclose(f(rope(q, pos + offset)).norm(dim=-1), f(q).norm(dim=-1), atol=vec_tol)

    # 3. Scores depend only on relative position: shifting every position changes nothing.
    s0 = f(rope(q, pos)) @ f(rope(k, pos)).transpose(-1, -2)
    s1 = f(rope(q, pos + offset)) @ f(rope(k, pos + offset)).transpose(-1, -2)
    assert torch.allclose(s0, s1, atol=score_tol), (s0 - s1).abs().max()

    # 4. Incremental decoding matches the full sequence (catches cache position bugs).
    full = f(rope(k, pos + offset))
    step = torch.cat([f(rope(k[..., t:t+1, :], pos[t:t+1] + offset)) for t in range(seq)], dim=-2)
    assert torch.allclose(full, step, atol=vec_tol)

rope = make_rope(128)
for dtype in (torch.float32, torch.bfloat16):
    for offset in (1_000, 5_000, 120_000):          # up to your longest served position
        check_rope(rope, offset=offset, dtype=dtype)

For anything ported, add a fifth test comparing outputs with the reference implementation on identical inputs; a disagreement by a channel permutation means a layout mismatch. The shift test is the most valuable: it catches wrong signs, wrong frequency formulas and mismatched query and key positions. Note that the harness computes angles in float32; the next section shows why.

Worked example: the bug that appears after a few thousand tokens

A team ports a model to a new inference engine and runs everything in bfloat16 to save memory. Short-prompt evaluations match the reference. Long-document question answering degrades sharply once prompts pass a few thousand tokens. The port's unit tests, which compared against the reference on 64-token inputs at positions 0 to 63, all pass.

The cause is the angle computation. bfloat16 has 8 bits of significand precision, so between 4,096 and 8,192 representable numbers are 32 apart. The engine built the position tensor in bfloat16, so position 5,000 was stored as 4,992. For pair 0, whose frequency is 1 radian per token, that is an error of 8 radians, a completely wrong angle; for slow pairs the relative error is the same and the absolute error smaller. Tokens near each other in a long prompt collapsed onto identical rounded positions and became positionally indistinguishable to the fast channels that positional heads rely on.

Integers up to 256 are exact in bfloat16, which is why the short tests passed. Running the harness above in bfloat16 with an offset of 5,000 separates the cases cleanly: in a simulation with this seed, a correct implementation's scores moved by about 0.16 from rounding alone, while the version with bfloat16 positions moved by more than 30. The fix is to compute positions and angles in float32, or precompute cos and sin tables in float32, and cast only the rotated activations. The lesson for the harness: run it at the longest positions you will serve and in the dtype you will serve in.

Choosing the base when you pretrain

If you are training a model rather than serving one, the base is a hyperparameter you pick before seeing the cost of picking it wrong. The constraint from the clock picture is that the slowest pairs should not wrap around within the longest context the model will be trained on, and ideally the base should leave headroom for later extension.

Two recipes are common. The first is adjusted base frequency (ABF), described by Xiong and colleagues in Effective Long-Context Scaling of Foundation Models (2023): pretrain at a short context with the original base, then continue pretraining on long sequences with the base raised, from 10,000 to 500,000 in their Llama 2 Long work, so the model learns the new, slower ladder. The second is to start with a large base from the beginning; Llama 3 uses 500,000. Either way, the base is part of the model: a checkpoint is only valid with the base, and any scaling configuration, it was trained with.

# A staged long-context schedule (illustrative values, not a published recipe)
stages = [
    dict(name="pretrain",     seq_len=8_192,   rope_theta=500_000,   tokens="most of budget"),
    dict(name="long-context", seq_len=131_072, rope_theta=500_000,   tokens="small fraction",
         data="long documents mixed with short ones so short-context quality holds"),
]
# Variant (ABF): pretrain with rope_theta=10_000, then raise it for the long-context stage.
# After any stage that changes rope_theta, rerun the property tests and long-context evals.

Packed sequences need position ids that restart at each document (see RoPE architecture). Evaluate long context with retrieval across the full window, not only perplexity, which can look healthy while mid-window recall is poor.

Beyond one dimension: images, video and M-RoPE

Images have two axes and video three. Axial RoPE splits each head's channel pairs into groups rotated by the row index and by the column index, so relative position becomes a 2-D offset between patches.

Qwen2-VL's multimodal RoPE (M-RoPE) applies this inside a language model. The rotary embedding is decomposed into temporal, height and width components, and each token carries three position ids. For text the three ids are identical, so M-RoPE reduces to ordinary 1-D RoPE. For an image the temporal id is constant and height and width ids follow the patch's row and column. For video the temporal id advances with each frame. The sketch below builds the ids for a text-image-text sequence; the exact offsets in a real model come from its own preprocessing code.

import torch

def mrope_ids(n_text_before, grid_h, grid_w, n_text_after):
    """Return a (3, seq) tensor of (temporal, height, width) ids. Illustrative."""
    ids = []
    for p in range(n_text_before):                 # text: all three ids equal
        ids.append((p, p, p))
    start = n_text_before
    for r in range(grid_h):                        # image: constant t, row and column ids
        for c in range(grid_w):
            ids.append((start, start + r, start + c))
    nxt = max(max(t) for t in ids) + 1             # resume text after the largest id used
    for p in range(n_text_after):
        ids.append((nxt + p,) * 3)
    return torch.tensor(ids).T

print(mrope_ids(3, 2, 3, 2))

The design consequence is that position ids are no longer a simple arange: they depend on the input's structure, so every component that touches them, including packing, caching and speculative decoding, must carry three ids per token.

Failure modes

  • Angles in low precision. Positions or angles held in bfloat16 or float16 corrupt long sequences. Compute in float32.
  • Base mismatch. Serving with a different rope_theta than training changes every angle. Read it from the checkpoint's config; never hard-code it.
  • Layout mismatch. Interleaved pairs against half-split pairs after a port. The reference-comparison test catches it.
  • Position ids out of sync with the cache. Decoding step t must use position t, including after cache eviction or prompt reuse. See the KV cache.
  • Running past the trained length. Slow pairs meet unseen angles and quality falls off a cliff. Use a scaling method the model was tuned with, or stay inside the window.
  • Uneven band changes. Truncating or scaling frequencies affects positional and semantic channels differently. Re-run long and short evaluations after any change.

Trade-offs

ChoiceStrengthWeakness
RoPE, small baseSharp local position signal, proven at short contextSlow pairs wrap or go unseen at long context
RoPE, large baseMore headroom for long contextCoarser position signal in mid-band channels
Partial rotary or p-RoPESome channels fully position-free for content matchingFewer channels carry position; another hyperparameter
ALiBiSimple linear bias, extrapolates gracefullyFixed recency bias; less common in recent large models

What to do next

  1. Add the four property tests to continuous integration for every attention implementation you own, run at your longest served position and in your serving dtype.
  2. Add a reference-comparison test for every ported checkpoint.
  3. Audit where positions and angles are computed and make sure it is float32.
  4. Record rope_theta and any scaling configuration as part of each checkpoint's identity, and fail loading if they are missing.
  5. If you pretrain, decide the long-context plan up front: a large base from the start, or ABF in a later stage, with long-context retrieval evaluations.
  6. For multimodal models, check that packing, caching and decoding all carry the full multi-axis position ids.
Key takeaway: RoPE turns position into a bank of clocks, one per channel pair, whose periods range from a few tokens to far beyond the context. Trained models use the fast clocks for positional patterns and the slow ones much like position-free channels, which is why the base matters and why changes to it act unevenly. Test implementations by their properties at long positions in the serving dtype, compute angles in float32, treat the base as part of the checkpoint, and carry multi-axis ids through every component when you extend RoPE to images and video.