For most of the transformer's history, the attention question had one answer: multi-head softmax attention, later with fewer key-value heads. By 2026 that answer has split. Open models released through late 2025, which form the baseline this year builds on, ship at least five distinct designs: grouped-query attention, multi-head latent attention, local-global interleaving, indexer-driven sparse attention, and hybrids that replace three of every four attention layers with a gated linear recurrence. One lab even moved back to full attention after shipping a linear variant.
This article is a field guide to that landscape. The maths of each mechanism is covered elsewhere on this site; here the focus is what each design costs at serving and training time, what goes wrong in production, and how to choose. It includes a fleet-sizing worked example at 128K context, code for the two newest mechanisms, and a checklist.
Why attention keeps changing
Two costs drive every variant. During prefill, attention compute grows with the square of the sequence length. During decoding, every generated token must read the key-value (KV) cache for all previous tokens in every layer, so memory capacity limits how many requests fit on a GPU and memory bandwidth limits how fast each one decodes. At short contexts the MLP layers dominate and the attention design barely matters. At 128K tokens and beyond, with agents and long documents now routine, attention memory and bandwidth set the price of every request.
The variants attack different terms. Sharing or compressing KV heads shrinks bytes per token. Windows cap the number of tokens a layer remembers. Sparse selection cuts the tokens each query reads without cutting what is stored. Linear recurrences replace the growing cache with a fixed-size state. The per-mechanism maths is in the attention variants maths article; the rest of this page is about consequences.
The 2026 map: what shipped where
| Family | Shipped in (examples) | What it saves | What it costs |
|---|---|---|---|
| GQA | the default in most dense open models | KV bytes by the group factor | little; mild quality loss versus full MHA |
| MLA | DeepSeek-V2/V3 family, Kimi Linear's full layers | KV stored as a small latent | more complex kernels; RoPE handled separately |
| Local:global interleave | Gemma 3 (5:1, 1,024-token windows); gpt-oss (alternating dense and banded layers) | KV capped in local layers | long-range recall rests on few global layers |
| Indexer sparse | DeepSeek-V3.2-Exp: lightning indexer picks 2,048 tokens per query | attention compute and bandwidth at long context | full KV still stored; extra indexer pass |
| Gated linear hybrid | Qwen3-Next and Kimi Linear, both 3:1 linear-to-full | most layers keep a constant-size state | new kernels, cache handling and failure modes |
| Full attention, by choice | MiniMax-M2, after M1's lightning attention | nothing; it is the baseline | pays full cost for predictability |
Read the last row carefully. MiniMax's team said linear attention proved tricky in production models, and returned to full attention for M2. Efficient attention is not a free lunch, and the frontier is genuinely unsettled. Treat any single benchmark showing parity as a starting hypothesis, not a conclusion.
Indexer-based sparse attention
DeepSeek Sparse Attention (DSA) splits attention into two passes. A lightning indexer with a few small heads, cheap enough to run in FP8, scores every earlier token against the current query. A top-k step keeps the best 2,048 tokens, and the main attention runs only over those. The indexer is still quadratic in principle but much cheaper per pair, so total cost drops sharply at long context. Detailed patterns for sparse attention are in the sparse attention article. In outline:
# Per query position t, decoding. Shapes: idx_q [H_i, d_i], idx_k [T, d_i], w [H_i]
def indexer_scores(idx_q, idx_k, w):
# cheap scores: per indexer head, ReLU(q . k), weighted and summed over heads
s = torch.relu(idx_q @ idx_k.T) # [H_i, T]
return (w[:, None] * s).sum(0) # [T]
def dsa_step(q, K, V, idx_q, idx_k, w, k_top=2048):
T = K.shape[0]
if T <= k_top:
sel = torch.arange(T)
else:
sel = indexer_scores(idx_q, idx_k, w).topk(k_top).indices
# main attention over the selected rows only
att = torch.softmax((q @ K[sel].T) / K.shape[-1] ** 0.5, dim=-1)
return att @ V[sel]The operational point is in the code: K and V for all T tokens must still be resident, because any of them may be selected next step. Sparse selection lowers compute and the bytes read per step, not the bytes stored, so it does not raise how many concurrent long requests fit in memory. It also needs the indexer trained to agree with the dense attention it replaces; DeepSeek trained the indexer to match dense attention before switching to sparse training. A poorly aligned indexer silently drops the one token that mattered.
Gated linear hybrids and output gating
Qwen3-Next and Kimi Linear both use Gated DeltaNet-style layers for three of every four blocks, with full attention (gated attention in Qwen3-Next, MLA in Kimi Linear) in the fourth. Each linear layer keeps a fixed matrix state per head instead of a cache. The delta rule treats that state as an associative memory: before writing a new key-value pair it removes what the memory currently returns for that key, and a decay gate forgets old content. With state S of shape [d_v, d_k], decay alpha and write strength beta in (0, 1):
# Gated delta rule, one head, recurrent form (training uses a chunked parallel form)
def gated_delta(q, k, v, alpha, beta): # q,k: [T, d_k], v: [T, d_v]
S = torch.zeros(v.shape[1], k.shape[1])
out = []
for t in range(q.shape[0]):
kt = k[t] / k[t].norm() # keys are L2-normalised
S = alpha[t] * S # forget
S = S - beta[t] * torch.outer(S @ kt, kt) # erase old value at kt
S = S + beta[t] * torch.outer(v[t], kt) # write new value
out.append(S @ q[t])
return torch.stack(out)Kimi Delta Attention refines the gate from one scalar per head to one value per channel. Kimi reports up to 75% less KV cache and up to six times higher decoding throughput at very long contexts; treat these as the vendor's figures for its own setup. Why keep any full layers? A fixed-size state is lossy: exact retrieval of an arbitrary earlier token, such as a phone number from page 3 of a long contract, is precisely what a compressed memory does badly, and the periodic full layer restores it. The trade-offs of mixing the two are covered in hybrid architectures.
Output gating is the quieter 2025 change. Qwen's gated attention multiplies each head's attention output by a learned sigmoid gate computed from the input before the output projection. The gate lets a head emit nearly nothing when it has nothing useful to say, reducing the pressure that otherwise creates attention sinks and huge activations, and improving training stability. It is cheap and orthogonal to everything else here. Background on why sinks form is in attention sinks.
Worked example: memory per request at 128K
Take a 48-layer model with model width 4,096, 32 query heads of dimension 128 and 8 KV heads, serving 128K-token contexts in BF16 (2 bytes per value). Compare per-request memory for four stacks:
L, T, H_kv, d, B = 48, 131_072, 8, 128, 2
kv_layer = 2 * H_kv * d * B * T # K and V for one layer: 512 MiB
dense = L * kv_layer # 48 full GQA layers: 24.0 GiB
local = (L // 6) * kv_layer + (L - L // 6) * 2 * H_kv * d * B * 1024
# 8 global + 40 window(1,024): ~4.2 GiB
# hybrid 3:1: 12 full layers + 36 linear layers, 32 heads x 128 x 128 state each
state = 32 * 128 * 128 * 4 # fp32 state per linear layer: 2 MiB
hybrid = 12 * kv_layer + 36 * state # ~6.1 GiB, flat in T for the linear part
sparse = dense # DSA-style top-k: still 24 GiB storedPer request, dense GQA needs 24 GiB, so an 80 GB accelerator holding 40 GB of weights serves one such request with little room to spare. The 5:1 local-global stack needs about 4.2 GiB and the 3:1 hybrid about 6.1 GiB, so the same card serves six to eight. Top-k sparse stores the full 24 GiB but reads far less per step, so it decodes each request faster without fitting more of them. Kimi Linear's full layers use MLA, which compresses those 12 layers further. These are capacity numbers, not quality numbers: whether the 4-to-6 GiB stacks answer 128K-context questions as well as the 24 GiB one must be measured on your tasks.
Bandwidth tells the same story per step. Each decoded token reads the stored KV of every full layer, so dense GQA reads roughly 24 GiB per token per request at full context, while the hybrid reads about a quarter of that plus tiny states. The KV cache article covers paging and quantising what remains, and the MLA article covers compressing it.
Operating and choosing
- Serving engines must support the mix. A hybrid model needs separate memory pools for growing KV and fixed states, and prefix caching must snapshot recurrent state at block boundaries rather than reuse pages. Check your engine's support for the exact architecture before committing; a fallback path can be many times slower.
- Speculative decoding gets harder. Rejected draft tokens must be rolled back. That is trivial for a KV cache (truncate) but needs saved state checkpoints for recurrent layers.
- Precision matters in recurrent state. States are usually kept in FP32 even when weights are BF16 or FP8; quantising states aggressively accumulates error over long sequences.
- Evaluate the failure you fear. Needle-in-a-haystack tests are easy for sparse and hybrid designs. Test multi-hop retrieval, long code edits and exact copying of long spans at your target length.
- KV quantisation stacks with everything. Storing the remaining full-layer cache in FP8 halves it again for any of these designs, and DeepSeek-V3.2-Exp's published serving path supports an FP8 cache. Validate quality at your longest contexts, where quantisation error has the most tokens to accumulate across.
- Converting is possible but not free. The original GQA paper converted a multi-head checkpoint by mean-pooling each group's key and value heads, then continued pre-training for a small fraction of the original compute to recover quality. Converting to a hybrid or MLA design is a larger surgery: new parameters must be distilled or trained, so budget it as a training project, not a config change.
- Training cost differs from serving cost. Chunked linear-attention kernels and sparse kernels are less mature than FlashAttention; check throughput in your framework, not in the paper.
Choosing: for contexts under about 32K, dense GQA with a good kernel is simplest and nearly as cheap. For long-context serving where you control the model, a local-global stack or a 3:1 linear hybrid buys the most concurrency. For long-context compute with exact recall, indexer sparsity on top of MLA is the strongest published option. If your stack cannot support a new layer type end to end, full attention plus KV quantisation is a respectable answer, as MiniMax's reversal shows.
What to do next
- Write down your context length distribution and concurrency target; the right variant depends on both.
- Compute per-request KV and state memory for candidate models with the script above, using their real configs.
- Confirm your serving engine supports each candidate's layer types, prefix caching and speculative decoding.
- Build a long-context evaluation with multi-hop retrieval and exact copying at your target length, not only needles.
- Measure throughput and quality at 8K, 32K and 128K before committing.
- Track new releases by asking three questions: which layers keep full KV, what is stored per token, and what is read per step.