A gate is a learned number between two bounds, usually produced from the input, that multiplies some other signal. That definition sounds too small to deserve an article, yet gates now appear at five distinct places in a modern decoder block, and each placement fixes a different problem: attention heads that cannot say nothing, adapters that would wreck a pretrained model on step one, residual branches that destabilise deep stacks, and linear-time layers that cannot forget.

This article treats gating as one design tool with several uses: each placement from first principles, PyTorch for the ones you are most likely to add, a numeric example, kernels, failure modes and diagnostics.

Advertisement

What counts as a gate, and why multiplication

A residual block adds: x + f(x). Addition can only push a signal; it cannot switch one off unless the branch learns to produce an exact cancellation. A gate multiplies: g * f(x) with g in (0, 1) for a sigmoid or (-1, 1) for a tanh. Multiplication gives the network a cheap way to say how much of this signal, right now, for this token. Because g is usually computed from the current input, the same weights can pass a signal for one token and suppress it for the next.

Gates differ along three axes: what is multiplied (a head output, a recurrent state, a residual branch, a choice among experts), granularity (per layer, per head or per channel), and initialisation, which decides whether the model behaves sensibly on the first step of training.

Where gates sit in a modern decoder block: five multiplicative switches, each with its own jobHidden state xresidual streamAttention / mixersoftmax or linearOutput gatey * sigmoid(xW)W_oout projyDecay gateS = a * S + v klinear layersResidual scalarx + alpha * f(x)add to residualFFN: GLU gateact(xW1) * xW3MoE routertop-k of softmax(xWr)orGated cross-attention (adapters)x + tanh(alpha) * xattn(x, image), alpha = 0Diagnosticsgate histograms, first-token attention mass
Gate placements in a decoder block. Not every model uses every gate; a hybrid such as Qwen3-Next uses output-gated softmax attention in some layers and decay-gated linear layers in others.

The GLU feed-forward gate, briefly

The feed-forward network in most current LLMs is a gated linear unit: W2(act(x W1) * (x W3)), where the activation is SiLU for SwiGLU or GELU for GeGLU. One projection proposes features and the other decides, per token and channel, how much of each to keep; the hidden width is usually cut to about two thirds of 4d to hold the parameter count. Details are in SwiGLU architecture. This gate is channelwise, input-dependent and starts neither open nor closed.

Advertisement

The router as a gate over experts, briefly

A mixture-of-experts layer replaces the single FFN with many and lets a router choose. The router is a gate with a discrete twist: it computes a score per expert, keeps the top k, and multiplies each chosen expert's output by its normalised score. That multiplication is what lets gradient reach the router, since top-k selection is not differentiable. Balancing and capacity belong to MoE Transformer architecture; the rest of this article covers gates that scale a single signal.

The attention output gate

Softmax attention forces each head to spend exactly one unit of attention per query. A head that has nothing useful to contribute for a given token still has to put that mass somewhere. Trained models solve this by dumping it on a token whose value vector is small, very often the first token in the sequence. That habit is the attention sink, described in Attention sinks: it works, but it couples every head to one position, produces very large activations, and complicates streaming and long-context extrapolation.

The fix studied in the NeurIPS 2025 paper Gated Attention for Large Language Models (Qiu et al., Qwen team) is to multiply the scaled dot-product attention output by sigmoid(x W_g) before the output projection. The gate is computed from the hidden state that produced the query, so a head can output almost nothing for a token directly, rather than by aiming at a sink. Of the positions the paper compares, this one is the most effective; it reports more stable training, tolerance of larger learning rates, sparse gate values and the disappearance of the attention sink. Qwen3-Next uses this gated attention in its softmax layers.

Elementwise gating, one gate per channel of every head, costs a d x d projection per layer, about one more Q, K, V or O matrix. Headwise gating, one scalar per head, costs d x n_heads. The code supports both. Note that the gate reads x, not y, and sits before W_o so each head is gated separately.

import torch
import torch.nn as nn
import torch.nn.functional as F

class GatedSelfAttention(nn.Module):
    """Causal self-attention with an elementwise sigmoid gate on the SDPA output,
    applied per head channel before the output projection."""
    def __init__(self, d_model: int, n_heads: int, headwise: bool = False):
        super().__init__()
        self.h, self.dh = n_heads, d_model // n_heads
        self.qkv = nn.Linear(d_model, 3 * d_model, bias=False)
        # Elementwise: one gate per channel (d_model x d_model extra weights).
        # Headwise: one scalar per head (d_model x n_heads extra weights).
        self.gate = nn.Linear(d_model, n_heads if headwise else d_model, bias=True)
        self.out = nn.Linear(d_model, d_model, bias=False)
        self.headwise = headwise

    def forward(self, x):                       # x: [B, T, D]
        B, T, D = x.shape
        q, k, v = self.qkv(x).view(B, T, 3, self.h, self.dh).unbind(2)
        q, k, v = (t.transpose(1, 2) for t in (q, k, v))     # [B, H, T, dh]
        y = F.scaled_dot_product_attention(q, k, v, is_causal=True)
        y = y.transpose(1, 2)                                # [B, T, H, dh]
        g = torch.sigmoid(self.gate(x))                      # query-dependent
        g = g.view(B, T, self.h, 1) if self.headwise else g.view(B, T, self.h, self.dh)
        return self.out((y * g).reshape(B, T, D))

Worked example: one head with nothing to say

Take a head that, for the current token, finds no relevant earlier token. Without a gate, suppose it puts 0.90 of its attention on position 0, whose value vector has norm about 0.1, and spreads 0.10 over thirty other positions whose value vectors have norm about 1 and point in assorted directions. The output norm is roughly 0.90 * 0.1 + 0.10 * (about 0.2 after partial cancellation), about 0.11. The head has achieved near-silence, but only by keeping position 0's value small and its key attractive, for every head that uses the trick.

With an output gate, the head can attend wherever its scores land, say evenly over the thirty-one positions, giving an output norm of perhaps 0.2, and the gate for this token emits 0.03. The contribution is now about 0.006 and position 0 is no longer special. Where the head is useful, the gate opens towards 1. In a trained gated model you should see many gate values near zero and first-token attention mass close to that of any other position.

Gated cross-attention and residual scalars

Adding layers to a model that already works raises a different problem. A randomly initialised cross-attention branch inserted into a frozen language model adds noise to every residual stream on step one, degrading outputs before it has learned anything. Flamingo's answer was the gated cross-attention block: x + tanh(alpha) * xattn(x, image), with alpha a scalar initialised to zero; the new block's feed-forward layer gets its own zero-initialised tanh gate too. At initialisation the whole model is exactly the original language model.

Notice the gradient structure. With alpha = 0, tanh(alpha) = 0, so the cross-attention weights receive zero gradient on the first step. The scalar alpha does receive gradient, proportional to how much the branch output would help, so it moves first and opens the gate; then the branch weights learn. With a tiny learning rate for alpha, the adapter can look stuck.

Applied to every residual branch, the same idea gives ReZero (a scalar starting at zero) and LayerScale (a per-channel vector starting small), which make deep stacks trainable by starting each block near the identity. One wrapper covers all three:

class GatedBranch(nn.Module):
    """Wraps any sublayer f so the block starts as (almost) the identity."""
    def __init__(self, f: nn.Module, d_model: int, mode: str = "tanh"):
        super().__init__()
        self.f, self.mode = f, mode
        if mode == "tanh":          # Flamingo-style: one scalar, starts at exactly 0
            self.alpha = nn.Parameter(torch.zeros(1))
        elif mode == "layerscale":  # per-channel, starts small but non-zero
            self.alpha = nn.Parameter(torch.full((d_model,), 1e-4))
        else:                       # ReZero: one scalar, starts at 0
            self.alpha = nn.Parameter(torch.zeros(1))

    def forward(self, x, *args):
        scale = torch.tanh(self.alpha) if self.mode == "tanh" else self.alpha
        return x + scale * self.f(x, *args)

Decay gates in linear attention and recurrent layers

Linear attention replaces the softmax with a running state S that accumulates v k^T per token and is read with the query, so decoding needs constant memory instead of a growing KV cache. The plain version never forgets, so the state fills with stale associations.

A decay gate multiplies the state by a learned, input-dependent factor before each write. Gated Linear Attention (GLA) uses a per-key-channel decay vector; Mamba's selective state-space layers make the decay input-dependent through the step size. Gated DeltaNet combines a scalar decay with the delta rule: before writing v under key k, it erases whatever the state currently returns for k, scaled by a write strength beta. The decay handles forgetting in bulk; the delta rule handles precise overwriting.

# Recurrent (decode-time) form of two decay-gated linear mixers, one head.
# Plain linear attention would be S = S + outer(v, k): it never forgets.
# S is a [d_v, d_k] state matrix; k_t, q_t are [d_k]; v_t is [d_v].

def gla_step(S, q, k, v, alpha):              # alpha in (0,1)^d_k, data-dependent
    S = S * alpha[None, :] + outer(v, k)      # per key-channel forgetting
    return S, S @ q

def gated_delta_step(S, q, k, v, a, beta):    # a in (0,1) scalar decay, beta in (0,1)
    k = k / norm(k)
    S = a * (S - beta * outer(S @ k, k))      # decay, then erase what k used to point at
    S = S + beta * outer(v, k)                # write the new association
    return S, S @ q

Qwen3-Next's published layout repeats three Gated DeltaNet layers followed by one gated softmax attention layer, each followed by an MoE block, so a single model carries decay gates, output gates and expert routing at once. The design logic is that the linear layers carry most of the sequence cheaply and the periodic softmax layers recover the exact retrieval that a fixed-size state cannot.

How gates meet the hardware

Almost every gate is an elementwise sigmoid or tanh and a multiply: memory-bound work where the cost is reading and writing the tensor. As separate kernels, an output gate adds a full read and write of the attention output. Fuse it: widen the QKV GEMM to produce the gate projection too, and apply sigmoid and multiply in the attention kernel's epilogue or the output projection's prologue. A standalone sigmoid kernel per layer in the profiler means the gate costs more than it should.

Training cannot afford the token-by-token recurrence of decay gates, so these layers use a chunkwise-parallel form: within a chunk, cumulative decays turn the recurrence into masked matrix multiplications on tensor cores; across chunks, a small state is carried. Products of numbers below one underflow fast in bf16, so implementations use cumulative sums of log-decays in fp32.

Failure modes

  • Dropped gate weights. A converter or inference engine unaware of the output gate silently ignores its projection; the model produces fluent but degraded text. Assert every checkpoint tensor is consumed.
  • Saturated gates. A sigmoid pinned near 0 or 1 passes almost no gradient; check the learning rate and that the gate input is normalised.
  • Gates that never open. Weight decay or a tiny learning rate keeps a zero-initialised tanh scalar near zero. Exclude gate scalars from weight decay.
  • Forgetting too fast. Small decays shrink a linear layer's memory to a few dozen tokens; long-context failures in hybrid models often live here.
  • Numerical underflow. Decay products in bf16 reach zero within a chunk and produce NaNs. Use log-space sums.

Gate values have meaning, so log them. Per layer, track the mean output-gate value and the fraction below 0.1 (a gate stuck near 0.5 everywhere is learning nothing), the attention mass on position 0, the residual scalars over training, and decay per head. These four plots catch most of the failures above before an evaluation run does.

Trade-offs

GateFixesCostsUse when
Attention output gateAttention sinks, instability, large activationsOne projection per layer (elementwise) or almost nothing (headwise)Training from scratch, especially long context
Zero-init tanh branchDamage when adding layers to a trained modelSlow start until the gate opensAdapters, cross-modal layers on a frozen LM
ReZero / LayerScaleInstability in very deep stacksExtra scalars, one more thing to tuneDeep models or high learning rates
Decay gateLinear attention cannot forgetChunked kernels, careful numericsLinear-time or hybrid architectures
MoE routerCompute per token versus capacityBalancing, communication, memoryLarge models with sparse compute

What to do next

  1. List every gate in the model you train or serve, with its placement, granularity and initial value, using the diagram above as a template.
  2. If you train from scratch, run a small ablation with the elementwise output gate from the code above and compare loss curves and first-token attention mass.
  3. If you add layers to a frozen model, wrap each new branch in a zero-initialised tanh gate, exclude the scalar from weight decay, and log its value.
  4. Add the four diagnostic plots to your training dashboard: gate value distribution, first-token attention, residual scalars and decay per layer.
  5. In your inference stack, assert that every checkpoint tensor is consumed, so a missing gate fails loudly.
  6. Profile one layer and confirm that the gate's sigmoid and multiply are fused rather than separate kernels.
Key takeaway: A gate is a learned, bounded, usually input-dependent multiplier, and modern transformers use it in several places for different reasons. The output gate on attention lets a head say nothing without an attention sink; zero-initialised tanh gates let you add layers to a trained model without breaking it; residual scalars stabilise deep stacks; decay gates let linear-time layers forget. Know where each gate sits, how it starts and what its values look like, fuse it into neighbouring kernels, and treat a gate stuck at one value as a bug.