"Attention Is All You Need" (Vaswani and colleagues, 2017) replaced recurrence in machine translation with attention alone and became the template for nearly every large language model since. Most people know it through diagrams redrawn many times. Reading the paper itself is still worth it, because the original model differs from today's in ways that matter when you reproduce results, read older code or reason about why modern choices were made.
This article reads the paper closely. It covers the architecture exactly as specified, the argument behind each design choice with numbers you can rerun, a faithful implementation of one layer, the full training recipe, a parameter count, what the ablations showed, and what survived into 2026 models. The modern block with full shapes is covered in anatomy of a transformer block; this page stays with the 2017 design and why it looked the way it did.
The problem the paper set out to fix
In 2017 the best translation systems were recurrent encoder-decoders with attention. A recurrent network computes the state for token t from the state for token t-1, so a sentence of n tokens needs n sequential steps that cannot run in parallel, however many GPU cores you have. Information from token 1 reaches token n only by passing through every state in between, which makes long-range dependencies hard to learn.
The paper's claim is that attention by itself can do the whole job. Self-attention lets every position look at every other position in one step: the computation inside a layer is parallel across positions, and any two tokens are connected by a path of length one. The paper's Table 1 states the trade:
| Layer type | Work per layer | Sequential steps | Longest path between tokens |
|---|---|---|---|
| Self-attention | O(n² · d) | O(1) | O(1) |
| Recurrent | O(n · d²) | O(n) | O(n) |
| Convolutional, kernel k | O(k · n · d²) | O(1) | O(logk n) |
| Self-attention restricted to r neighbours | O(r · n · d) | O(1) | O(n/r) |
The paper notes that self-attention is cheaper than recurrence whenever n is smaller than d, which held for sentence-level translation: sentences of a few dozen tokens against d = 512. The quadratic term in n was not the bottleneck then. It became the bottleneck once contexts grew to many thousands of tokens, which is why later work such as FlashAttention exists.
The architecture as specified
The model is an encoder-decoder. The encoder is a stack of six identical layers, each with a multi-head self-attention sublayer and a position-wise feed-forward sublayer. The decoder is a stack of six layers with three sublayers: masked self-attention over the target produced so far, encoder-decoder attention whose queries come from the decoder and whose keys and values come from the encoder output, and a feed-forward network. All sublayers produce vectors of width dmodel = 512 so the residual additions line up.
Four details are easy to miss. First, every sublayer is wrapped as LayerNorm(x + Sublayer(x)), with dropout applied to the sublayer output before the add: the normalisation sits after the residual connection. Second, the source embedding, the target embedding and the pre-softmax output projection share one weight matrix, and the embeddings are multiplied by √dmodel. Third, position enters only once, at the bottom, by adding fixed sinusoids, PE(pos, 2i) = sin(pos / 100002i/d) and PE(pos, 2i+1) = cos(the same), so each pair of dimensions is a clock with its own wavelength. Fourth, the decoder input is the target shifted right by one, and its self-attention mask stops position i from seeing positions after i, so training on whole sentences in parallel never leaks the answer.
The feed-forward sublayer is two linear maps with a ReLU between, expanding 512 to 2,048 and back, applied independently at each position.
Why the dot product is divided by the square root of d_k
Attention computes softmax(QKT / √dk) V. If the components of a query and a key are independent with mean 0 and variance 1, their dot product over dk dimensions has variance dk. The paper gives this argument in a footnote. Running it with 4,000 random pairs gives a standard deviation of 1.99 for dk = 4, 7.89 for dk = 64 and 22.54 for dk = 512; after dividing by √dk all three are 0.99 to 1.0.
The scale matters because softmax saturates. Ten logits with standard deviation √512 put 0.97 of the probability on a single key in one seeded draw; the same logits scaled down give a maximum of 0.31. A saturated softmax has gradients near zero for every key except the winner, so the layer stops learning which key it should attend to. Dividing by √dk keeps the logits at unit scale at initialisation whatever the head size.
Multi-head attention splits dmodel = 512 into h = 8 heads of dk = dv = 64. Each head has its own projections and attends independently; the head outputs are concatenated and projected back to 512. The cost is about that of one head of width 512, but each head can learn a different relation.
A faithful layer in code
This PyTorch encoder layer follows the paper rather than modern practice: post-norm, ReLU, biases everywhere and dropout of 0.1 on each sublayer output. The decoder layer adds a masked self-attention and a cross-attention sublayer with the same wrapping.
import math, torch, torch.nn as nn, torch.nn.functional as F
class MultiHead(nn.Module):
def __init__(self, d=512, h=8):
super().__init__()
self.h, self.dk = h, d // h
self.q, self.k, self.v, self.o = (nn.Linear(d, d) for _ in range(4))
def forward(self, x, mem=None, mask=None): # mask: True where attention is allowed
mem = x if mem is None else mem
B, T, D = x.shape
split = lambda t: t.view(B, -1, self.h, self.dk).transpose(1, 2)
q, k, v = split(self.q(x)), split(self.k(mem)), split(self.v(mem))
att = q @ k.transpose(-2, -1) / math.sqrt(self.dk)
if mask is not None:
att = att.masked_fill(~mask, float("-inf"))
out = att.softmax(-1) @ v
return self.o(out.transpose(1, 2).reshape(B, T, D))
class EncoderLayer(nn.Module):
def __init__(self, d=512, h=8, ff=2048, p=0.1):
super().__init__()
self.att, self.drop = MultiHead(d, h), nn.Dropout(p)
self.ff = nn.Sequential(nn.Linear(d, ff), nn.ReLU(), nn.Linear(ff, d))
self.n1, self.n2 = nn.LayerNorm(d), nn.LayerNorm(d)
def forward(self, x, pad_mask): # post-norm: LayerNorm(x + Sublayer(x))
x = self.n1(x + self.drop(self.att(x, mask=pad_mask)))
return self.n2(x + self.drop(self.ff(x)))
def sinusoids(n, d=512):
pos = torch.arange(n).unsqueeze(1)
div = torch.exp(torch.arange(0, d, 2) * (-math.log(10000.0) / d))
pe = torch.zeros(n, d)
pe[:, 0::2], pe[:, 1::2] = torch.sin(pos * div), torch.cos(pos * div)
return pe
The training recipe
The architecture gets the attention, but the recipe is what made it train. The paper used WMT 2014 English-German, about 4.5 million sentence pairs with a shared byte-pair vocabulary of about 37,000 tokens, and English-French, 36 million sentences with a 32,000 word-piece vocabulary. Batches were built by sequence length to hold about 25,000 source and 25,000 target tokens. Training ran on one machine with eight P100 GPUs: the base model took about 0.4 seconds per step for 100,000 steps, or 12 hours, and the big model 1.0 seconds per step for 300,000 steps, or 3.5 days.
| Setting | Base | Big |
|---|---|---|
| Layers, dmodel, dff, heads | 6, 512, 2048, 8 | 6, 1024, 4096, 16 |
| Dropout | 0.1 | 0.3 for English-German, 0.1 for English-French |
| Optimizer | Adam, β1 = 0.9, β2 = 0.98, ε = 1e-9 | same |
| Label smoothing | 0.1 | 0.1 |
| Checkpoints averaged | last 5, written every 10 minutes | last 20 |
| Decoding | beam 4, length penalty α = 0.6 | same |
The learning rate rises linearly for warmup_steps = 4,000 and then decays with the inverse square root of the step number, scaled by dmodel-0.5. Often called the Noam schedule, it peaks at 6.99e-4 for the base model and 4.94e-4 for the big one; for the base model it is 1.75e-4 at step 1,000, 4.42e-4 at step 10,000 and 1.40e-4 at step 100,000.
def lrate(step, d_model=512, warmup=4000):
step = max(step, 1)
return d_model ** -0.5 * min(step ** -0.5, step * warmup ** -1.5)
def smoothed_loss(logits, target, eps=0.1, pad_id=0):
# Spread eps of the probability mass uniformly; ignore padding positions.
logp = F.log_softmax(logits, dim=-1)
nll = -logp.gather(-1, target.unsqueeze(-1)).squeeze(-1)
uniform = -logp.mean(dim=-1)
loss = (1 - eps) * nll + eps * uniform
keep = target.ne(pad_id)
return loss[keep].mean()The paper reports that label smoothing "hurts perplexity, as the model learns to be more unsure, but improves accuracy and BLEU score". Checkpoint averaging is cheap ensembling: averaging the weights of the last few checkpoints smooths optimiser noise at no inference cost. With this recipe the big model reached 28.4 BLEU on English-German newstest2014 and 41.8 on English-French in the revised arXiv version, at a fraction of the training cost of earlier systems.
Worked example: counting the parameters
Counting parameters by hand is the fastest check that you understand a model. Take the base configuration with a 37,000-token shared vocabulary and biases on every linear layer:
- One attention sublayer: four 512 × 512 projections plus biases, 4 × (262,144 + 512) = 1,050,624.
- One feed-forward sublayer: 512 × 2,048 + 2,048 + 2,048 × 512 + 512 = 2,099,712.
- LayerNorm: 1,024 per norm (scale and shift).
- Encoder layer, one attention, one feed-forward and two norms: 3,152,384. Six of them: 18.9 million.
- Decoder layer, two attentions, one feed-forward and three norms: 4,204,032. Six of them: 25.2 million.
- Shared embedding and output matrix: 37,000 × 512 = 18.9 million, counted once because it is tied.
The total is 63.1 million, against the 65 million the paper reports; the gap depends on the exact vocabulary size and bookkeeping. The same arithmetic for the big model gives 214.2 million against the paper's 213 million. The embedding is almost a third of the base model, which is why tying it mattered, and the decoder is a third larger than the encoder because of its extra attention sublayer. If you count very differently, check whether your implementation unties the embeddings or drops biases.
What the ablations showed
Table 3 varies one thing at a time on the development set. "While single-head attention is 0.9 BLEU worse than the best setting, quality also drops off with too many heads": with total width fixed, more heads means narrower heads, and very narrow heads lose capacity. Reducing the key size dk hurt quality, which the authors read as a hint that a compatibility function richer than a dot product might help. Bigger models were better, and dropout was essential to avoid overfitting. Replacing the sinusoids with learned position embeddings gave "nearly identical results"; the authors kept sinusoids in the hope that they would extrapolate to longer sequences.
That hope did not hold up well. Absolute sinusoids added once at the input extrapolate poorly in practice, which is one reason later models moved to relative schemes such as rotary embeddings; see positional encodings and rotary position embeddings.
What survived and what changed
The core idea survived intact: scaled dot-product attention, multiple heads, residual streams, position-wise feed-forward networks and teacher-forced parallel training with a causal mask. Almost every default around it has changed in typical large decoder-only models, although no single list fits every model:
| 2017 choice | Common in 2026 models | Why it changed |
|---|---|---|
| Encoder-decoder | Decoder-only for language models; encoder-decoder still used for some translation and speech systems | One stack trained on next-token prediction scales simply over any text |
| Post-norm | Pre-norm, usually RMSNorm | Post-norm needs careful warmup and becomes unstable in deep stacks; pre-norm keeps a clean residual path |
| Sinusoids added at the input | Rotary embeddings applied to queries and keys | Relative position inside attention, and better behaviour at long context |
| ReLU feed-forward | Gated feed-forward such as SwiGLU | Better quality for the same compute in published comparisons |
| Full multi-head attention | Grouped-query or latent attention | Shrinks the KV cache that dominates inference memory; see multi-head latent attention |
| Dropout 0.1 | Often none in large pretraining | With one pass over huge data, overfitting is not the main risk |
| Adam with the Noam schedule | AdamW, warmup, then cosine or similar decay | Decoupled weight decay and schedules that end at a known step |
The cross-attention sublayer, which modern language models dropped, lives on in encoder-decoder and multimodal systems; cross-attention covers it.
Pitfalls when reproducing the paper
- Post-norm without warmup: the original layout diverges or stalls if you remove the warmup. Keep the schedule when you keep post-norm.
- Masks: a padding mask applied to queries instead of keys, or a causal mask off by one, produces a model that trains to a suspiciously low loss and then generates garbage. Test that changing a future token never changes an earlier output.
- Embedding scale: with tied weights, forgetting the √dmodel factor leaves token embeddings tiny next to the sinusoids, and position swamps content.
- Batch size in tokens: the recipe assumes about 50,000 tokens per step. With less memory, use gradient accumulation to reach it rather than reusing the learning rate on smaller batches.
- Fully masked rows: a row with every key masked yields NaN from softmax; make sure no query attends to nothing.
- BLEU comparisons: scores from that era depend on tokenisation and evaluation scripts. Report sacreBLEU with its signature when you compare against them.
What to do next
- Read sections 3 and 5 of the paper with the diagram above beside you, and match each sentence to a line of the code.
- Count the parameters of the base and big models yourself before running anything.
- Train a small post-norm model, then switch it to pre-norm without warmup and compare the loss curves.
- Write the future-token test for your causal mask and keep it in your test suite.
- Swap the sinusoids for rotary embeddings and compare quality on sequences longer than you trained on.
- List the deviations from 2017 in the model you use at work, with the reason for each.