"Attention Is All You Need" (Vaswani and colleagues, 2017) replaced recurrence in machine translation with attention alone and became the template for nearly every large language model since. Most people know it through diagrams redrawn many times. Reading the paper itself is still worth it, because the original model differs from today's in ways that matter when you reproduce results, read older code or reason about why modern choices were made.

This article reads the paper closely. It covers the architecture exactly as specified, the argument behind each design choice with numbers you can rerun, a faithful implementation of one layer, the full training recipe, a parameter count, what the ablations showed, and what survived into 2026 models. The modern block with full shapes is covered in anatomy of a transformer block; this page stays with the 2017 design and why it looked the way it did.

Advertisement

The problem the paper set out to fix

In 2017 the best translation systems were recurrent encoder-decoders with attention. A recurrent network computes the state for token t from the state for token t-1, so a sentence of n tokens needs n sequential steps that cannot run in parallel, however many GPU cores you have. Information from token 1 reaches token n only by passing through every state in between, which makes long-range dependencies hard to learn.

The paper's claim is that attention by itself can do the whole job. Self-attention lets every position look at every other position in one step: the computation inside a layer is parallel across positions, and any two tokens are connected by a path of length one. The paper's Table 1 states the trade:

Layer typeWork per layerSequential stepsLongest path between tokens
Self-attentionO(n² · d)O(1)O(1)
RecurrentO(n · d²)O(n)O(n)
Convolutional, kernel kO(k · n · d²)O(1)O(logk n)
Self-attention restricted to r neighboursO(r · n · d)O(1)O(n/r)

The paper notes that self-attention is cheaper than recurrence whenever n is smaller than d, which held for sentence-level translation: sentences of a few dozen tokens against d = 512. The quadratic term in n was not the bottleneck then. It became the bottleneck once contexts grew to many thousands of tokens, which is why later work such as FlashAttention exists.

The architecture as specified

The 2017 Transformer: post-norm encoder-decoder, N = 6 layers each sidesource tokensembed x sqrt(512) + sinusoidself-attention, 8 headsLayerNorm(x + Sublayer(x))feed-forward 512-2048-512ReLU, then add and normx6 encoder layersmemoryone vector per source tokentarget tokens, shifted rightsame shared embeddingmasked self-attentionposition i sees only positions up to iencoder-decoder attentionQ from decoder, K and V from memoryfeed-forwardadd and norm after each sublayerx6 decoder layerslinear + softmaxweights tied to the embeddingK, V to every decoder layer
Each sublayer is wrapped as LayerNorm(x + Sublayer(x)): normalisation after the residual add, which later models moved before the sublayer.

The model is an encoder-decoder. The encoder is a stack of six identical layers, each with a multi-head self-attention sublayer and a position-wise feed-forward sublayer. The decoder is a stack of six layers with three sublayers: masked self-attention over the target produced so far, encoder-decoder attention whose queries come from the decoder and whose keys and values come from the encoder output, and a feed-forward network. All sublayers produce vectors of width dmodel = 512 so the residual additions line up.

Four details are easy to miss. First, every sublayer is wrapped as LayerNorm(x + Sublayer(x)), with dropout applied to the sublayer output before the add: the normalisation sits after the residual connection. Second, the source embedding, the target embedding and the pre-softmax output projection share one weight matrix, and the embeddings are multiplied by √dmodel. Third, position enters only once, at the bottom, by adding fixed sinusoids, PE(pos, 2i) = sin(pos / 100002i/d) and PE(pos, 2i+1) = cos(the same), so each pair of dimensions is a clock with its own wavelength. Fourth, the decoder input is the target shifted right by one, and its self-attention mask stops position i from seeing positions after i, so training on whole sentences in parallel never leaks the answer.

The feed-forward sublayer is two linear maps with a ReLU between, expanding 512 to 2,048 and back, applied independently at each position.

Advertisement

Why the dot product is divided by the square root of d_k

Attention computes softmax(QKT / √dk) V. If the components of a query and a key are independent with mean 0 and variance 1, their dot product over dk dimensions has variance dk. The paper gives this argument in a footnote. Running it with 4,000 random pairs gives a standard deviation of 1.99 for dk = 4, 7.89 for dk = 64 and 22.54 for dk = 512; after dividing by √dk all three are 0.99 to 1.0.

The scale matters because softmax saturates. Ten logits with standard deviation √512 put 0.97 of the probability on a single key in one seeded draw; the same logits scaled down give a maximum of 0.31. A saturated softmax has gradients near zero for every key except the winner, so the layer stops learning which key it should attend to. Dividing by √dk keeps the logits at unit scale at initialisation whatever the head size.

Multi-head attention splits dmodel = 512 into h = 8 heads of dk = dv = 64. Each head has its own projections and attends independently; the head outputs are concatenated and projected back to 512. The cost is about that of one head of width 512, but each head can learn a different relation.

A faithful layer in code

This PyTorch encoder layer follows the paper rather than modern practice: post-norm, ReLU, biases everywhere and dropout of 0.1 on each sublayer output. The decoder layer adds a masked self-attention and a cross-attention sublayer with the same wrapping.

import math, torch, torch.nn as nn, torch.nn.functional as F

class MultiHead(nn.Module):
    def __init__(self, d=512, h=8):
        super().__init__()
        self.h, self.dk = h, d // h
        self.q, self.k, self.v, self.o = (nn.Linear(d, d) for _ in range(4))

    def forward(self, x, mem=None, mask=None):        # mask: True where attention is allowed
        mem = x if mem is None else mem
        B, T, D = x.shape
        split = lambda t: t.view(B, -1, self.h, self.dk).transpose(1, 2)
        q, k, v = split(self.q(x)), split(self.k(mem)), split(self.v(mem))
        att = q @ k.transpose(-2, -1) / math.sqrt(self.dk)
        if mask is not None:
            att = att.masked_fill(~mask, float("-inf"))
        out = att.softmax(-1) @ v
        return self.o(out.transpose(1, 2).reshape(B, T, D))

class EncoderLayer(nn.Module):
    def __init__(self, d=512, h=8, ff=2048, p=0.1):
        super().__init__()
        self.att, self.drop = MultiHead(d, h), nn.Dropout(p)
        self.ff = nn.Sequential(nn.Linear(d, ff), nn.ReLU(), nn.Linear(ff, d))
        self.n1, self.n2 = nn.LayerNorm(d), nn.LayerNorm(d)

    def forward(self, x, pad_mask):                    # post-norm: LayerNorm(x + Sublayer(x))
        x = self.n1(x + self.drop(self.att(x, mask=pad_mask)))
        return self.n2(x + self.drop(self.ff(x)))

def sinusoids(n, d=512):
    pos = torch.arange(n).unsqueeze(1)
    div = torch.exp(torch.arange(0, d, 2) * (-math.log(10000.0) / d))
    pe = torch.zeros(n, d)
    pe[:, 0::2], pe[:, 1::2] = torch.sin(pos * div), torch.cos(pos * div)
    return pe

The training recipe

The architecture gets the attention, but the recipe is what made it train. The paper used WMT 2014 English-German, about 4.5 million sentence pairs with a shared byte-pair vocabulary of about 37,000 tokens, and English-French, 36 million sentences with a 32,000 word-piece vocabulary. Batches were built by sequence length to hold about 25,000 source and 25,000 target tokens. Training ran on one machine with eight P100 GPUs: the base model took about 0.4 seconds per step for 100,000 steps, or 12 hours, and the big model 1.0 seconds per step for 300,000 steps, or 3.5 days.

SettingBaseBig
Layers, dmodel, dff, heads6, 512, 2048, 86, 1024, 4096, 16
Dropout0.10.3 for English-German, 0.1 for English-French
OptimizerAdam, β1 = 0.9, β2 = 0.98, ε = 1e-9same
Label smoothing0.10.1
Checkpoints averagedlast 5, written every 10 minuteslast 20
Decodingbeam 4, length penalty α = 0.6same

The learning rate rises linearly for warmup_steps = 4,000 and then decays with the inverse square root of the step number, scaled by dmodel-0.5. Often called the Noam schedule, it peaks at 6.99e-4 for the base model and 4.94e-4 for the big one; for the base model it is 1.75e-4 at step 1,000, 4.42e-4 at step 10,000 and 1.40e-4 at step 100,000.

def lrate(step, d_model=512, warmup=4000):
    step = max(step, 1)
    return d_model ** -0.5 * min(step ** -0.5, step * warmup ** -1.5)

def smoothed_loss(logits, target, eps=0.1, pad_id=0):
    # Spread eps of the probability mass uniformly; ignore padding positions.
    logp = F.log_softmax(logits, dim=-1)
    nll = -logp.gather(-1, target.unsqueeze(-1)).squeeze(-1)
    uniform = -logp.mean(dim=-1)
    loss = (1 - eps) * nll + eps * uniform
    keep = target.ne(pad_id)
    return loss[keep].mean()

The paper reports that label smoothing "hurts perplexity, as the model learns to be more unsure, but improves accuracy and BLEU score". Checkpoint averaging is cheap ensembling: averaging the weights of the last few checkpoints smooths optimiser noise at no inference cost. With this recipe the big model reached 28.4 BLEU on English-German newstest2014 and 41.8 on English-French in the revised arXiv version, at a fraction of the training cost of earlier systems.

Worked example: counting the parameters

Counting parameters by hand is the fastest check that you understand a model. Take the base configuration with a 37,000-token shared vocabulary and biases on every linear layer:

  • One attention sublayer: four 512 × 512 projections plus biases, 4 × (262,144 + 512) = 1,050,624.
  • One feed-forward sublayer: 512 × 2,048 + 2,048 + 2,048 × 512 + 512 = 2,099,712.
  • LayerNorm: 1,024 per norm (scale and shift).
  • Encoder layer, one attention, one feed-forward and two norms: 3,152,384. Six of them: 18.9 million.
  • Decoder layer, two attentions, one feed-forward and three norms: 4,204,032. Six of them: 25.2 million.
  • Shared embedding and output matrix: 37,000 × 512 = 18.9 million, counted once because it is tied.

The total is 63.1 million, against the 65 million the paper reports; the gap depends on the exact vocabulary size and bookkeeping. The same arithmetic for the big model gives 214.2 million against the paper's 213 million. The embedding is almost a third of the base model, which is why tying it mattered, and the decoder is a third larger than the encoder because of its extra attention sublayer. If you count very differently, check whether your implementation unties the embeddings or drops biases.

What the ablations showed

Table 3 varies one thing at a time on the development set. "While single-head attention is 0.9 BLEU worse than the best setting, quality also drops off with too many heads": with total width fixed, more heads means narrower heads, and very narrow heads lose capacity. Reducing the key size dk hurt quality, which the authors read as a hint that a compatibility function richer than a dot product might help. Bigger models were better, and dropout was essential to avoid overfitting. Replacing the sinusoids with learned position embeddings gave "nearly identical results"; the authors kept sinusoids in the hope that they would extrapolate to longer sequences.

That hope did not hold up well. Absolute sinusoids added once at the input extrapolate poorly in practice, which is one reason later models moved to relative schemes such as rotary embeddings; see positional encodings and rotary position embeddings.

What survived and what changed

The core idea survived intact: scaled dot-product attention, multiple heads, residual streams, position-wise feed-forward networks and teacher-forced parallel training with a causal mask. Almost every default around it has changed in typical large decoder-only models, although no single list fits every model:

2017 choiceCommon in 2026 modelsWhy it changed
Encoder-decoderDecoder-only for language models; encoder-decoder still used for some translation and speech systemsOne stack trained on next-token prediction scales simply over any text
Post-normPre-norm, usually RMSNormPost-norm needs careful warmup and becomes unstable in deep stacks; pre-norm keeps a clean residual path
Sinusoids added at the inputRotary embeddings applied to queries and keysRelative position inside attention, and better behaviour at long context
ReLU feed-forwardGated feed-forward such as SwiGLUBetter quality for the same compute in published comparisons
Full multi-head attentionGrouped-query or latent attentionShrinks the KV cache that dominates inference memory; see multi-head latent attention
Dropout 0.1Often none in large pretrainingWith one pass over huge data, overfitting is not the main risk
Adam with the Noam scheduleAdamW, warmup, then cosine or similar decayDecoupled weight decay and schedules that end at a known step

The cross-attention sublayer, which modern language models dropped, lives on in encoder-decoder and multimodal systems; cross-attention covers it.

Pitfalls when reproducing the paper

  • Post-norm without warmup: the original layout diverges or stalls if you remove the warmup. Keep the schedule when you keep post-norm.
  • Masks: a padding mask applied to queries instead of keys, or a causal mask off by one, produces a model that trains to a suspiciously low loss and then generates garbage. Test that changing a future token never changes an earlier output.
  • Embedding scale: with tied weights, forgetting the √dmodel factor leaves token embeddings tiny next to the sinusoids, and position swamps content.
  • Batch size in tokens: the recipe assumes about 50,000 tokens per step. With less memory, use gradient accumulation to reach it rather than reusing the learning rate on smaller batches.
  • Fully masked rows: a row with every key masked yields NaN from softmax; make sure no query attends to nothing.
  • BLEU comparisons: scores from that era depend on tokenisation and evaluation scripts. Report sacreBLEU with its signature when you compare against them.

What to do next

  1. Read sections 3 and 5 of the paper with the diagram above beside you, and match each sentence to a line of the code.
  2. Count the parameters of the base and big models yourself before running anything.
  3. Train a small post-norm model, then switch it to pre-norm without warmup and compare the loss curves.
  4. Write the future-token test for your causal mask and keep it in your test suite.
  5. Swap the sinusoids for rotary embeddings and compare quality on sequences longer than you trained on.
  6. List the deviations from 2017 in the model you use at work, with the reason for each.
Key takeaway: The 2017 Transformer is scaled dot-product attention in a post-norm encoder-decoder with sinusoidal positions, tied embeddings and a recipe built on warmup, label smoothing, dropout and checkpoint averaging. The attention core survived unchanged; nearly everything around it has since been replaced for stability, long context or inference cost, and knowing which is which is what lets you read old code and reproduce old results.