A transformer is a function that takes a sequence of tokens and returns, for every position, a probability distribution over the next token. That is the entire interface of GPT, Llama, Claude-class models and their relatives. Chat, code completion and tool use are all built on top of this one capability, applied over and over.

This introduction builds the architecture from first principles. It computes attention by hand, walks the residual stream, compares the model families, separates training from inference, and ends with a runnable PyTorch model and the cost formulas you need to size one. Each topic has a deeper page on this site; this one is the map that connects them.

Token ids[The, cat, sat]Embeddingids to d-dim vectors+ positionlearned, sinusoid or RoPETransformer block (repeated L times)Norm, then causal self-attentionNorm, then MLP (feed-forward)each sublayer adds its output to the residual streamx = x + Attn(norm(x)); x = x + MLP(norm(x))Final normUnembeddingd to vocab logitsSoftmaxnext-token probsA decoder-only transformer: everything between embedding and unembedding reads and writes one residual stream.
The decoder-only transformer pipeline: embed, add position, refine through L residual blocks, then map back to vocabulary probabilities.
Advertisement

The problem attention solved

Before 2017 the standard sequence model was the recurrent network. An RNN reads tokens one at a time and squeezes everything it has seen into a fixed-size hidden state. That design has two problems. Training is serial: step 500 cannot be computed until step 499 is done, so a GPU's thousands of cores sit mostly idle. And information from early tokens has to survive hundreds of updates to reach late ones, which in practice it often does not.

The transformer, introduced in "Attention Is All You Need" (Vaswani et al., 2017), removes recurrence entirely. Every position looks directly at every earlier position through attention, so the path between any two tokens is one step long. During training every position is computed at the same time, as large matrix multiplications, which is exactly the workload GPUs are built for. That parallelism is the real reason transformers scaled: you could finally turn more compute and data into a better model efficiently.

The problem attention solved

Before 2017 the standard sequence model was the recurrent network. An RNN reads tokens one at a time and squeezes everything it has seen into a fixed-size hidden state. That design has two problems. Training is serial: step 500 cannot be computed until step 499 is done, so a GPU's thousands of cores sit mostly idle. And information from early tokens has to survive hundreds of updates to reach late ones, which in practice it often does not.

The transformer, introduced in "Attention Is All You Need" (Vaswani et al., 2017), removes recurrence entirely. Every position looks directly at every earlier position through attention, so the path between any two tokens is one step long. During training every position is computed at the same time, as large matrix multiplications, which is exactly the workload GPUs are built for. That parallelism is the real reason transformers scaled: you could finally turn more compute and data into a better model efficiently.

Advertisement

The whole pipeline at a glance

Follow one input through the diagram above. Text is split into tokens by a tokenizer (see tokenization), and each token id selects a row of an embedding matrix, giving a vector of width d (768 in GPT-2 small, 4,096 in Llama 3 8B). The sequence is now a T by d matrix. Position information is added or injected. Then L identical blocks each refine the matrix. A final normalization and the unembedding matrix map each d-vector to a score per vocabulary entry, and softmax turns scores into probabilities.

The output has one distribution per position: position t predicts token t+1. During training all T predictions are scored at once; during generation only the last one is used.

Attention from first principles, computed by hand

Attention lets each position build its output as a weighted average of information from other positions, where the weights are computed from the content. Each position's vector is projected three ways: a query (what am I looking for), a key (what do I contain) and a value (what I pass on if selected). The score between position i and j is the dot product of query i and key j, divided by the square root of the key dimension so scores do not grow with width. Softmax over j turns scores into weights, and the output is the weighted sum of values:

Attention(Q, K, V) = softmax( Q K^T / sqrt(d_k) + M ) V

M is the causal mask: 0 where j <= i, minus infinity where j > i,
so no position can see the future.

Now by hand, with three tokens and d_k = 2. Let the keys be k1 = [1, 0], k2 = [0, 1], k3 = [1, 1] and the values v1 = [1, 0], v2 = [0, 2], v3 = [3, 1]. Take token 3, with query q3 = [1, 0]. Its raw scores against the three keys are 1, 0 and 1. Divide by the square root of 2: 0.707, 0, 0.707. Exponentiate: 2.028, 1.000, 2.028, which sum to 5.056. The weights are therefore 0.401, 0.198 and 0.401. The output is 0.401 x [1, 0] + 0.198 x [0, 2] + 0.401 x [3, 1] = [1.604, 0.797]. Token 3 has pulled mostly from tokens 1 and 3, whose keys matched its query, and a little from token 2.

Token 1 is the opposite extreme. The mask hides tokens 2 and 3, its single weight is 1, and its output is just v1. That asymmetry is what makes the model causal: each position's output depends only on its own past, so one forward pass can train T next-token predictions without any of them seeing the answer.

Multi-head attention runs this several times in parallel with smaller projections (for example 12 heads of 64 dimensions in place of one of 768), so different heads can track different relationships, such as the previous token or the subject of the sentence. Their outputs are concatenated and projected back to d. The full derivation and its numerical-stability details are in scaled dot-product attention.

The MLP and the residual stream

Attention moves information between positions. The second sublayer, the MLP (feed-forward network), transforms each position independently. It expands the vector to a wider hidden size (classically 4d), applies a nonlinearity, and projects back. About two thirds of a standard block's parameters sit here, and much of a model's stored factual knowledge appears to live in these weights.

Both sublayers are wrapped in a residual connection: each one reads the current vector, computes something, and adds its result back. It is useful to picture a single residual stream running from the embedding to the unembedding, with every sublayer reading from it and writing a small update into it. Residuals keep gradients flowing through deep stacks, because the identity path passes gradients straight back. A normalization layer before each sublayer ("pre-norm") keeps the scale of what each sublayer reads under control. The block-level shapes and parameter accounting are worked through in the complete transformer block.

Why position must be added

Attention by itself is order-blind. Shuffle the input tokens and each token's attention output is unchanged, just moved to the new position, because dot products and weighted sums do not care where a vector sits. "Dog bites man" and "man bites dog" would look identical. Position must therefore be supplied explicitly.

The original paper added fixed sinusoidal vectors to the embeddings. GPT-2 learned one position vector per index, which caps context at the trained length. Most current open models use rotary position embeddings (RoPE), which rotate query and key vectors by an angle proportional to their position, so that the query-key dot product depends on the relative distance between tokens. RoPE is also the basis of most context-extension tricks. The options are compared in positional embeddings.

Three families of transformer

FamilyAttentionTrained toExamplesUse for
Encoder-onlyBidirectional: every token sees every tokenFill in masked tokensBERT, RoBERTaClassification, embeddings, retrieval
Decoder-onlyCausalPredict the next tokenGPT series, Llama, MistralGeneration, chat, code, general assistants
Encoder-decoderEncoder bidirectional; decoder causal plus cross-attention to the encoderMap an input sequence to an output sequenceOriginal transformer, T5Translation, summarization, speech-to-text

Decoder-only models won the general-purpose race because a single objective (next-token prediction on raw text) scales without labelled data, and because any task can be phrased as "continue this text".

Training versus inference

Training feeds a whole sequence at once. The target is the input shifted by one token; the causal mask ensures position t only sees tokens up to t; the loss is the average cross-entropy over all positions. Everything is parallel across positions, so one forward and backward pass trains thousands of predictions.

Inference is serial. Generation produces one token, appends it, and runs again. Recomputing keys and values for the whole prefix every step would be wasteful, so servers keep a KV cache: the keys and values of every past token, per layer. Each new step then computes only the new token's query, key and value, and attends over the cache. This splits inference into two phases: prefill processes the prompt in parallel and is compute-bound, while decode produces tokens one at a time and is limited by memory bandwidth, since each step must stream all the weights and the cache. Most serving optimizations target one of these two phases. The cache arithmetic is in KV cache math.

What changed between 2017 and today

2017 / GPT-2 eraTypical nowWhy it changed
Post-norm (normalize after the residual add)Pre-normDeep stacks train stably without delicate warm-up
LayerNormRMSNormDrops mean-centring; cheaper and works as well
Sinusoidal or learned absolute positionsRoPERelative positions; extendable context
ReLU or GELU MLPGated MLP such as SwiGLUBetter quality per parameter
Multi-head attention, one K and V per headGrouped-query attention (several query heads share K and V)Shrinks the KV cache several times over
Dense MLPMixture-of-experts in many large modelsMore parameters for the same compute per token
Bias terms everywhereMostly no biasesSimpler, no measurable loss

None of these changes alters the skeleton. A 2026 model is still embeddings, a residual stream, alternating attention and MLP sublayers, and an unembedding. Learn the skeleton once and every new architecture paper reads as a diff against it.

A complete transformer in fifty lines

The model below is a complete, trainable decoder-only transformer in PyTorch 2.x. It uses pre-norm, learned positions, GELU, weight tying, and PyTorch's fused scaled_dot_product_attention with the causal flag. Train it on a text file of tokens and it will produce plausible text in that style.

import torch
import torch.nn as nn
import torch.nn.functional as F


class Block(nn.Module):
    def __init__(self, d, n_heads):
        super().__init__()
        self.n_heads = n_heads
        self.norm1 = nn.LayerNorm(d)
        self.qkv = nn.Linear(d, 3 * d, bias=False)
        self.proj = nn.Linear(d, d, bias=False)
        self.norm2 = nn.LayerNorm(d)
        self.mlp = nn.Sequential(nn.Linear(d, 4 * d), nn.GELU(), nn.Linear(4 * d, d))

    def forward(self, x):
        B, T, D = x.shape
        q, k, v = self.qkv(self.norm1(x)).split(D, dim=-1)
        # (B, T, D) -> (B, heads, T, head_dim)
        q, k, v = (t.view(B, T, self.n_heads, D // self.n_heads).transpose(1, 2) for t in (q, k, v))
        a = F.scaled_dot_product_attention(q, k, v, is_causal=True)
        x = x + self.proj(a.transpose(1, 2).reshape(B, T, D))   # residual 1
        return x + self.mlp(self.norm2(x))                       # residual 2


class TinyLM(nn.Module):
    def __init__(self, vocab, d=256, n_heads=4, n_layers=4, max_len=512):
        super().__init__()
        self.max_len = max_len
        self.tok = nn.Embedding(vocab, d)
        self.pos = nn.Embedding(max_len, d)
        self.blocks = nn.ModuleList(Block(d, n_heads) for _ in range(n_layers))
        self.norm = nn.LayerNorm(d)
        self.head = nn.Linear(d, vocab, bias=False)
        self.head.weight = self.tok.weight                       # weight tying

    def forward(self, idx):                                      # idx: (B, T) token ids
        T = idx.shape[1]
        x = self.tok(idx) + self.pos(torch.arange(T, device=idx.device))
        for block in self.blocks:
            x = block(x)
        return self.head(self.norm(x))                           # (B, T, vocab) logits


def train_step(model, opt, batch):                               # batch: (B, T + 1) ids
    inputs, targets = batch[:, :-1], batch[:, 1:]                # shift by one
    logits = model(inputs)
    loss = F.cross_entropy(logits.reshape(-1, logits.size(-1)), targets.reshape(-1))
    opt.zero_grad()
    loss.backward()
    opt.step()
    return loss.item()


@torch.no_grad()
def generate(model, idx, n_new, temperature=1.0):
    for _ in range(n_new):
        logits = model(idx[:, -model.max_len:])[:, -1, :] / temperature
        nxt = torch.multinomial(F.softmax(logits, dim=-1), num_samples=1)
        idx = torch.cat([idx, nxt], dim=1)
    return idx

Read it against the diagram: tok and pos are the embedding stage, each Block is one pass of norm, attention, residual, norm, MLP, residual, and head is the unembedding. train_step is the shifted-target, all-positions-at-once training described above. generate deliberately has no KV cache: it reruns the whole prefix every step, which is correct but slow. A sensible first run is AdamW at a learning rate around 3e-4 with batches of 32 sequences of 256 tokens; watch the loss fall below the log of the vocabulary size within a few hundred steps.

Costs: parameters, compute and KV cache

Three formulas size almost any decision. Parameters: each standard block has about 4d squared in attention projections and 8d squared in a 4x MLP, so roughly 12 x L x d squared overall, plus the embedding matrix (vocabulary x d). For GPT-2 small (L = 12, d = 768, vocabulary 50,257) that gives 85M in blocks plus 39M in embeddings, about 124M, matching the published size. Compute: a forward pass costs about 2 FLOPs per parameter per token, and training about 6, plus an attention term that grows with sequence length. KV cache: per token, 2 (K and V) x layers x KV heads x head dimension x bytes. Llama 3 8B has 32 layers, 8 KV heads of dimension 128, so in BF16 that is 2 x 32 x 8 x 128 x 2 = 131,072 bytes, or 128 KiB per token. An 8,192-token context costs 1 GiB per sequence, which is why grouped-query attention mattered: with 32 KV heads it would be 4 GiB.

Attention's score matrix grows with the square of sequence length, which is why long context is expensive and why kernels such as FlashAttention avoid storing that matrix in slow memory.

Common pitfalls

  • Off-by-one targets. Forgetting to shift targets makes the model copy its input; loss collapses to near zero and the model has learned nothing.
  • Missing or wrong mask. Without the causal mask, training looks excellent and generation is garbage, because the model learned to read the answer.
  • Exceeding the position range. Learned position tables have a fixed length; indexing past it crashes or, with some implementations, silently reuses positions.
  • Softmax overflow in low precision. Scores must be computed or normalized in higher precision; fused kernels do this for you, hand-written ones often do not.

What to do next

  1. Redo the three-token attention example with a different query, and check your result with a few lines of NumPy.
  2. Run the TinyLM code on a small text corpus with a byte-level or character vocabulary; confirm the loss drops.
  3. Add a KV cache to generate and verify its output matches the uncached version token for token.
  4. Swap in RMSNorm, RoPE and a SwiGLU MLP one at a time, and compare loss curves.
  5. Compute parameters and KV-cache size for one open model's config file using the formulas above.
  6. Continue with the linked deep dives on attention, the transformer block and KV cache math.
Key takeaway: A transformer maps tokens to next-token distributions using a residual stream refined by alternating attention, which mixes information across positions, and MLPs, which transform each position. The causal mask makes training parallel across positions, while generation is serial and relies on a KV cache. Modern models change components such as RMSNorm, RoPE, SwiGLU and grouped-query attention, not the skeleton, and roughly 12 L d squared parameters, 2 FLOPs per parameter per token and the KV-cache formula size most decisions.