BERT, from Devlin and colleagues at Google in 2018, showed that one pretrained Transformer encoder, fine-tuned with a single extra layer, could set the state of the art on question answering, natural language inference and named-entity tagging at once. Its key move was a training objective that let every token see context on both sides. A left-to-right language model cannot do that, because a token that can see the future can simply read off the answer. BERT avoided the problem by hiding the tokens it asks the model to predict.

Decoder models now dominate generation, but encoders in the BERT family still do much of the classification, retrieval and tagging in production, because they are small, fast and produce one contextual vector per token. This article explains how BERT turns text into inputs, how its two pretraining objectives work and which one later turned out not to matter, the exact recipe, how to fine-tune it, and what replaced each design choice. The mask itself is compared with the causal mask in Causal vs Bidirectional Attention.

The architecture

BERT is the encoder half of the original Transformer with no changes to the block: multi-head self-attention and a feed-forward layer, each wrapped in a residual connection followed by LayerNorm (the post-LN arrangement), with GELU in the feed-forward layer. The paper released two sizes.

ModelLayers LHidden HHeads AFeed-forwardParameters
BERT-base12768123,072110M
BERT-large241,024164,096340M

Attention has no causal mask, so the representation of every token depends on the whole sequence, left and right. The only mask is a padding mask that stops real tokens attending to pad positions. The cost of attention is quadratic in sequence length, and BERT caps inputs at 512 tokens. How attention computes those weights is covered in Attention Mechanism, in depth.

BERT pretraining: one encoder, two heads[CLS]mydog[MASK]cute[SEP]helikesplay[SEP]sum of token embedding + segment embedding (A or B) + learned position embeddingEmbedding layer: 30k WordPiece vocab, LayerNorm, dropoutTransformer encoder x Lbidirectional self-attention: every token attends to every token, no causal maskNSP head[CLS] -> pooler (tanh) -> IsNext?MLM headdense + GELU + LayerNorm -> tied vocab projectionloss only at the ~15% selected positions, e.g. predict 'is' at [MASK]binary loss on sentence pairsFine-tuning drops both heads and adds a small task head on [CLS] or on each token.
Inputs are summed embeddings; the encoder is fully bidirectional; pretraining attaches an MLM head to every selected position and an NSP head to [CLS].

From text to inputs

Text becomes input in four steps. First, WordPiece tokenization with a vocabulary of about 30,000 subwords splits rare words into pieces marked with a continuation prefix, so "playing" may become play and ##ing. Tokenizer design is covered in Tokenizers. Second, a [CLS] token is put first and a [SEP] token closes each segment, so a sentence pair becomes [CLS] A [SEP] B [SEP]. Third, each position gets three embeddings that are summed: the token embedding, a segment embedding saying whether it belongs to A or B, and a learned position embedding for positions 0 to 511. Fourth, the sum goes through LayerNorm and dropout into the encoder.

The final hidden state at [CLS] is meant to summarize the sequence for classification. That only works after fine-tuning: the raw pretrained [CLS] vector is a poor sentence embedding, a finding that motivated Sentence-BERT's siamese fine-tuning.

Masked language modelling

Masked language modelling picks 15% of the WordPiece positions in each sequence and asks the model to predict the original token at each, using a softmax over the vocabulary. The loss counts only those positions. The trick is what happens to a selected position: 80% of the time it is replaced by [MASK], 10% of the time by a random token, and 10% of the time left unchanged.

The reason is a mismatch between pretraining and use. [MASK] never appears in fine-tuning data, so a model trained only on masked inputs could learn to produce good representations only at [MASK] positions. Because a selected token may be unchanged or random, the model cannot tell which visible tokens are being tested, so it must keep a good contextual representation of every token. The random replacements are only 1.5% of all tokens, too few to hurt language understanding.

Here is the masking as a collate function, written the way most libraries implement it:

import torch

def mask_tokens(input_ids, tokenizer, p_select=0.15):
    labels = input_ids.clone()
    special = torch.tensor(
        [tokenizer.get_special_tokens_mask(ids, already_has_special_tokens=True)
         for ids in input_ids.tolist()], dtype=torch.bool)
    probs = torch.full(labels.shape, p_select)
    probs.masked_fill_(special | (input_ids == tokenizer.pad_token_id), 0.0)
    selected = torch.bernoulli(probs).bool()
    labels[~selected] = -100                      # ignored by cross-entropy

    to_mask = torch.bernoulli(torch.full(labels.shape, 0.8)).bool() & selected
    input_ids[to_mask] = tokenizer.mask_token_id

    to_rand = torch.bernoulli(torch.full(labels.shape, 0.5)).bool() & selected & ~to_mask
    input_ids[to_rand] = torch.randint(len(tokenizer), labels.shape)[to_rand]
    return input_ids, labels                      # remaining 10%: left unchanged

Note the 0.5: after 80% have been masked, half of the remaining 20% gives the 10% random share. Worked example: for my dog is cute, suppose is is selected. Four times out of five the input reads my dog [MASK] cute; once in ten it reads my dog apple cute; once in ten it is unchanged. In every case the label at that position is is and every other label is -100. The MLM head predicts it from the final hidden state with a dense layer, GELU and LayerNorm, then an output projection whose weights are tied to the input token embeddings.

Masking is expensive in one way: only 15% of positions give a learning signal, so BERT needs many steps. The original code also masked statically, once during preprocessing, duplicating the data ten times with different masks. Later work masks dynamically in the collate function as above, and Google's later release masked whole words rather than single subword pieces, which removes an easy shortcut of guessing ##ing from play.

Next sentence prediction, and why it went away

The second objective, next sentence prediction, built pairs where B is the true next segment 50% of the time and a random segment from the corpus otherwise, and trained a binary classifier on [CLS]. The motivation was tasks such as question answering and inference that reason about two texts.

It did not survive. RoBERTa (Liu et al., 2019) found that removing NSP and filling each input with contiguous full sentences matched or improved downstream results. The likely reason is that a random segment usually comes from a different document, so NSP mostly teaches topic detection, which MLM already covers. ALBERT replaced it with sentence-order prediction, two consecutive segments either in order or swapped, which forces the model to learn coherence rather than topic.

The pretraining recipe

The published pretraining recipe is worth knowing in numbers, because later models are described as changes to it. Data: BooksCorpus (800M words) and English Wikipedia (2,500M words, text passages only). Batch: 256 sequences of up to 512 tokens, about 128,000 tokens per batch, for 1,000,000 steps, roughly 40 epochs over the 3.3 billion words. Optimizer: Adam with learning rate 1e-4, beta1 0.9, beta2 0.999, L2 weight decay 0.01, 10,000 warm-up steps, then linear decay, and dropout 0.1 everywhere. To save compute, 90% of steps used sequences of length 128 and only the last 10% used length 512 to learn the later position embeddings. The decoupled weight-decay form most people use today is explained in AdamW Math.

RoBERTa showed this recipe was undertrained. Keeping the architecture and training longer on about 160 GB of text, with batches of 8,000 sequences, dynamic masking and no NSP, gave large gains. If you pretrain an encoder today, start from RoBERTa-style settings rather than the original ones.

What pretraining costs, and probing a checkpoint

The recipe also tells you what pretraining costs, which is worth computing before you decide to do it. A training step costs about 6 FLOPs per parameter per token, ignoring attention, which is small at these lengths. The short phase is 900,000 steps of 256 x 128 = 32,768 tokens, about 2.95e10 tokens; the long phase is 100,000 steps of 131,072 tokens, about 1.31e10. That is 4.26e10 tokens in total, and for BERT-base's 110M parameters about 6 x 1.1e8 x 4.26e10, or 2.8e19 FLOPs. At an achieved 400 TFLOP/s on one modern accelerator that is roughly 20 hours; BERT-large is about three times more. So reproducing BERT is affordable, but the result would be a 2018-quality model. Fine-tuning an existing checkpoint costs minutes, which is why it is the default.

Before fine-tuning anything, probe the checkpoint with its own pretraining head. The fill-mask pipeline runs the MLM head and returns the most likely tokens for the masked position:

from transformers import pipeline

fill = pipeline("fill-mask", model="bert-base-uncased")
for r in fill("The doctor told the patient to take the [MASK] twice a day."):
    print(f'{r["token_str"]:>12}  {r["score"]:.3f}')

Plausible completions with sensible probabilities show the tokenizer, checkpoint and head are wired correctly. Garbage usually means a mismatched tokenizer or a model loaded without its MLM head. The same probe is a cheap way to check domain fit: if the model cannot complete sentences from your domain, continued MLM pretraining on in-domain text before fine-tuning tends to help.

Fine-tuning

Fine-tuning adds one layer and trains everything end to end. For sentence or pair classification, a linear layer reads the [CLS] hidden state. For tagging, a linear layer reads every token's hidden state. For extractive question answering, two vectors score each token as answer start and end. The paper's search space is small: batch size 16 or 32, learning rate 5e-5, 3e-5 or 2e-5, and 2 to 4 epochs. With the Hugging Face libraries a classifier is a few lines:

from datasets import load_dataset
from transformers import (AutoTokenizer, AutoModelForSequenceClassification,
                          TrainingArguments, Trainer)

tok = AutoTokenizer.from_pretrained("bert-base-uncased")
model = AutoModelForSequenceClassification.from_pretrained("bert-base-uncased", num_labels=2)

ds = load_dataset("glue", "sst2")
ds = ds.map(lambda b: tok(b["sentence"], truncation=True, max_length=128), batched=True)

args = TrainingArguments(output_dir="out", learning_rate=2e-5, num_train_epochs=3,
                         per_device_train_batch_size=32, warmup_ratio=0.1,
                         weight_decay=0.01, eval_strategy="epoch", seed=42)
Trainer(model=model, args=args, train_dataset=ds["train"],
        eval_dataset=ds["validation"], processing_class=tok).train()

Fine-tuning BERT-large on small datasets is known to be unstable: some seeds diverge or stay at chance. Run several seeds, use warm-up, and keep the best on a validation set. For sentence similarity, use a model trained for embeddings rather than raw [CLS]; a worked application is in Sentence Similarity with BERT.

What came after

ModelWhat it changedWhy
RoBERTa (2019)no NSP, dynamic masking, more data, bigger batches, longer trainingBERT was undertrained
ALBERT (2019)factorized embeddings, cross-layer weight sharing, sentence-order predictionfewer parameters
ELECTRA (2020)a small generator corrupts tokens; the encoder classifies every token as original or replacedlearns from 100% of positions, not 15%
DeBERTa (2020)disentangled content and position attentionbetter use of relative position
ModernBERT (2024)RoPE, GeGLU, alternating local and global attention, 8,192-token contextmodern training tricks applied to an encoder

When choosing an encoder today, the original BERT checkpoints are rarely the best option; a RoBERTa, DeBERTa-v3 or ModernBERT checkpoint of the same size usually fine-tunes better, and the fine-tuning code is unchanged.

Failure modes

  • Truncation hides the answer. At a 512-token limit, long documents lose their tail. Use a sliding window with stride, or a long-context encoder.
  • Tokenizer and checkpoint mismatch. An uncased tokenizer with a cased model silently maps text to the wrong IDs. Always load both from the same checkpoint name.
  • Raw [CLS] used as an embedding. Retrieval quality is poor; use a model trained for sentence embeddings.
  • Masking special tokens or padding. A custom collate function that selects [CLS], [SEP] or pad positions wastes signal and can destabilize training.
  • Divergent fine-tuning runs. Too high a learning rate or no warm-up; lower it and run several seeds.
  • Using BERT to generate text. An MLM is not a left-to-right generator; use a decoder.

What to do next

  1. Tokenize a few sentences with a BERT tokenizer and inspect the WordPiece splits, special tokens and segment IDs.
  2. Run the mask function above on a batch and confirm about 15% of labels are not -100, split 80/10/10.
  3. Fine-tune bert-base-uncased on SST-2 with the script, three seeds, and record the spread.
  4. Repeat with a RoBERTa or ModernBERT checkpoint of similar size and compare accuracy and speed.
  5. For retrieval, swap in a sentence-embedding model instead of raw [CLS].
  6. Read the causal versus bidirectional comparison to see why decoders replaced BERT for generation.
Key takeaway: BERT is a Transformer encoder trained to fill in hidden tokens using context on both sides. Masked language modelling with the 80/10/10 rule is what made that possible; next sentence prediction turned out to be unnecessary. Fine-tuning adds one layer on top. Successors changed the data, the objective and the attention, but the recipe of pretrain bidirectionally, then fine-tune, is still how encoders are built.