BERT, from Devlin and colleagues at Google in 2018, showed that one pretrained Transformer encoder, fine-tuned with a single extra layer, could set the state of the art on question answering, natural language inference and named-entity tagging at once. Its key move was a training objective that let every token see context on both sides. A left-to-right language model cannot do that, because a token that can see the future can simply read off the answer. BERT avoided the problem by hiding the tokens it asks the model to predict.
Decoder models now dominate generation, but encoders in the BERT family still do much of the classification, retrieval and tagging in production, because they are small, fast and produce one contextual vector per token. This article explains how BERT turns text into inputs, how its two pretraining objectives work and which one later turned out not to matter, the exact recipe, how to fine-tune it, and what replaced each design choice. The mask itself is compared with the causal mask in Causal vs Bidirectional Attention.
The architecture
BERT is the encoder half of the original Transformer with no changes to the block: multi-head self-attention and a feed-forward layer, each wrapped in a residual connection followed by LayerNorm (the post-LN arrangement), with GELU in the feed-forward layer. The paper released two sizes.
| Model | Layers L | Hidden H | Heads A | Feed-forward | Parameters |
|---|---|---|---|---|---|
| BERT-base | 12 | 768 | 12 | 3,072 | 110M |
| BERT-large | 24 | 1,024 | 16 | 4,096 | 340M |
Attention has no causal mask, so the representation of every token depends on the whole sequence, left and right. The only mask is a padding mask that stops real tokens attending to pad positions. The cost of attention is quadratic in sequence length, and BERT caps inputs at 512 tokens. How attention computes those weights is covered in Attention Mechanism, in depth.
From text to inputs
Text becomes input in four steps. First, WordPiece tokenization with a vocabulary of about 30,000 subwords splits rare words into pieces marked with a continuation prefix, so "playing" may become play and ##ing. Tokenizer design is covered in Tokenizers. Second, a [CLS] token is put first and a [SEP] token closes each segment, so a sentence pair becomes [CLS] A [SEP] B [SEP]. Third, each position gets three embeddings that are summed: the token embedding, a segment embedding saying whether it belongs to A or B, and a learned position embedding for positions 0 to 511. Fourth, the sum goes through LayerNorm and dropout into the encoder.
The final hidden state at [CLS] is meant to summarize the sequence for classification. That only works after fine-tuning: the raw pretrained [CLS] vector is a poor sentence embedding, a finding that motivated Sentence-BERT's siamese fine-tuning.
Masked language modelling
Masked language modelling picks 15% of the WordPiece positions in each sequence and asks the model to predict the original token at each, using a softmax over the vocabulary. The loss counts only those positions. The trick is what happens to a selected position: 80% of the time it is replaced by [MASK], 10% of the time by a random token, and 10% of the time left unchanged.
The reason is a mismatch between pretraining and use. [MASK] never appears in fine-tuning data, so a model trained only on masked inputs could learn to produce good representations only at [MASK] positions. Because a selected token may be unchanged or random, the model cannot tell which visible tokens are being tested, so it must keep a good contextual representation of every token. The random replacements are only 1.5% of all tokens, too few to hurt language understanding.
Here is the masking as a collate function, written the way most libraries implement it:
import torch
def mask_tokens(input_ids, tokenizer, p_select=0.15):
labels = input_ids.clone()
special = torch.tensor(
[tokenizer.get_special_tokens_mask(ids, already_has_special_tokens=True)
for ids in input_ids.tolist()], dtype=torch.bool)
probs = torch.full(labels.shape, p_select)
probs.masked_fill_(special | (input_ids == tokenizer.pad_token_id), 0.0)
selected = torch.bernoulli(probs).bool()
labels[~selected] = -100 # ignored by cross-entropy
to_mask = torch.bernoulli(torch.full(labels.shape, 0.8)).bool() & selected
input_ids[to_mask] = tokenizer.mask_token_id
to_rand = torch.bernoulli(torch.full(labels.shape, 0.5)).bool() & selected & ~to_mask
input_ids[to_rand] = torch.randint(len(tokenizer), labels.shape)[to_rand]
return input_ids, labels # remaining 10%: left unchangedNote the 0.5: after 80% have been masked, half of the remaining 20% gives the 10% random share. Worked example: for my dog is cute, suppose is is selected. Four times out of five the input reads my dog [MASK] cute; once in ten it reads my dog apple cute; once in ten it is unchanged. In every case the label at that position is is and every other label is -100. The MLM head predicts it from the final hidden state with a dense layer, GELU and LayerNorm, then an output projection whose weights are tied to the input token embeddings.
Masking is expensive in one way: only 15% of positions give a learning signal, so BERT needs many steps. The original code also masked statically, once during preprocessing, duplicating the data ten times with different masks. Later work masks dynamically in the collate function as above, and Google's later release masked whole words rather than single subword pieces, which removes an easy shortcut of guessing ##ing from play.
Next sentence prediction, and why it went away
The second objective, next sentence prediction, built pairs where B is the true next segment 50% of the time and a random segment from the corpus otherwise, and trained a binary classifier on [CLS]. The motivation was tasks such as question answering and inference that reason about two texts.
It did not survive. RoBERTa (Liu et al., 2019) found that removing NSP and filling each input with contiguous full sentences matched or improved downstream results. The likely reason is that a random segment usually comes from a different document, so NSP mostly teaches topic detection, which MLM already covers. ALBERT replaced it with sentence-order prediction, two consecutive segments either in order or swapped, which forces the model to learn coherence rather than topic.
The pretraining recipe
The published pretraining recipe is worth knowing in numbers, because later models are described as changes to it. Data: BooksCorpus (800M words) and English Wikipedia (2,500M words, text passages only). Batch: 256 sequences of up to 512 tokens, about 128,000 tokens per batch, for 1,000,000 steps, roughly 40 epochs over the 3.3 billion words. Optimizer: Adam with learning rate 1e-4, beta1 0.9, beta2 0.999, L2 weight decay 0.01, 10,000 warm-up steps, then linear decay, and dropout 0.1 everywhere. To save compute, 90% of steps used sequences of length 128 and only the last 10% used length 512 to learn the later position embeddings. The decoupled weight-decay form most people use today is explained in AdamW Math.
RoBERTa showed this recipe was undertrained. Keeping the architecture and training longer on about 160 GB of text, with batches of 8,000 sequences, dynamic masking and no NSP, gave large gains. If you pretrain an encoder today, start from RoBERTa-style settings rather than the original ones.
What pretraining costs, and probing a checkpoint
The recipe also tells you what pretraining costs, which is worth computing before you decide to do it. A training step costs about 6 FLOPs per parameter per token, ignoring attention, which is small at these lengths. The short phase is 900,000 steps of 256 x 128 = 32,768 tokens, about 2.95e10 tokens; the long phase is 100,000 steps of 131,072 tokens, about 1.31e10. That is 4.26e10 tokens in total, and for BERT-base's 110M parameters about 6 x 1.1e8 x 4.26e10, or 2.8e19 FLOPs. At an achieved 400 TFLOP/s on one modern accelerator that is roughly 20 hours; BERT-large is about three times more. So reproducing BERT is affordable, but the result would be a 2018-quality model. Fine-tuning an existing checkpoint costs minutes, which is why it is the default.
Before fine-tuning anything, probe the checkpoint with its own pretraining head. The fill-mask pipeline runs the MLM head and returns the most likely tokens for the masked position:
from transformers import pipeline
fill = pipeline("fill-mask", model="bert-base-uncased")
for r in fill("The doctor told the patient to take the [MASK] twice a day."):
print(f'{r["token_str"]:>12} {r["score"]:.3f}')Plausible completions with sensible probabilities show the tokenizer, checkpoint and head are wired correctly. Garbage usually means a mismatched tokenizer or a model loaded without its MLM head. The same probe is a cheap way to check domain fit: if the model cannot complete sentences from your domain, continued MLM pretraining on in-domain text before fine-tuning tends to help.
Fine-tuning
Fine-tuning adds one layer and trains everything end to end. For sentence or pair classification, a linear layer reads the [CLS] hidden state. For tagging, a linear layer reads every token's hidden state. For extractive question answering, two vectors score each token as answer start and end. The paper's search space is small: batch size 16 or 32, learning rate 5e-5, 3e-5 or 2e-5, and 2 to 4 epochs. With the Hugging Face libraries a classifier is a few lines:
from datasets import load_dataset
from transformers import (AutoTokenizer, AutoModelForSequenceClassification,
TrainingArguments, Trainer)
tok = AutoTokenizer.from_pretrained("bert-base-uncased")
model = AutoModelForSequenceClassification.from_pretrained("bert-base-uncased", num_labels=2)
ds = load_dataset("glue", "sst2")
ds = ds.map(lambda b: tok(b["sentence"], truncation=True, max_length=128), batched=True)
args = TrainingArguments(output_dir="out", learning_rate=2e-5, num_train_epochs=3,
per_device_train_batch_size=32, warmup_ratio=0.1,
weight_decay=0.01, eval_strategy="epoch", seed=42)
Trainer(model=model, args=args, train_dataset=ds["train"],
eval_dataset=ds["validation"], processing_class=tok).train()Fine-tuning BERT-large on small datasets is known to be unstable: some seeds diverge or stay at chance. Run several seeds, use warm-up, and keep the best on a validation set. For sentence similarity, use a model trained for embeddings rather than raw [CLS]; a worked application is in Sentence Similarity with BERT.
What came after
| Model | What it changed | Why |
|---|---|---|
| RoBERTa (2019) | no NSP, dynamic masking, more data, bigger batches, longer training | BERT was undertrained |
| ALBERT (2019) | factorized embeddings, cross-layer weight sharing, sentence-order prediction | fewer parameters |
| ELECTRA (2020) | a small generator corrupts tokens; the encoder classifies every token as original or replaced | learns from 100% of positions, not 15% |
| DeBERTa (2020) | disentangled content and position attention | better use of relative position |
| ModernBERT (2024) | RoPE, GeGLU, alternating local and global attention, 8,192-token context | modern training tricks applied to an encoder |
When choosing an encoder today, the original BERT checkpoints are rarely the best option; a RoBERTa, DeBERTa-v3 or ModernBERT checkpoint of the same size usually fine-tunes better, and the fine-tuning code is unchanged.
Failure modes
- Truncation hides the answer. At a 512-token limit, long documents lose their tail. Use a sliding window with stride, or a long-context encoder.
- Tokenizer and checkpoint mismatch. An uncased tokenizer with a cased model silently maps text to the wrong IDs. Always load both from the same checkpoint name.
- Raw [CLS] used as an embedding. Retrieval quality is poor; use a model trained for sentence embeddings.
- Masking special tokens or padding. A custom collate function that selects
[CLS],[SEP]or pad positions wastes signal and can destabilize training. - Divergent fine-tuning runs. Too high a learning rate or no warm-up; lower it and run several seeds.
- Using BERT to generate text. An MLM is not a left-to-right generator; use a decoder.
What to do next
- Tokenize a few sentences with a BERT tokenizer and inspect the WordPiece splits, special tokens and segment IDs.
- Run the mask function above on a batch and confirm about 15% of labels are not -100, split 80/10/10.
- Fine-tune bert-base-uncased on SST-2 with the script, three seeds, and record the spread.
- Repeat with a RoBERTa or ModernBERT checkpoint of similar size and compare accuracy and speed.
- For retrieval, swap in a sentence-embedding model instead of raw [CLS].
- Read the causal versus bidirectional comparison to see why decoders replaced BERT for generation.