Most teams that ship a small language model never pretrain one; they fine-tune or distill an open base model, and that is usually the right call. Pretraining from scratch earns its cost in a narrower set of cases: the target domain uses a vocabulary or script that existing tokenizers handle badly, licensing requires full control over the training data, the model must fit an unusual shape such as a very small on-device budget or a long context at tiny width, or the point is research into the training recipe itself. When one of those applies, the hard part is rarely writing a transformer; it is freezing a chain of early decisions in the right order and building a pipeline whose stages hand clean artifacts to each other.

This article covers the architecture decisions and the pipeline as a whole, for models from roughly 100 million to 1 billion parameters. The data pipeline gets its own treatment in Pretraining Data for Small Language Models, and the inner training loop, including checkpoints and loss-spike recovery, in The SLM Pretraining Loop in Pseudocode.

Advertisement

Start from the budget, not the architecture

The first decision is how much compute you have and how the model will be used, because those two facts pin the parameter count N and the token count D. Training compute for a dense decoder is well approximated by C ≈ 6·N·D floating-point operations: two for the forward multiply-add per parameter per token, four for the backward pass. A 350M-parameter model trained on 100B tokens needs about 6 × 3.5e8 × 1e11 = 2.1e20 FLOPs. If your accelerators sustain an effective 400 TFLOP/s each after utilization losses, that is about 525,000 accelerator-seconds, roughly 146 accelerator-hours, before restarts, evaluation and ablations, which together commonly add a third or more.

Compute-optimal scaling results suggest on the order of 20 tokens per parameter minimizes loss for a fixed training budget. Small models are almost never trained there, and deliberately so. A small model is chosen because inference is cheap, and over its lifetime it will serve far more tokens than it trains on, so the sensible trade is to spend extra training compute to push a small model's quality up rather than to train a larger model that costs more on every request. Public small models are routinely trained at hundreds or thousands of tokens per parameter; returns diminish but keep coming. Decide the target D early, because it sets how much data the pipeline in article two must deliver and how much repetition you will tolerate.

Architecture choices that matter at small scale

The mainstream decoder recipe has converged, and for a from-scratch small model there are few reasons to deviate: a decoder-only transformer, pre-normalization with RMSNorm, rotary position embeddings, a gated feed-forward block such as SwiGLU, no bias terms in linear layers, and a final norm before the output projection. SwiGLU adds a third weight matrix, which is why its hidden size is usually set near 8/3 of the model width instead of 4 times.

Four choices deserve deliberate thought at small scale. The first is grouped-query attention. Using fewer key-value heads than query heads shrinks the KV cache, which is the dominant memory cost for long-context inference on small devices, and costs little quality; a ratio of four query heads per key-value head is a common middle ground. The second is tied input and output embeddings. At 1B parameters and above untied embeddings are cheap relative to the whole; at 150M they are not, and tying them frees a large fraction of the parameter budget for transformer layers. The third is depth versus width. For a fixed parameter count, small models tend to benefit from being somewhat deeper and narrower than a naive scaled-down large model, though very deep and very narrow stacks train slowly because each layer adds sequential latency and tiny matrices use accelerators poorly. Keep the head dimension at 64 or 128 so kernels stay efficient, and pick a width that is a multiple of 128.

The fourth is vocabulary size, and it is where small models differ most from large ones. An embedding matrix costs vocabulary times width parameters. The next section makes the arithmetic concrete, but the short version is that a large multilingual vocabulary can consume a third or more of a 150M model, spending capacity on rare tokens that are seen too few times to learn well. Choose the vocabulary size together with the tokenizer's compression on your data, not by copying a large model's.

Advertisement

Parameter counting before you commit

Write the parameter count as a function of the config and print the breakdown before training anything. It catches mistakes such as an accidental untied head, a feed-forward size that is not a multiple of the tile size, or a vocabulary that eats the budget. The same function gives the non-embedding parameter count, which is the number scaling relationships care about, and the per-token FLOPs used for utilization accounting later.

def param_count(cfg):
    d, L, V = cfg.d_model, cfg.n_layers, cfg.vocab_size
    hd = d // cfg.n_heads                         # head dim, keep at 64 or 128
    kv = cfg.n_kv_heads * hd                      # GQA: fewer K/V heads
    attn = d * d + 2 * d * kv + d * d             # Wq, Wk, Wv, Wo
    ffn = 3 * d * cfg.ffn_hidden                  # SwiGLU: gate, up, down
    norms = 2 * d                                 # two RMSNorm gains per layer
    per_layer = attn + ffn + norms
    embed = V * d
    head = 0 if cfg.tie_embeddings else V * d
    total = L * per_layer + embed + head + d      # + final norm
    return dict(total=total, non_embedding=L * per_layer,
                embed_share=(embed + head) / total)

cfg = Config(d_model=768, n_layers=24, n_heads=12, n_kv_heads=4,
             ffn_hidden=2048, vocab_size=32768, tie_embeddings=True,
             context=2048)
print(param_count(cfg))   # about 176M total, 151M non-embedding, embed share ~14%

With these numbers a 32,768-token vocabulary costs about 25M parameters at width 768. Switch to a 128K vocabulary and the embedding alone is about 100M, two-thirds the size of the 151M transformer stack, and untying the head would double it. That does not mean small vocabularies always win: a larger vocabulary compresses text into fewer tokens, which lowers the compute per document and effectively lengthens the context. Measure bytes per token on a held-out sample of your mixture for two or three candidate sizes and pick the knee, where compression stops improving much per added row.

Tokenizer: train it on the mixture, then freeze it

Train the tokenizer on a sample drawn with the same domain weights you plan to train with, not on whatever corpus is closest to hand, because a tokenizer trained mostly on web English will fragment code and other languages into many tokens and quietly shrink their effective share of training. Byte-level BPE or a unigram model with byte fallback guarantees that any input can be encoded. Split digits into single tokens so arithmetic sees consistent units, keep whitespace runs that matter for code, and reserve special tokens you do not need yet, including chat-role markers, tool-call delimiters and a handful of spares, because adding tokens after pretraining means resizing the embedding and training new rows from nothing.

Then freeze it. The tokenizer's content hash becomes part of every downstream artifact's identity: token shards, checkpoints and evaluation caches. Changing the tokenizer after shards exist invalidates all of them, and a mismatch between the tokenizer used for shards and the one shipped with the model is one of the most expensive silent bugs in the pipeline, because the model still produces fluent-looking output on common tokens.

The pipeline as stages with artifacts

Stages, and the artifact each one hands to the nextBudget and configN, D, context, vocabData pipelinefiltered, deduped docsTokenizertrained, then frozenToken shardspacked, index, hashesSmoke testsoverfit, init loss, MFUProxy ablations5-10% size, same codeMain runcheckpoints, eval hooksCooldown / annealLR decay, best dataBase model evalloss by domain, tasksRelease bundleweights + tokenizerPost-trainingSFT, preference tuningGate failed?fix upstream, rerunEvery arrow carries a versioned artifact with a content hash; a stage never reads another stage's scratch space.Config hash + tokenizer hash + shard manifest hash identify a run; a checkpoint without all three is not resumable.
The from-scratch pipeline. Each stage consumes and produces a versioned, hashed artifact, and a failed gate sends work back upstream instead of patching downstream.

It helps to treat pretraining as a build system. Each stage has declared inputs, produces an immutable output with a manifest, and has a gate that must pass before the next stage starts.

def run_pipeline(plan):
    raw     = stage("collect",   plan.sources)                   # snapshots + licenses
    docs    = stage("filter",    raw, plan.filters)              # article 2
    tok     = stage("tokenizer", sample(docs, plan.mix), plan.vocab_size)
    shards  = stage("pack",      docs, tok, plan.mix, plan.context)
    gate(smoke_tests(plan.cfg, shards, tok))                     # below
    choice  = stage("ablate",    proxy_runs(plan, shards))       # mix, peak LR
    ckpts   = stage("pretrain",  plan.cfg.with_(choice), shards) # article 3
    base    = stage("cooldown",  ckpts.last_stable(), shards.high_quality())
    gate(evaluate_base(base, tok, plan.eval_suite))
    return bundle(base, tok, plan.cfg, shards.manifest_hash)

def stage(name, *inputs):
    key = hash_of(name, *[i.content_hash for i in inputs])
    if store.exists(key):             # rerun is a no-op when nothing changed
        return store.get(key)
    out = STAGES[name](*inputs)
    store.put(key, out, manifest=describe(out))
    return out

Content-addressed stages make reruns cheap and make provenance automatic: a base model can be traced back to the exact shard manifest, tokenizer and filter versions that produced it. They also enforce the most important discipline, which is that a problem found downstream is fixed upstream. If evaluation shows the model has memorized a benchmark, the fix belongs in decontamination, and the pipeline recomputes everything that depended on it rather than someone patching a shard by hand.

Smoke tests before the expensive run

A week of compute is too costly to spend discovering a bug that ten minutes could have found. Before the main run, a short battery of checks confirms the model, data path and loop behave. None of them prove the run will succeed; all of them catch failures that are otherwise found late.

  • Initial loss. With standard initialization, loss at step zero should be close to ln(V), about 10.4 for a 32K vocabulary. Much higher means the output layer's initialization or scaling is wrong.
  • Overfit one batch. Repeating a single batch should drive loss close to zero within a few hundred steps. If it does not, gradients are not reaching some parameters or the labels are misaligned by one position.
  • Label shift and masking. Decode a packed sequence and its targets side by side and confirm targets are inputs shifted by one and that document boundaries behave as intended.
  • Throughput and utilization. Measure tokens per second and model FLOPs utilization at the real batch shape; a figure far below comparable setups points to a data loader bottleneck or unfused kernels.
  • Resume equivalence. Train 200 steps, checkpoint at 100, resume and compare losses from step 101 onward; they should match closely. Article three explains what the checkpoint must contain for this to hold.

Proxy ablations: decide cheaply, then commit

Some decisions cannot be settled by reasoning: the peak learning rate, the data mixture weights, whether a filter helps. Settle them with proxy runs, models perhaps a tenth the size of the target trained on a proportionally reduced token budget, using the exact same code, tokenizer and shards. Change one thing per ablation and run two seeds for close calls, because a small model's seed-to-seed spread is often as large as the effect being tested.

Mixture and filter rankings usually transfer from proxy to target reasonably well. The optimal learning rate does not transfer directly across widths under standard parameterization; it generally falls as the model widens. Sweep at two proxy sizes and extrapolate, or use a parameterization designed for width transfer, then confirm with a short full-size run.

Cooldown and the base model

Many current recipes end pretraining with a cooldown, a phase where the learning rate decays to a small value while the mixture shifts toward the highest-quality data: curated text, code, math and instruction-like material that is still not chat formatted. A schedule with a long constant phase and a separate decay, described in article three, lets you branch several cooldowns from one checkpoint and compare them, instead of committing to a mixture for the entire run.

The base model is then evaluated as a base model: held-out loss by domain, next-token tasks scored by likelihood rather than generation, long-context retrieval at the trained length, and a contamination report. Instruction-following benchmarks are not meaningful yet; that is post-training's job.

Handing off to post-training

The release bundle is what the post-training team consumes, and it should be complete enough that nobody has to ask how the model was made. It contains the weights, the frozen tokenizer with its reserved special tokens documented, the model config, the trained context length and RoPE settings, the shard manifest hash and mixture, the evaluation report, and known weaknesses. Reserved tokens matter here: if the chat template's role markers were reserved during pretraining, supervised fine-tuning can use them directly, and their embedding rows exist, even if they were never seen.

Failure modes

  • Copying a large model's vocabulary. The embedding eats the parameter budget and rare tokens stay undertrained.
  • Tokenizer drift. Shards and the shipped model use different tokenizer versions; output looks fluent but degrades.
  • Transferring learning rate by assumption. A rate tuned on a proxy diverges or underfits at full width.
  • No resume test. The first hardware failure reveals that data order or schedule state was not checkpointed.
  • Stages sharing scratch space. A manual fix in one stage's output is overwritten or never propagated, and nobody can reproduce the model.
  • Chasing chat benchmarks on a base model. Mixture decisions get made on tasks the base model cannot yet do, so the signal is noise.

Trade-offs

Pretraining from scratch buys control over tokenizer, data, architecture and licensing at a cost measured in compute, engineering time and weeks of calendar. Continued pretraining of an open base model on domain data buys much of the domain benefit for a fraction of the cost, but inherits the base model's tokenizer and data provenance. Choose from-scratch only when the tokenizer, data rights or architecture genuinely have to be yours, and then invest in the pipeline discipline above.

Key takeaway: Pin parameters and tokens from the compute budget and the inference economics, expect to train far past 20 tokens per parameter, use the converged decoder recipe with grouped-query attention and tied embeddings, count parameters so the vocabulary does not eat a small model, train the tokenizer on the mixture and freeze it, run the pipeline as hashed stages with gates, spend ten minutes on smoke tests before spending weeks on the run, settle mixture and learning rate with proxy ablations, and hand post-training a complete bundle.