Mixture-of-experts (MoE) replaces a transformer's single feed-forward block with many smaller ones, the experts, and a router that sends each token to a few of them. A model can then hold many parameters while spending the compute of a much smaller one per token. For frontier models the motivation is training cost. For small models, the ones meant to run on a laptop, a phone or a single modest GPU, the arithmetic changes: the binding constraints become how much memory the device has and how fast it can stream weights, usually at batch size one.

This article assumes the router and load-balancing maths from the MoE maths article and focuses on what changes at small scale: how open small MoEs are actually shaped, how to balance experts without hurting quality, how to build one from a dense checkpoint, what it costs to run on a device, and how to fine-tune one without breaking the router.

Advertisement

Why small changes the MoE trade-off

A dense model's parameters are all used for every token. An MoE model has total parameters, which must all be stored, and active parameters, which are read and multiplied for each token. On a data-centre cluster, total parameters are spread over many GPUs and the win is fewer FLOPs per token. On a single device, every parameter must fit in local memory, so an MoE with 7 billion total parameters needs the memory of a 7B dense model even if it runs like a 1B model.

What it buys you on the device is speed at batch one. Autoregressive decoding of one sequence is limited by memory bandwidth: to produce each token the device must read every active weight once. An MoE reads only the chosen experts, so decode speed tracks active parameters, while quality tracks something closer to total parameters. Small MoE is therefore a bet that you have spare memory capacity but limited bandwidth and compute, which describes many laptops and unified-memory devices well, and phones less well.

Prefill, processing a long prompt, is different. Hundreds of tokens routed in parallel touch nearly every expert, so prefill reads essentially all weights, but it is compute-bound anyway and its cost per token follows active FLOPs.

How small MoEs are actually shaped

Early MoEs used a few large experts, each the size of the dense FFN, with top-1 or top-2 routing. Small open MoEs have mostly moved to fine-grained experts: many small experts with more of them active. Splitting each expert into smaller pieces and activating proportionally more keeps compute constant but gives the router far more combinations, which the DeepSeekMoE work showed improves specialisation. Many designs also add shared experts that every token passes through, so common knowledge lives in one place instead of being duplicated across routed experts.

Two open examples show the range. OLMoE-1B-7B uses 64 small experts per MoE layer with 8 active per token, for 6.9B total and 1.3B active parameters, trained from scratch on 5 trillion tokens with all data, code and logs released. Qwen1.5-MoE-A2.7B has 14.3B total and 2.7B active parameters, routes each token to 4 of 60 experts plus 4 shared experts that are always on, and was upcycled from the dense Qwen-1.8B rather than trained from scratch. The SLM architectures comparison places MoE alongside the other axes of small-model design, such as attention variant and vocabulary size.

A small fine-grained MoE block with a shared expertToken hidden hafter attentionRouterscores + balance biasTop-k selectk of 64, bias only hereShared expertalways onRouted experts (small FFNs)E0E1E2E3E4E5E6E7E8E9E10E11E12E13E14E15... 64 in total, k active per tokenWeighted sumgates from raw scores+ residual, next layer
One small MoE layer: the router scores all experts, a balance bias influences only which k experts are selected, the selected experts' outputs are weighted by the raw scores and added to the always-on shared expert, then the residual.
Advertisement

A small MoE layer in PyTorch

The module below is a readable reference implementation, not a fast kernel. It uses sigmoid router scores, a balance bias for selection, gate weights from the raw scores, a shared expert and SwiGLU experts.

import torch
import torch.nn as nn
import torch.nn.functional as F

class Expert(nn.Module):
    def __init__(self, d, hidden):
        super().__init__()
        self.w_gate = nn.Linear(d, hidden, bias=False)
        self.w_up = nn.Linear(d, hidden, bias=False)
        self.w_down = nn.Linear(hidden, d, bias=False)

    def forward(self, x):                                # SwiGLU FFN
        return self.w_down(F.silu(self.w_gate(x)) * self.w_up(x))

class SmallMoE(nn.Module):
    def __init__(self, d=1024, n_experts=64, k=8, expert_hidden=256,
                 shared_hidden=1024, bias_lr=1e-3):
        super().__init__()
        self.k, self.n, self.bias_lr = k, n_experts, bias_lr
        self.router = nn.Linear(d, n_experts, bias=False)
        self.experts = nn.ModuleList(Expert(d, expert_hidden) for _ in range(n_experts))
        self.shared = Expert(d, shared_hidden)
        self.register_buffer("balance_bias", torch.zeros(n_experts))   # not a parameter

    def forward(self, x):                                # x: [tokens, d]
        scores = torch.sigmoid(self.router(x))           # affinity per expert
        _, idx = torch.topk(scores + self.balance_bias, self.k, dim=-1)  # bias picks...
        gates = scores.gather(-1, idx)                   # ...but raw scores weight
        gates = gates / gates.sum(-1, keepdim=True)
        out = self.shared(x)
        for e in idx.unique().tolist():                  # loop over experts, not tokens
            tok, slot = (idx == e).nonzero(as_tuple=True)
            y = gates[tok, slot, None] * self.experts[e](x[tok])
            out = out.index_add(0, tok, y)               # out-of-place: autograd-safe
        if self.training:
            with torch.no_grad():                        # balance outside the gradient
                load = torch.bincount(idx.flatten(), minlength=self.n).float()
                self.balance_bias += self.bias_lr * torch.sign(load.mean() - load)
        return out

Two design points matter. The loop runs over experts, gathering each expert's tokens into one batch, because a loop over tokens would launch thousands of tiny matrix multiplies. Production implementations go further and sort tokens by expert into one grouped GEMM. And the balance bias lives in a buffer updated under torch.no_grad(): it steers selection without creating a gradient that fights the language-modelling loss.

Balancing without fighting the loss

Routers left alone collapse: a few experts win early, get more training, win more, and the rest go idle. The classic remedy, an auxiliary load-balancing loss, is derived in the maths article; its weakness is that it adds a gradient unrelated to language modelling, and too large a coefficient measurably hurts quality while too small fails to balance.

DeepSeek-V3 popularised an alternative often called auxiliary-loss-free balancing, as in the code above. Each expert gets a bias that is added to its score only when choosing the top-k. After each step, overloaded experts' biases go down and underloaded ones go up by a small fixed amount. The gate weights that scale expert outputs still come from the unbiased scores, so the model's function is not distorted, only the routing decisions. DeepSeek-V3 still kept a very small sequence-level balance loss as a safeguard against extreme imbalance within one sequence.

At small scale balance matters for a second reason: on a device, an overloaded expert is not a straggler on another GPU but simply wasted capacity, parameters you store and never use. Track per-expert token share through training; a healthy fine-grained model shows no expert near zero and no expert far above its fair share.

Building one from a dense model: sparse upcycling

Training an MoE from scratch costs the full pretraining budget. Sparse upcycling, introduced by Komatsuzaki and colleagues, starts from a trained dense checkpoint instead: copy each FFN into N experts, add a freshly initialised router, and continue training. The attention layers, embeddings and norms carry over unchanged. Qwen1.5-MoE-A2.7B is an upcycled model, reported to reach its quality at a fraction of the compute of training a comparable dense model.

  1. Choose which layers become MoE; many designs convert every FFN, some only every other layer to save memory.
  2. For fine-grained experts, split each dense FFN's hidden dimension into segments, so four 256-wide experts come from one 1,024-wide FFN, then replicate segments to reach the expert count.
  3. Break symmetry. Identical copies receive identical gradients if routing is uniform; add small noise to expert weights or re-initialise part of each expert.
  4. Initialise the router small so early routing is near uniform, then continue pretraining with balancing active from the first step.
  5. Evaluate against the dense parent at equal extra training tokens, not just at the end, to confirm the MoE is earning its memory.

Worked example: the batch-1 budget on a device

Suppose a device has 8 GB of usable memory and about 100 GB/s of memory bandwidth, a plausible laptop or unified-memory figure. With 4-bit weights, about half a byte per parameter, the snippet below computes the memory footprint and the bandwidth ceiling on decode speed for three models.

def batch1_decode_ceiling(total_params, active_params, bytes_per_param, mem_gb_s):
    footprint_gb = total_params * bytes_per_param / 1e9      # must fit in memory
    per_token_gb = active_params * bytes_per_param / 1e9     # weights streamed per token
    return footprint_gb, mem_gb_s / per_token_gb             # upper bound, tokens/s

# 4-bit weights (~0.5 bytes/param), 100 GB/s memory bandwidth, ignoring KV cache
for name, total, active in [("dense 1.3B", 1.3e9, 1.3e9),
                            ("dense 7B",   7.0e9, 7.0e9),
                            ("MoE 6.9B/1.3B", 6.9e9, 1.3e9)]:
    fp, tps = batch1_decode_ceiling(total, active, 0.5, 100)
    print(f"{name:14s} footprint {fp:4.2f} GB  ceiling {tps:5.0f} tok/s")
# dense 1.3B     footprint 0.65 GB  ceiling   154 tok/s
# dense 7B       footprint 3.50 GB  ceiling    29 tok/s
# MoE 6.9B/1.3B  footprint 3.45 GB  ceiling   154 tok/s

The MoE fits in the same memory as the dense 7B model but has the decode ceiling of the 1.3B model, and its quality in published comparisons sits well above dense models with the same active parameters. The dense 1.3B model is smallest, and the dense 7B is slowest. These are ceilings: attention reads the KV cache too, routing and scattered expert access cost efficiency, and embedding tables count toward active parameters. Measured speed is typically well below the ceiling, but the ratios hold.

If the device has only 3 GB free, the MoE no longer fits and the picture flips: the dense 1.3B model or a distilled model is the better choice. See the SLM quantisation guide for the bits-per-weight trade-offs behind the half-byte assumption.

Running experts that do not fit: offloading and quantisation

When total parameters exceed fast memory, runtimes keep attention, embeddings and shared experts resident and page routed experts from slower memory, for example from CPU RAM to GPU memory or from flash to RAM. This works only when routing has locality: consecutive tokens reusing recent experts. Locality exists but is limited, so an expert cache with least-recently-used eviction often hits well within a topic and poorly at boundaries, and each miss costs a transfer far slower than compute. Offloading is a way to run a model that does not fit, not a way to make it fast.

Quantisation fits MoE unusually well. Each routed expert sees a fraction of tokens, and the shared expert and attention see all of them, so a common split keeps shared and attention weights at higher precision and quantises routed experts more aggressively. Calibrate quantisation with enough data to activate every expert; an expert that saw no calibration tokens gets poor scales and silently degrades the rare inputs that route to it.

Batching changes the economics. At batch one, decode reads k experts per layer; with sixteen concurrent sequences it may read most of them, and the bandwidth advantage over a dense model of equal total size shrinks. Small MoEs shine for single-user local inference, and the MoE serving article covers the multi-GPU, many-user case.

Fine-tuning a small MoE

Fine-tuning on a narrow dataset can shift routing sharply: a few experts suited to the new domain absorb most tokens and the rest drift. Three habits keep it stable. Freeze the router or give it a much lower learning rate than the experts for the first run. Keep balancing active, with bias updates or a small auxiliary loss. And log per-expert token share and router entropy alongside the loss, because collapse shows there before it shows in evaluation scores.

Parameter-efficient methods need thought. LoRA on attention and shared experts is cheap and usually enough. LoRA on every routed expert multiplies adapter count by the number of experts and gives each adapter only its share of tokens. Distilling a small MoE into a dense student, covered in the distillation article, is an option when the target device cannot hold the MoE's total parameters.

Failure modes

SymptomCauseFix
A few experts take most tokensRouter collapse, balancing too weakBias balancing or raise aux coefficient; check init
Quality drop after adding aux lossCoefficient too largeLower it or switch to bias-based balancing
Upcycled experts stay identicalNo symmetry breakingAdd noise or partial re-initialisation
Local decode far below expectationsExperts paged from slow memoryFit the model in fast memory or use a dense model
Rare topics degrade after quantisingExperts missing from calibration dataBroader calibration set; check per-expert coverage
Fine-tune loss fine, eval collapsesRouting shifted to few expertsFreeze or slow router; monitor entropy
Slow reference implementationPer-token expert loopGroup tokens per expert; grouped GEMM kernels

Trade-offs: small MoE or dense

SituationBetter choiceWhy
Memory to spare, bandwidth-bound, single userSmall MoEDecode speed of active size, quality nearer total size
Tight memory budgetDense model of that sizeMoE total parameters must still fit
Many concurrent users on one GPUDense, or MoE with careBatching erodes the per-token bandwidth saving
You own a good dense checkpointUpcycleReuses pretraining; cheaper than scratch
Heavy domain fine-tuning plannedDense, or MoE with frozen routerDense fine-tuning is simpler to stabilise

What to do next

  1. Write down your target device's usable memory and bandwidth, and compute footprint and decode ceiling for each candidate model, as in the worked example.
  2. Run OLMoE-1B-7B or another open small MoE locally and measure tokens per second against a dense model of equal active size.
  3. Log per-expert token share on your own prompts to see how balanced routing really is on your workload.
  4. If you quantise, check that every expert receives calibration tokens and compare accuracy on rare-topic prompts.
  5. Before fine-tuning, freeze the router for a first run and track router entropy alongside loss.
  6. If memory is the constraint, compare against distilling into a dense student before committing to MoE.
Key takeaway: At small scale an MoE trades memory for speed: every expert must fit on the device, but decoding reads only the active ones, so it runs like its active size while scoring closer to its total size. Current small MoEs use many fine-grained experts plus shared experts, balance with selection biases rather than heavy auxiliary losses, and are often upcycled from dense checkpoints. Budget memory and bandwidth first, quantise with full expert coverage, and keep the router stable when fine-tuning.