Most of the strongest open-weight language models released since late 2024 are mixtures of experts. DeepSeek-V3, Qwen3's largest models, Llama 4, gpt-oss and Kimi K2 all replace most dense feed-forward layers with many smaller expert networks and a router that sends each token to a few of them. If you choose, fine-tune or serve models in 2026, you are making MoE decisions whether you intend to or not.

This article explains what MoE buys from first principles, lays out how the major open MoE models are actually configured, read from their published configuration files on 2026-10-03, draws the design trends from that table, and works through the arithmetic that decides whether an MoE fits your hardware and workload. It closes with fine-tuning guidance, failure modes and a checklist. The table covers a selection of releases from December 2024 to 2026, not every model; check any model it omits the same way, from its own config file.

One MoE layer in one picture

One MoE layer, fine-grained routing with a shared expertToken hidden hafter attentionRouterscores s = sigma(Wh)Bias b_iselection onlyShared expertalways onExpert 17selectedExpert 102selectedExpert 3..256idle for hExpert 230selectedWeighted sumgates x outputstop-k on s+bCompute per token follows the selected experts; memory must hold all of them.
A router scores every expert, a balancing bias influences only which experts are chosen, the selected experts and a shared expert process the token, and their outputs are combined with gate weights.

What MoE buys, from first principles

A dense transformer applies every parameter to every token, so training and inference FLOPs scale with parameter count: roughly 2N FLOPs per token for a forward pass through N parameters. MoE breaks that link. Each MoE layer holds E expert feed-forward networks, and a learned router picks k of them per token. The model's total parameters set its capacity to store knowledge; its active parameters, the shared layers plus k experts per MoE layer, set the compute per token.

That is why gpt-oss-120b can have 116.8 billion parameters and only 5.1 billion active per token, according to its model card. The catch is that the hardware still has to hold all 116.8 billion, and serving many users at once eventually touches nearly all of them. MoE trades memory capacity and communication for compute. Whether that trade is good depends on which resource you are short of.

The router is the new moving part. It is usually a single linear layer that produces one score per expert for each token; the top-k scores pick the experts, and normalised scores weight their outputs. Because the choice is discrete, the router learns only through the gate weights of the experts it picked, so an expert that is never chosen never receives gradient and never improves. Left alone, training drifts towards a few popular experts. Every production MoE therefore needs some balancing mechanism, and how that mechanism works is one of the main ways the 2026 models differ from their predecessors. The other differences are in how many experts there are, how large each one is, and whether some capacity is shared by every token.

The open MoE landscape, read from the configs

The table below comes from each model's own configuration file on Hugging Face, except where noted. Total and active counts are from the developers' reports or model cards.

Model (release)Routed experts, top-kShared expertLayersTotal / active
DeepSeek-V3 (Dec 2024)256, top-8161, first 3 dense671B / 37B
Qwen3-235B-A22B (Apr 2025)128, top-8none94, all MoE235B / 22B
Llama 4 Maverick (Apr 2025)128, top-11alternating dense and MoE400B / 17B
Kimi K2 (Jul 2025)384, top-8161, first 1 denseabout 1T / 32B
gpt-oss-120b (Aug 2025)128, top-4none36, all MoE116.8B / 5.1B
Qwen3-Next-80B-A3B (Sep 2025)512, top-10148, hybrid attention80B / about 3B
DeepSeek-V4-Flash (2026)256, top-6143, hybrid compressed attention284B / 13B
Kimi K3 (2026)896, top-16293, 69 linear + 24 MLA2.8T / 104B

Llama 4 Maverick's configuration is gated, so its row comes from Meta's published description. DeepSeek-V4 also has a Pro variant, which its model card gives as 1.6T total and 49B active. For comparison, Mixtral 8x7B from December 2023 used 8 large experts with top-2 routing. The distance from 8 experts to 896 in under three years is the headline story of MoE design.

Six design trends

Experts got smaller and more numerous. DeepSeek-V3's experts have an intermediate size of 2,048 against a hidden size of 7,168; Qwen3-Next uses 512 experts of intermediate size 512. Fine-grained experts give the router more combinations, letting each expert specialise more narrowly at the same active compute.

Shared experts absorb common knowledge. DeepSeek-V3, Kimi K2, Llama 4 Maverick and Qwen3-Next route every token through one always-on expert, so routed experts do not each relearn the basics. Qwen3-235B and gpt-oss show this is not mandatory.

Balancing moved out of the loss. Older MoEs added an auxiliary loss that penalised uneven expert load, which competes with the language-modelling objective. DeepSeek-V3's config records topk_method: noaux_tc: a per-expert bias is added to router scores only when choosing the top-k, and nudged down for overloaded experts and up for idle ones after each step. The gate weights that scale the expert outputs come from the unbiased scores. Qwen3-235B still uses a small auxiliary loss, coefficient 0.001 in its config.

Routing respects the network. DeepSeek-V3 groups its 256 experts into 8 groups and lets each token pick from only 4 of them (n_group 8, topk_group 4), which caps how many machines a token's activations must travel to during expert-parallel training and serving.

Experts are stored in low precision. gpt-oss ships its MoE weights in MXFP4, about 4.25 bits per parameter, while the router, attention and embeddings stay in higher precision. Because experts are over 90 percent of the parameters, that is what lets the 120b model fit a single 80 GB GPU.

Attention is changing underneath. Qwen3-Next pairs its MoE with a hybrid stack that uses full attention only every fourth layer, and gpt-oss alternates sliding-window (128 tokens) and full-attention layers. The 2026 models go further: DeepSeek-V4's model card describes a hybrid of two compressed attention mechanisms for its one-million-token context, and Kimi K3's lists 69 Kimi Delta Attention layers, a linear-attention design, alongside 24 gated MLA layers. MoE cuts feed-forward cost; long context needs attention to get cheaper too.

Routing as it is done now, in code

The full MoE block is covered in MoE Transformer Architecture, in depth. Here is the part that changed most: bias-based, group-limited top-k selection in the style of DeepSeek-V3, simplified to the essentials.

import torch

def route(h, W, bias, k=8, n_group=8, topk_group=4, scale=2.5):
    # h: [T, d] tokens, W: [E, d] router weights, bias: [E] balancing bias
    s = torch.sigmoid(h @ W.T)                     # affinity scores, [T, E]
    sel = s + bias                                 # bias affects choice only
    T, E = sel.shape
    g = sel.view(T, n_group, E // n_group)
    group_score = g.topk(2, dim=-1).values.sum(-1) # rank groups by their best experts
    keep = group_score.topk(topk_group, dim=-1).indices
    mask = torch.zeros(T, n_group, device=h.device).scatter_(1, keep, 1.0)
    sel = sel.masked_fill(mask.repeat_interleave(E // n_group, 1) == 0, float("-inf"))
    idx = sel.topk(k, dim=-1).indices              # chosen experts, [T, k]
    gate = s.gather(1, idx)                        # weights use unbiased scores
    gate = gate / gate.sum(-1, keepdim=True) * scale
    return idx, gate

@torch.no_grad()
def update_bias(bias, idx, E, gamma=1e-3):
    load = torch.bincount(idx.flatten(), minlength=E).float()
    bias += gamma * torch.sign(load.mean() - load) # raise idle, lower busy experts

The bias update runs once per training step on the batch's expert counts and never receives a gradient. The step size here is illustrative; tune it so expert load converges within a few hundred steps without oscillating.

Worked example: dense or MoE on one 8-GPU node

Suppose you must serve a coding assistant on one node of eight 80 GB GPUs, and you are choosing between a dense 32B model, Qwen3-235B-A22B, and DeepSeek-V3. Start with weights. At 8 bits per parameter, the dense model needs about 32 GB, Qwen3-235B about 235 GB and DeepSeek-V3 about 671 GB. The node has 640 GB, so DeepSeek-V3 at 8 bits does not fit with room for the KV cache, while Qwen3-235B leaves roughly 400 GB for cache and activations.

Now compute. Per generated token, the dense model performs about 2 x 32B = 64 GFLOPs and Qwen3-235B about 2 x 22B = 44 GFLOPs, so the MoE is cheaper per token despite being seven times larger. But decode at small batch is bound by memory bandwidth, not FLOPs, and the question is how many bytes are read per step. With 128 experts and top-8, the chance that a given expert is unused by a batch of B tokens in a layer is (1 - 8/128)^B. At B = 1 that is 94 percent, so only active weights are read. At B = 32 it is 0.9375^32, about 13 percent, so 87 percent of experts are read every step and the MoE moves nearly all of its 235 GB per step, against 32 GB for the dense model.

The lesson: MoE is a clear win when you are compute bound, as in training and large-batch prefill, and when quality per FLOP matters. At moderate decode batch on a single node it is bandwidth hungry, which is why serious MoE serving spreads experts across many GPUs with expert parallelism so each GPU reads only its own experts. Layout choices are worked through in MoE serving architecture, in depth and multi-node launch in MoE Expert Parallelism Deployment.

Fine-tuning an MoE without breaking it

Fine-tuning an MoE adds risks a dense model does not have. A small, narrow dataset can push the router to send nearly every token to a handful of experts, which collapses capacity and, under expert parallelism, overloads the GPUs holding those experts. Practical rules that hold across the models above:

  • Keep load balancing on. If the model was trained with an auxiliary loss, keep it at the original coefficient; if it used bias balancing, keep updating the bias or freeze it, but do not let routing drift unmonitored.
  • Consider freezing the router for small datasets. You lose some adaptation, but routing stays as pretrained and stable.
  • Put LoRA adapters on attention and the shared expert first. Adapters on hundreds of routed experts multiply adapter parameters, and most experts see few fine-tuning tokens.
  • Log per-expert token counts per layer every few hundred steps. A healthy run keeps the max-to-mean load ratio near its pretraining value.
  • Evaluate on general benchmarks as well as the target task; routing changes can silently damage unrelated skills.

Failure modes

  • Routing collapse. A few experts take most tokens, quality drops and one GPU becomes the bottleneck. Watch load histograms; restore balancing.
  • Token dropping. Capacity-limited implementations discard tokens that overflow an expert's buffer, which shows up as rare, unexplained quality loss under load. Prefer dropless kernels or raise the capacity factor.
  • Memory surprise. Teams size hardware from active parameters and discover the model does not load. Size from total parameters plus KV cache.
  • Bandwidth surprise. Throughput per GPU is far below the active-parameter estimate at moderate batch, because nearly all experts are read every step.
  • Communication stalls. All-to-all exchanges between expert-parallel ranks dominate step time on slow interconnects; keep expert parallelism within fast links where possible.
  • Quantisation damage. Quantising the router as aggressively as the experts can flip expert choices; gpt-oss keeps its router out of MXFP4 for that reason.

Trade-offs: when MoE is the wrong choice

Choose a dense model when you serve at small batch on limited memory, when you need predictable latency per request, or when your fine-tuning data is small and you cannot afford routing risk. Choose an MoE when quality per training or inference FLOP is the constraint, when you serve at large batch across enough GPUs to spread the experts, or when a model of the quality you need only exists as an MoE. Small MoEs for single devices, where memory rather than compute dominates, are a different trade covered in Small MoE Architecture, and the parameter arithmetic is collected in MoE math architecture.

What to do next

  1. For any MoE you are considering, open its config.json and record routed experts, top-k, shared experts, dense layers and expert width; do not rely on blog summaries.
  2. Size memory from total parameters at your chosen precision, then add KV cache for your target context and concurrency.
  3. Compute the fraction of experts touched at your real decode batch with (1 - k/E)^B, and estimate bytes read per step from it.
  4. Benchmark against a dense model of similar active size on your own prompts, measuring quality, tokens per second per GPU and p99 latency.
  5. If fine-tuning, log per-expert load from step zero, keep balancing on, and try router freezing on small datasets.
  6. If serving across GPUs, map expert-parallel groups to your fastest interconnect and measure all-to-all time separately from compute.
  7. Keep routers and other small, sensitive tensors at higher precision when quantising.
Key takeaway: In 2026 most large open models are mixtures of experts with many small experts, often a shared expert, top-k routing balanced by a bias rather than a loss, routing limited to a few devices, and experts stored in low precision. MoE cuts compute per token but not memory, and at moderate batch it reads nearly every expert, so size from total parameters, measure bandwidth at your real batch, keep balancing on when fine-tuning, and verify every model detail from its own config.