A mixture-of-experts (MoE) model replaces each dense feed-forward block with many smaller expert networks and a router that sends every token to a few of them. The attraction for serving is simple: a model with 30 billion parameters that computes like a 3 billion parameter model. The catch is that you must still store all 30 billion, and the routing decisions change how memory, bandwidth and communication behave in ways a dense-model playbook does not predict.
This article takes the operator's view. It sizes one real model from its configuration file, shows why the advertised active parameter count misleads capacity planning, compares the two main ways to lay the model across GPUs, configures an engine for each, and lists the failures you will meet. The GPU-level mechanics of dispatch, combine and skew are covered in MoE serving on GPUs, and the router and parameter math in MoE math; here we use their results rather than re-deriving them.
What changes when the model is sparse
In a dense transformer every token touches every weight, so the cost model is easy: weight bytes are read once per decode step and shared by the whole batch, and compute scales with batch size. In an MoE layer each token picks its top-k experts from a router score. Different tokens in a batch pick different experts, so the weights read in a step depend on the batch, and the work per expert is uneven.
Three consequences follow for serving. Memory is set by total parameters, not active ones, so an MoE model needs the same GPU capacity as a dense model of its total size. Compute per token is set by active parameters, so prefill is cheap relative to size. And decode, which is limited by memory bandwidth, reads whichever experts the batch touched, which at realistic batch sizes is most of them. The engine then adds a fourth: how experts are split across GPUs decides whether each layer needs an all-reduce or an all-to-all, and whose memory holds the KV cache.
A worked model: Qwen3-30B-A3B by the numbers
Take Qwen3-30B-A3B, whose published configuration lists a hidden size of 2,048, 48 layers, 128 experts per layer with 8 selected per token, an expert intermediate size of 768, 32 query heads and 4 key-value heads of dimension 128, and a 151,936-token vocabulary with untied input and output embeddings. Every MoE fact an operator needs can be computed from those fields:
cfg = dict(hidden=2048, layers=48, experts=128, topk=8, moe_inter=768,
q_heads=32, kv_heads=4, head_dim=128, vocab=151936, tied=False)
def sizes(c, bytes_per_param=2):
expert = 3 * c["hidden"] * c["moe_inter"] # gate, up, down
attn = c["hidden"] * c["head_dim"] * (2 * c["q_heads"] + 2 * c["kv_heads"])
router = c["hidden"] * c["experts"]
embed = c["vocab"] * c["hidden"] * (1 if c["tied"] else 2)
total = c["layers"] * (c["experts"] * expert + attn + router) + embed
active = c["layers"] * (c["topk"] * expert + attn + router) + embed
kv_per_token = 2 * c["layers"] * c["kv_heads"] * c["head_dim"] * bytes_per_param
return total, active, kv_per_token
def expected_unique_experts(c, batch):
"""Uniform-routing estimate of distinct experts a decode step touches per layer."""
p_miss = (1 - c["topk"] / c["experts"]) ** batch
return c["experts"] * (1 - p_miss)
total, active, kv = sizes(cfg)
print(f"total {total/1e9:.1f}B, active {active/1e9:.2f}B, KV {kv/1024:.0f} KiB/token")
for b in (1, 8, 32, 64):
print(b, round(expected_unique_experts(cfg, b), 1))
# total 30.5B, active 3.35B, KV 96 KiB/token
# 1 8.0 | 8 51.6 | 32 111.8 | 64 125.9Each expert is three 2,048 by 768 matrices, about 4.7 million parameters, so one layer's 128 experts hold 604 million and all 48 layers hold 29.0 billion. Attention adds about 0.9 billion and the two embedding tables 0.6 billion, for 30.5 billion in total. Only 8 experts run per token, so the active count is about 3.3 billion, which is where the A3B in the name comes from.
| Quantity | Value | What it drives |
|---|---|---|
| Total parameters | 30.5B, 56.9 GiB in bf16 | Minimum GPU memory; the model does not fit a 48 GB card unquantized |
| Active parameters per token | about 3.3B | Prefill FLOPs, and decode bytes at batch 1 only |
| KV cache per token (bf16) | 96 KiB | Concurrent context you can hold after weights |
| KV heads | 4 | Tensor parallelism beyond 4 duplicates KV heads instead of splitting them |
Why 3B active is not 3B read per step
At batch size 1 a decode step reads about 3.3 billion parameters, roughly 6.7 GB in bf16, which is the promise of MoE. But with 64 sequences decoding together, each choosing 8 of 128 experts per layer, the uniform-routing estimate in the code above says the step touches about 126 distinct experts per layer: effectively the whole model. At batch 8 it is already about 52 of 128. Real routers are not uniform, which lowers the number a little for skewed traffic, but the conclusion holds: in the batch range where serving is economical, decode bandwidth tracks total parameters, not active ones.
That changes two planning assumptions. First, an MoE model is not a small model for latency purposes once it is batched: time per output token is bounded below by reading most of the weights, so compare it to a dense model of similar total size, divided across however many GPUs share the read. Second, the win shows up in compute: once all experts are being read anyway, extra tokens in the batch add expert FLOPs at roughly the cost of a 3B model, so throughput per GPU at large batch is where MoE beats a dense model of equal quality. Plan for high concurrency, not low latency at low load.
Choosing a layout: tensor parallel or expert parallel
On one 80 GB GPU the model fits with room for about 13 GiB of KV cache at 90 percent memory utilization, roughly 140,000 tokens of context in total. That is fine for a development box and thin for production, so most deployments use at least two GPUs, and there are two ways to split the model.
Tensor parallelism (TP). Every weight matrix, attention and experts alike, is split across GPUs, and each layer ends with an all-reduce. Each GPU stores half the weights, about 28.5 GiB, and half the KV heads, so KV cost per GPU halves too. It is the default in most engines and works well inside one NVLink domain. Its limits: every request occupies every GPU, the all-reduce sits on the critical path of every layer, and splitting each expert's small 768-wide matrices produces thin, inefficient GEMMs. With 4 KV heads, TP of 8 must replicate heads rather than split them, wasting KV memory.
Data-parallel attention with expert parallelism (DP+EP). Each GPU keeps a full copy of the attention layers and embeddings, about 1.5 billion parameters, and owns a slice of the experts: 64 of 128 each on two GPUs. Requests are spread across GPUs, and each GPU holds the KV cache only for its own sequences. In every MoE layer, tokens travel by all-to-all to the GPU that owns their chosen experts and come back after. Expert matrices stay whole, so the expert GEMMs are larger and more efficient.
The rule of thumb: for a small MoE like this one on 2 to 8 GPUs of one node, TP is simpler and usually fine. DP+EP wins as expert count and GPU count grow, and it is the layout large MoE models are designed for. The collective cost is analysed in MoE all-to-all communication.
Configuring the engine
In vLLM both layouts are command-line choices. The flags below are from its expert-parallel deployment guide; the expert-parallel group size is tensor-parallel size times data-parallel size:
# Layout A: tensor parallel across two GPUs
vllm serve Qwen/Qwen3-30B-A3B \
--tensor-parallel-size 2 \
--max-model-len 32768 \
--gpu-memory-utilization 0.90
# Layout B: two data-parallel attention ranks, experts sharded across them (EP size = TP x DP = 2)
vllm serve Qwen/Qwen3-30B-A3B \
--data-parallel-size 2 \
--enable-expert-parallel \
--max-model-len 32768 \
--gpu-memory-utilization 0.90
# Larger deployments: pick an all-to-all backend and turn on the expert load balancer.
# 128 logical + 16 redundant = 144 physical experts, 18 per GPU on EP=8.
vllm serve <large-moe-model> \
--tensor-parallel-size 1 --data-parallel-size 8 --enable-expert-parallel \
--all2all-backend deepep_low_latency \
--enable-eplb \
--eplb-config '{"window_size": 1000, "step_interval": 3000, "num_redundant_experts": 16}'Three parts of this need judgement. The all-to-all backend defaults to an all-gather and reduce-scatter implementation that is simple and portable; the DeepEP backends, one tuned for high-throughput prefill and one for low-latency decode, are designed for large multi-node EP and need their library installed and a supported network. The expert-parallel load balancer (EPLB) records how many tokens each expert receives over a sliding window and periodically re-places experts, optionally replicating hot ones as redundant copies; the physical expert count must still divide evenly over the EP group. Finally, KV memory under DP+EP is per rank, so a long-context request is limited by one GPU's KV budget, not the pool's.
Continuous batching and paged KV memory work as for dense models, and they matter more here because MoE throughput only arrives at high concurrency; see paged attention for how the KV pool is managed.
Benchmarking the choice
Do not pick a layout from first principles alone. Run the same traffic shape against both and compare throughput at the same latency targets, not raw peak throughput:
# Same traffic shape against both layouts; compare at equal p99 TTFT / TPOT targets.
# Option names move between releases: check `vllm bench serve --help` on your version.
vllm bench serve --model Qwen/Qwen3-30B-A3B \
--dataset-name random --random-input-len 2000 --random-output-len 300 \
--num-prompts 2000 --max-concurrency 64Sweep concurrency from low to past saturation and record time to first token and time per output token at p50 and p99. Expect TP to win at low concurrency, because every GPU works on every request, and DP+EP to pull ahead as concurrency grows, if the all-to-all is not the bottleneck. Use prompts whose length and topic mix resemble production: routing depends on content, so a synthetic random-token prompt can route unrealistically evenly and flatter the EP layout.
Also measure prefill and decode separately. Prefill batches thousands of tokens per step and is compute-bound, while decode moves a few tokens per sequence and is latency-bound; teams running large MoE models often split them onto separate pools with different parallel layouts, as described in prefill-decode disaggregation.
Failure modes you will see in production
- Out of memory at load despite small active size. Capacity was planned from active parameters. Plan from total parameters, plus KV cache, plus activation workspace.
- One slow rank sets the pace. Under EP a layer finishes only when the busiest GPU finishes its experts. A traffic shift, such as a burst of code or of one language, concentrates tokens on a few experts. Watch per-rank expert load, and use EPLB or redundant experts when imbalance persists.
- Throughput collapses when crossing nodes. The all-to-all went from NVLink to the inter-node network. Keep EP inside a node unless the model needs more memory than a node has, and when it does, use backends built for it.
- Long requests rejected under DP+EP. Each rank has its own KV pool, so a context that fits the aggregate may not fit one rank. Cap maximum length per rank or route long requests to a TP pool.
- Quality drift after quantization. Experts that receive few calibration tokens get poor scales. Check per-expert calibration coverage, and keep the router at higher precision.
Operational guidance and trade-offs
Treat an MoE deployment as a memory-heavy, throughput-oriented service. Size GPUs from total parameters plus the concurrency you need to reach good utilization; buy latency with more GPUs rather than smaller batches. Export expert load per rank, all-to-all time per step and KV utilization per rank alongside the usual latency metrics, because the first two do not exist for dense models and are where MoE regressions appear.
The trade-off space is: TP is simpler, uses one fabric pattern and gives the lowest latency for a small model on one node; DP+EP scales further and gives better GEMM efficiency but adds routing skew and all-to-all as new failure surfaces; weight quantization of the experts cuts the memory that dominates MoE cost and is often the first optimization to reach for.
What to do next
- Pull your model's config file and compute total parameters, active parameters and KV bytes per token with the script above.
- Size memory from the total plus the KV budget for your target concurrency, not from the active count.
- Stand up the TP layout on one node first and record p50 and p99 latency across a concurrency sweep with production-like prompts.
- Try DP+EP with the same traffic, and keep it only if it wins at your latency target.
- Add per-rank expert load and all-to-all timing to your dashboards, and enable EPLB only after you have seen persistent imbalance.
- Evaluate expert weight quantization with per-expert calibration checks before buying more GPUs.