A mixture-of-experts layer replaces one feed-forward block with many, and a router sends each token to a few of them. The promise is more parameters at the same compute per token. The catch is that the router is learned, and nothing in the language-modelling loss asks it to spread tokens evenly. Left alone, it sends most tokens to a few experts, which then overflow while the rest idle. Because experts live on different GPUs, the busiest one sets the step time for everyone.

This article is about load balancing as an engineering problem: why imbalance happens, how to measure it at the expert, GPU and node level, and the four levers that control it: capacity limits, auxiliary losses, the bias-based controller used by DeepSeek-V3, and expert placement and replication at serving time. It includes a runnable simulation of the bias controller, a worked example with real numbers from that simulation, and what to put on a dashboard. The derivation of router gradients and loss formulas is covered in MoE Routing Math, in depth; here the focus is keeping a training or serving cluster busy.

Why imbalance happens and what it costs

Imbalance is the default because routing is self-reinforcing. Early in training, an expert that receives slightly more tokens gets more gradient signal, improves faster, scores higher for more tokens, and receives more still. Without a counter-force, a few experts win and the rest stop learning, which is called routing collapse. Even a well-trained router stays uneven on real traffic: code, a single language or a long repeated document all concentrate tokens on the experts that specialise in them.

The cost shows up in three places.

  • Step time. With expert parallelism, each GPU computes its experts' tokens and then waits in the all-to-all for every other GPU. The step runs at the speed of the most loaded GPU. A max-to-mean load ratio of 1.5 means roughly a third of the cluster's expert compute is idle.
  • Dropped or padded tokens. Kernels that need static shapes give each expert a fixed buffer. Tokens beyond it are dropped (they skip the layer through the residual) and buffers that are not full are padding, so both overflow and underflow waste work.
  • Memory. Dropless implementations size buffers for the worst case seen, so a hot expert inflates activation memory and can cause out-of-memory errors on one rank only.

The communication side, where uneven routing becomes uneven all-to-all message sizes, is covered in MoE All-to-All Communication.

Measuring imbalance

Measure per layer and per step, because layers differ a lot and averages hide the worst one. Four numbers carry most of the information: the max-to-mean ratio of tokens per expert, the same ratio per GPU (which is what actually sets step time), the drop rate if you enforce capacity, and the count of experts receiving nothing. Routing entropy, normalised so that 1.0 means uniform, is a useful single trend line for collapse.

import numpy as np

def balance_metrics(chosen, n_experts, experts_per_gpu, capacity=None):
    """chosen: [tokens, k] expert ids for one MoE layer in one step."""
    load = np.bincount(chosen.ravel(), minlength=n_experts).astype(float)
    gpu = load.reshape(-1, experts_per_gpu).sum(axis=1)       # contiguous placement
    p = load / load.sum()
    out = {
        "expert_max_over_mean": load.max() / load.mean(),
        "gpu_max_over_mean": gpu.max() / gpu.mean(),
        "cv": load.std() / load.mean(),
        "entropy_ratio": -(p[p > 0] * np.log(p[p > 0])).sum() / np.log(n_experts),
        "dead_experts": int((load == 0).sum()),
    }
    if capacity is not None:
        out["drop_rate"] = np.maximum(load - capacity, 0).sum() / load.sum()
    return out

Two cautions. Statistics from a single small micro-batch are noisy; aggregate over the global batch or a window of steps before alerting. And the per-GPU ratio depends on placement: four hot experts that happen to sit on the same GPU are far worse than the same four spread across four GPUs, even though the per-expert ratio is identical.

Load balancing is a control loop around the routerTokenshidden statesRouter scoress = sigmoid(xW)Top-k on s + bbias picks expertsDispatch (all-to-all)capacity checkExpert FFNsone GEMM per expertLoad counterstokens per expertControllerb -= gamma*sign(err)PlacementEPLB, replicasbias updateslow loopFast loop: bias or auxiliary loss every step. Slow loop: expert placement and replication every few minutes.Gate weights for the expert outputs still come from s alone; the bias only changes which experts are chosen.
The two control loops. The fast loop changes routing every step; the slow loop changes where experts live.

Lever 1: capacity factors and drops

The oldest lever is a hard limit. The Switch Transformer defines expert capacity as the capacity factor times the tokens per batch divided by the number of experts; with top-k routing the numerator becomes tokens times k. A capacity factor of 1.0 leaves no slack; 1.25 allows a quarter more than a perfectly even share.

Worked numbers from the simulation later in this article: 4,096 tokens, 64 experts and top-8 routing make 32,768 assignments, an even share of 512 per expert, and a capacity of 640 at factor 1.25. With the untrained bias, the busiest expert received 856 tokens and 777 assignments, 2.37 percent, were dropped. After the controller settled, the busiest expert received 567 and nothing was dropped.

Capacity limits bound memory and keep shapes static, but drops are silent quality loss, and they concentrate on exactly the tokens the busy experts are best at. DeepSeek-V3 reports dropping no tokens in training or inference, relying on balancing instead. Grouped-GEMM kernels that accept variable token counts per expert make dropless execution practical; the remaining cost is the worst-case buffer memory.

Lever 2: auxiliary losses

The classic counter-force is an auxiliary loss added to the training loss. In the Switch form it is alpha times N times the sum over experts of f_i times P_i, where f_i is the fraction of tokens dispatched to expert i and P_i is the mean router probability for expert i. It is minimised by a uniform distribution, and its gradient flows through P_i because f_i is not differentiable. Switch used alpha = 0.01. ST-MoE added a router z-loss, the squared log-sum-exp of the router logits with coefficient 0.001, which keeps logits small and avoids numerical instability in the router rather than balancing load directly.

The auxiliary loss works, but it is a compromise. Its gradient competes with the language-modelling gradient, so a large alpha balances load at the cost of model quality and a small alpha lets imbalance grow. That tension motivated the next lever.

Lever 3: the bias controller without an auxiliary loss

DeepSeek-V3 balances load without an auxiliary loss on the main objective. Each routed expert gets a bias b_i that is added to its affinity score only when choosing the top-k experts. The gating value that multiplies the expert's output still comes from the original score, so the bias steers traffic without distorting what the model computes. After each step the bias of an overloaded expert is decreased by gamma and that of an underloaded expert increased by gamma. The report uses gamma = 0.001 for the first 14.3 trillion training tokens and 0 for the final 500 billion, freezing the biases, though not the router weights, for the end of training. It keeps a complementary sequence-wise balance loss with a very small coefficient, 0.0001, only to avoid extreme imbalance within a single sequence.

The bias is not trained by gradient descent; it is a sign-based integral controller. That is why it does not fight the language-modelling loss, and also why it has controller failure modes: a gamma too large makes the bias oscillate, too small makes it lag behind shifts in the data. The simulation below implements the rule on a synthetic router with four deliberately popular experts.

import numpy as np

def route(scores, bias, k):
    """scores: [T, E] affinities. Bias changes WHICH experts are picked, not their weights."""
    chosen = np.argsort(-(scores + bias), axis=1)[:, :k]       # selection uses s + b
    gate = np.take_along_axis(scores, chosen, axis=1)           # weights use s alone
    return chosen, gate / gate.sum(axis=1, keepdims=True)

def update_bias(bias, chosen, n_experts, gamma=1e-3):
    load = np.bincount(chosen.ravel(), minlength=n_experts)
    err = load.mean() - load                                     # positive: underloaded
    return bias + gamma * np.sign(err), load

rng = np.random.default_rng(0)
T, E, K, D = 4096, 64, 8, 32
W = rng.normal(size=(D, E)); W[:, :4] += 1.0                    # four popular experts
bias = np.zeros(E)
for step in range(301):
    x = rng.normal(size=(T, D)) / np.sqrt(D)
    scores = 1 / (1 + np.exp(-(x @ W)))
    chosen, gate = route(scores, bias, K)
    bias, load = update_bias(bias, chosen, E)
    if step % 50 == 0:
        print(step, round(load.max() / load.mean(), 2))

With seed 0 the expert max-to-mean ratio starts at 1.67, falls to 1.15 by step 50, and then hovers around 1.1. It does not reach 1.0: a fixed step on the sign of the error keeps every bias moving by gamma each step, and each batch is a fresh random sample. That residual is normal, and it is why capacity slack or dropless kernels are still needed on top of any balancer.

Which tokens to balance over

Over which tokens should balance be measured? The choice changes what the model learns. Balancing within each micro-batch, or each sequence, forces every small slice of data to use all experts evenly, which works against specialisation: a batch of code should be allowed to prefer the code experts. Work from the Qwen team on load-balancing loss in 2025 argued for computing the expert frequencies over the global batch, synchronised across data-parallel ranks, and reported better specialisation and model quality from it. DeepSeek-V3's design points the same way: its main balancer reacts to whole-batch load, and the sequence-wise term is kept tiny.

The systems view gives the opposite pressure. Step time depends on the load in each micro-batch on each GPU, not on the average over the global batch. A practical compromise is to compute the balancing signal over a large scope, keep capacity slack or dropless kernels for the per-step variance, and monitor the per-step per-GPU ratio separately.

Lever 4: placement and replication

Routing levers change which experts tokens choose. Placement levers change where those experts run, and they matter most at inference, where you cannot retrain the router and traffic shifts with users. Three are standard.

  • Spread hot experts. In the simulation, contiguous placement of 64 experts on 8 GPUs turned an expert ratio of 1.67 into a GPU ratio of 1.20; after balancing it was 1.02. Choosing the placement from measured load reduces the GPU ratio further at no training cost.
  • Node-limited routing. DeepSeek-V3 has 256 routed experts and one shared expert per MoE layer, activates 8 routed experts per token, and sends each token to at most 4 nodes. Limiting the fan-out bounds slow cross-node traffic.
  • Redundant experts. Replicate the hottest experts and split their traffic across copies. For prefill, DeepSeek-V3 deploys 32 redundant experts, with high-load experts chosen from online statistics and adjusted periodically. DeepSeek's open-source EPLB library computes such a plan from a load estimate.
import eplb   # github.com/deepseek-ai/EPLB

# weight: [layers, logical_experts] recent load per expert, e.g. a moving average of token counts
phy2log, log2phy, logcnt = eplb.rebalance_experts(
    weight, num_replicas=288, num_groups=8, num_nodes=4, num_gpus=32)
# phy2log: which logical expert each physical slot holds; logcnt: replicas per logical expert

EPLB has a hierarchical policy, which packs expert groups onto nodes evenly and then replicates within nodes, and a global policy for when nodes do not divide the groups evenly. Its README states that predicting expert load is out of scope and suggests a moving average of history; the argument values above are illustrative. Applying a new plan means moving weights between GPUs, so it runs on the slow loop. Deployment mechanics are in MoE Expert Parallelism Deployment, in depth and serving architecture in Mixture-of-Experts Serving Architecture in Depth.

Failure modes

  • Routing collapse. Entropy falls and dead experts appear early in training. Check that the balancer is active in every MoE layer and that the router runs in float32.
  • Bias oscillation. Loads flip between overloaded and underloaded each step; lower gamma or aggregate load over more tokens before each update.
  • Hidden drops. Evaluation runs with a different capacity factor than training, or drops are not logged at all. Always export the drop rate per layer.
  • Rank-local out-of-memory. One GPU fails while others have headroom; that is a hot expert in a dropless setup. Cap buffer size or replicate the expert.
  • Serving drift. Training-time balance does not survive production traffic; a single large customer can shift load. Refresh placement from live statistics.
  • Checkpoint mismatch. Restoring weights without the routing biases changes routing silently. Save the biases with the model.

Operating it

Put per-layer heatmaps of tokens per expert on the training dashboard, with the per-GPU max-to-mean ratio, drop rate, entropy ratio and dead-expert count as time series. Alert on a sustained rise in the GPU ratio, on any dead expert after warm-up, and on a drop rate above your tolerance. In serving, record per-expert load per replica and feed a moving average to the placement planner; re-plan when the predicted GPU ratio improves by a meaningful margin, not on every fluctuation, because each move costs weight transfers. The parallelism layout these metrics sit on is described in Expert Parallelism.

Trade-offs

LeverStrengthCost
Capacity factor with dropsStatic shapes, bounded memorySilent quality loss on hot experts
Dropless kernelsNo token lossWorst-case buffer memory, dynamic shapes
Auxiliary lossSimple, well studiedCompetes with the main loss; alpha needs tuning
Bias controllerDoes not distort gradients or gate valuesNeeds gamma tuning; residual per-step jitter
Global-batch statisticsAllows specialisationCross-rank sync; per-step imbalance remains
Placement and replicasFixes serving imbalance without retrainingExtra memory, weight movement

What to do next

  1. Export balance_metrics per layer per step from your training or serving loop, including the per-GPU ratio.
  2. Run the bias simulation above with your expert count, k and batch size, and sweep gamma to see oscillation and lag.
  3. Decide on dropless or capacity-limited execution and log the drop rate either way.
  4. If you use an auxiliary loss, compute its frequencies over the global batch and compare expert specialisation.
  5. For serving, collect per-expert load for a day and compute a replication plan with EPLB before buying more GPUs.
  6. Add alerts on GPU max-to-mean ratio, dead experts and drop rate, and save router biases in checkpoints.
Key takeaway: MoE routing drifts toward a few experts unless something pushes back, and the busiest GPU sets the step time. Measure tokens per expert and per GPU, drops and entropy per layer; balance with an auxiliary loss or a bias controller that changes selection but not gate values; measure balance over large token scopes; and fix serving imbalance with placement and replicated hot experts.