Mixture-of-experts models are sold on one number: active parameters. Mixtral 8x7B has 46.7 billion parameters but uses only about 12.9 billion per token, because each layer routes every token to 2 of its 8 expert feed-forward blocks. The tempting conclusion is that it costs about as much to serve as a 13B dense model. That is true in one regime, false in another, and the difference decides your cost per million tokens.
This article builds a cost model from first principles. It separates what you pay for in memory capacity (all parameters, always), in memory bandwidth (the experts the batch actually touches), and in compute (active parameters only). It gives a formula for how many experts a batch touches, turns that into tokens per second and dollars per million tokens on real GPU specifications, and walks through where the MoE advantage appears and where it vanishes. The model is written as runnable Python so you can substitute your own model, hardware and price.
Three bills: capacity, bandwidth, compute
An inference step pays three separate bills, and an MoE model changes each one differently.
- Capacity. Every expert must be resident somewhere, because any token may route to any expert. Mixtral in BF16 needs 93.4 GB for weights alone, more than one 80 GB H100, so the minimum deployment is two GPUs, the same as a 47B dense model. Whatever capacity remains holds the KV cache, which limits batch size.
- Bandwidth. In decode, each step must stream every weight it uses from HBM. An MoE layer reads only the experts that at least one token in the batch routed to. At batch 1 that is 2 of 8; at large batch it is all 8.
- Compute. Each token does about 2 FLOPs per active parameter, so FLOPs scale with active parameters at every batch size. This is the only bill on which MoE wins unconditionally.
Decode at modest batch sizes is bandwidth-bound, prefill is compute-bound. That one sentence explains most of the surprises below. For the per-operator version of the decode ledger on a dense model, see decode compute math.
How many experts does a batch touch?
How many distinct experts does a batch of B tokens touch in one layer? If routing is uniform and independent, each expert is missed by one token with probability 1 - k/E and by all B tokens with probability (1 - k/E)B. So the expected number touched is
U(B) = E * (1 - (1 - k/E) ** B)For Mixtral (E = 8, k = 2): U(1) = 2, U(4) = 5.47, U(8) = 7.20, U(16) = 7.92, and by 32 concurrent sequences every expert is read on essentially every step. Fine-grained designs push the saturation point out. With 256 experts and top-8 routing, as in DeepSeek-V3's routed experts, U(8) is about 57, U(32) about 163 and U(128) about 252 of 256. That is one reason fine-grained MoE keeps a bandwidth advantage at larger batches.
Real routing is not uniform. Skew toward popular experts lowers U, which reduces bytes read, but it concentrates work on a few experts, and under expert parallelism the busiest GPU sets the step time. Uniform routing is therefore the right planning assumption for bandwidth, and an imbalance factor should be added for latency. Balancing techniques are covered in MoE load balancing.
A cost model you can run
The model below computes, for one decode step, the bytes read (dense weights, touched experts and KV cache) and the FLOPs, then takes the roofline maximum of the two times. MBU and MFU are the fractions of peak bandwidth and compute you actually achieve; 0.7 and 0.5 are reasonable planning values for a tuned serving stack, and you should replace them with measurements. Communication between GPUs is ignored here and treated separately below.
from dataclasses import dataclass
@dataclass
class MoE:
layers: int = 32; d: int = 4096; ffn: int = 14336
experts: int = 8; top_k: int = 2
q_heads: int = 32; kv_heads: int = 8; head_dim: int = 128; vocab: int = 32000
def expert_params(self): # SwiGLU: gate, up, down
return 3 * self.d * self.ffn
def dense_params(self): # attention, router, embeddings, lm head
attn = 2 * self.d * self.q_heads * self.head_dim + 2 * self.d * self.kv_heads * self.head_dim
return self.layers * (attn + self.d * self.experts) + 2 * self.vocab * self.d
def active_params(self):
return self.dense_params() + self.layers * self.top_k * self.expert_params()
def kv_bytes_per_token(self, b=2):
return self.layers * 2 * self.kv_heads * self.head_dim * b
def unique_experts(E, k, B):
return E * (1 - (1 - k / E) ** B)
def decode_step(m, B, ctx, bw, flops, gpus, w_bytes=2, mbu=0.7, mfu=0.5):
U = unique_experts(m.experts, m.top_k, B)
weight_bytes = (m.dense_params() + m.layers * U * m.expert_params()) * w_bytes
kv_bytes = B * ctx * m.kv_bytes_per_token()
fl = 2 * B * m.active_params() + 4 * B * ctx * m.layers * m.q_heads * m.head_dim
t = max((weight_bytes + kv_bytes) / (bw * gpus * mbu), fl / (flops * gpus * mfu))
return B / t # tokens per second
def usd_per_mtok(tok_s, gpus, usd_per_gpu_hour):
return usd_per_gpu_hour * gpus / 3600 / tok_s * 1e6With this configuration the model reproduces the published sizes: 46.70B total and 12.88B active parameters, of which 1.6B are attention, router and embeddings and each expert is 176M. The KV cache costs 128 KiB per token.
Worked example: Mixtral on two H100s
Deploy Mixtral in BF16 on two H100 SXM GPUs (80 GB and 3.35 TB/s HBM3 each, 989 TFLOPS dense BF16), with an average context of 4,096 tokens and a rental price of $3.00 per GPU-hour. The price is an assumption; substitute yours. After 93.4 GB of weights, 66.6 GB remain, enough for about 111 concurrent sequences at that context with 10 per cent headroom. Running the model gives:
| Batch | Experts read per layer | Weight GB | KV GB | Memory ms | Compute ms | Tokens/s | $ per M tokens |
|---|---|---|---|---|---|---|---|
| 1 | 2.00 | 25.8 | 0.54 | 5.61 | 0.03 | 178 | 9.34 |
| 4 | 5.47 | 64.9 | 2.15 | 14.29 | 0.11 | 280 | 5.95 |
| 8 | 7.20 | 84.4 | 4.29 | 18.91 | 0.23 | 423 | 3.94 |
| 16 | 7.92 | 92.5 | 8.59 | 21.55 | 0.45 | 742 | 2.25 |
| 32 | 8.00 | 93.4 | 17.18 | 23.58 | 0.90 | 1,357 | 1.23 |
| 64 | 8.00 | 93.4 | 34.36 | 27.24 | 1.81 | 2,349 | 0.71 |
| 96 | 8.00 | 93.4 | 51.54 | 30.91 | 2.71 | 3,106 | 0.54 |
Compare two dense models with the same attention shape on the same hardware. A dense 12.9B model gives 178 tokens per second at batch 1, identical to the MoE, and 4,993 at batch 64 ($0.33 per million tokens). A dense 46.7B model gives 50 tokens per second at batch 1 ($33.38 per million) and 2,349 at batch 64 ($0.71), identical to the MoE. In words: at batch 1 the MoE costs like its active size; by batch 32 or so it reads every expert, and in a bandwidth-bound decode it costs like its total size. Compute time stays a small fraction of memory time throughout, so the MoE's FLOP saving buys nothing in decode on this hardware.
Where MoE actually saves money
So where does MoE save money? Three places.
- Prefill. Prompt processing is compute-bound, and compute scales with active parameters. At the same MFU of 0.5 across two H100s, prefill throughput is about 989e12 / (2 x 12.88e9), roughly 38,000 tokens per second, against about 10,600 for a dense 46.7B model, a 3.6x saving. Workloads dominated by long prompts and short answers, such as retrieval-augmented question answering, see most of this.
- Very large batches. Once batch size pushes decode into the compute-bound region, active parameters again set the cost. With the H100's ratio of roughly 295 FLOPs per byte of bandwidth, a dense model reaches that region only at batches in the hundreds; KV capacity usually runs out first, which is why large-scale MoE serving spreads experts across many GPUs to free memory for KV.
- Quality per unit of compute. The real argument for MoE is that a 46.7B-parameter model of knowledge is served with 13B worth of FLOPs per token. If the alternative is a dense model of similar quality, the comparison is against a larger dense model, not the 12.9B one.
To see how this plays out on a real request, price one with 2,000 prompt tokens and 300 output tokens on the same two GPUs at $6.00 per hour for the pair. Prefill at 38,393 tokens per second costs 86.8 millionths of a dollar; decode at the batch-64 rate of 2,349 tokens per second costs 212.9 millionths, for a total of about $0.30 per thousand requests. Decode is 71 per cent of the bill even though it handles only 13 per cent of the tokens. The dense 46.7B model pays the same 212.9 for decode but 314.8 for prefill, about $0.53 per thousand requests. So the MoE is 43 per cent cheaper on this mix, and all of the saving comes from prefill. Change the mix to 200 prompt tokens and 1,000 output tokens and the gap nearly disappears. This is why you must price your actual prompt and output length distribution rather than a single headline throughput number: the same model can look like a large saving or no saving at all depending on the traffic.
Expert parallelism: all-to-all and imbalance
With expert parallelism, each MoE layer adds two all-to-all exchanges per step: tokens are dispatched to the GPUs holding their experts, and results are combined back. Per token per layer, each exchange moves about k x d x bytes; for Mixtral that is 2 x 4096 x 2 = 16 KiB each way, or 512 KiB per token across 32 layers in each direction. Within an NVLink domain that is small; across nodes it can dominate decode latency, because all-to-all is latency-bound at small messages. Add it to the step time as a separate term, and multiply the expert compute term by an imbalance factor (the busiest GPU's load over the mean, measured on your own traffic, since it depends heavily on balancing). The mechanics are in MoE all-to-all and the deployment sizing in MoE expert-parallel deployment.
Cost levers
| Lever | Bill it cuts | Cost or risk |
|---|---|---|
| Larger batches (continuous batching) | Bandwidth per token | Latency per token rises; KV capacity caps it |
| Weight quantisation (FP8, INT4 experts) | Capacity and bandwidth | Quality checks needed per expert |
| KV quantisation or paging | Capacity, so bigger batches | Small accuracy and kernel complexity |
| Expert offload to CPU memory | Capacity | PCIe bandwidth makes misses very slow |
| Fine-grained experts | Bandwidth at mid batch | More routing and all-to-all overhead |
| Disaggregated prefill and decode | Lets each phase use its best batch | Extra KV transfer |
Batch size is usually the biggest lever, and continuous batching is how you keep it high. KV memory decides how far you can push it; see KV cache sizing.
Failure modes in cost planning
- Pricing by active parameters. Quoting 13B-dense costs for a 47B MoE underestimates decode cost by about 2x at production batch sizes.
- Ignoring capacity. A plan that fits active parameters on one GPU fails at load time.
- Benchmarking at batch 1. Single-stream tests show MoE at its best and say nothing about cost at the batch you will run.
- Uniform-routing latency estimates. Hot experts set step time under expert parallelism; measure the per-expert token histogram on real traffic.
- Forgetting the KV term. At long context, KV bytes can exceed weight bytes per step, and then MoE versus dense hardly matters.
What to do next
- Plug your model's configuration and your GPU's bandwidth, FLOPs and price into the cost model above and print the table for your target batch range.
- Measure real MBU and MFU at three batch sizes and replace the planning values.
- Log the per-layer expert histogram on production traffic and compute the actual experts touched per step and the imbalance factor.
- Split your cost estimate into prefill and decode using your real prompt and output length distribution.
- Price the levers: quantised experts, KV quantisation and larger batches, each with a quality and latency check.
- Re-run the model when you change hardware: the FLOPs-to-bandwidth ratio decides where MoE pays off.