Mixture-of-experts models are sold on one number: active parameters. Mixtral 8x7B has 46.7 billion parameters but uses only about 12.9 billion per token, because each layer routes every token to 2 of its 8 expert feed-forward blocks. The tempting conclusion is that it costs about as much to serve as a 13B dense model. That is true in one regime, false in another, and the difference decides your cost per million tokens.

This article builds a cost model from first principles. It separates what you pay for in memory capacity (all parameters, always), in memory bandwidth (the experts the batch actually touches), and in compute (active parameters only). It gives a formula for how many experts a batch touches, turns that into tokens per second and dollars per million tokens on real GPU specifications, and walks through where the MoE advantage appears and where it vanishes. The model is written as runnable Python so you can substitute your own model, hardware and price.

Three bills: capacity, bandwidth, compute

An inference step pays three separate bills, and an MoE model changes each one differently.

  • Capacity. Every expert must be resident somewhere, because any token may route to any expert. Mixtral in BF16 needs 93.4 GB for weights alone, more than one 80 GB H100, so the minimum deployment is two GPUs, the same as a 47B dense model. Whatever capacity remains holds the KV cache, which limits batch size.
  • Bandwidth. In decode, each step must stream every weight it uses from HBM. An MoE layer reads only the experts that at least one token in the batch routed to. At batch 1 that is 2 of 8; at large batch it is all 8.
  • Compute. Each token does about 2 FLOPs per active parameter, so FLOPs scale with active parameters at every batch size. This is the only bill on which MoE wins unconditionally.

Decode at modest batch sizes is bandwidth-bound, prefill is compute-bound. That one sentence explains most of the surprises below. For the per-operator version of the decode ledger on a dense model, see decode compute math.

How many experts does a batch touch?

How many distinct experts does a batch of B tokens touch in one layer? If routing is uniform and independent, each expert is missed by one token with probability 1 - k/E and by all B tokens with probability (1 - k/E)B. So the expected number touched is

U(B) = E * (1 - (1 - k/E) ** B)

For Mixtral (E = 8, k = 2): U(1) = 2, U(4) = 5.47, U(8) = 7.20, U(16) = 7.92, and by 32 concurrent sequences every expert is read on essentially every step. Fine-grained designs push the saturation point out. With 256 experts and top-8 routing, as in DeepSeek-V3's routed experts, U(8) is about 57, U(32) about 163 and U(128) about 252 of 256. That is one reason fine-grained MoE keeps a bandwidth advantage at larger batches.

Real routing is not uniform. Skew toward popular experts lowers U, which reduces bytes read, but it concentrates work on a few experts, and under expert parallelism the busiest GPU sets the step time. Uniform routing is therefore the right planning assumption for bandwidth, and an imbalance factor should be added for latency. Balancing techniques are covered in MoE load balancing.

A cost model you can run

The model below computes, for one decode step, the bytes read (dense weights, touched experts and KV cache) and the FLOPs, then takes the roofline maximum of the two times. MBU and MFU are the fractions of peak bandwidth and compute you actually achieve; 0.7 and 0.5 are reasonable planning values for a tuned serving stack, and you should replace them with measurements. Communication between GPUs is ignored here and treated separately below.

from dataclasses import dataclass

@dataclass
class MoE:
    layers: int = 32; d: int = 4096; ffn: int = 14336
    experts: int = 8; top_k: int = 2
    q_heads: int = 32; kv_heads: int = 8; head_dim: int = 128; vocab: int = 32000

    def expert_params(self):                       # SwiGLU: gate, up, down
        return 3 * self.d * self.ffn
    def dense_params(self):                        # attention, router, embeddings, lm head
        attn = 2 * self.d * self.q_heads * self.head_dim + 2 * self.d * self.kv_heads * self.head_dim
        return self.layers * (attn + self.d * self.experts) + 2 * self.vocab * self.d
    def active_params(self):
        return self.dense_params() + self.layers * self.top_k * self.expert_params()
    def kv_bytes_per_token(self, b=2):
        return self.layers * 2 * self.kv_heads * self.head_dim * b

def unique_experts(E, k, B):
    return E * (1 - (1 - k / E) ** B)

def decode_step(m, B, ctx, bw, flops, gpus, w_bytes=2, mbu=0.7, mfu=0.5):
    U = unique_experts(m.experts, m.top_k, B)
    weight_bytes = (m.dense_params() + m.layers * U * m.expert_params()) * w_bytes
    kv_bytes = B * ctx * m.kv_bytes_per_token()
    fl = 2 * B * m.active_params() + 4 * B * ctx * m.layers * m.q_heads * m.head_dim
    t = max((weight_bytes + kv_bytes) / (bw * gpus * mbu), fl / (flops * gpus * mfu))
    return B / t                                   # tokens per second

def usd_per_mtok(tok_s, gpus, usd_per_gpu_hour):
    return usd_per_gpu_hour * gpus / 3600 / tok_s * 1e6

With this configuration the model reproduces the published sizes: 46.70B total and 12.88B active parameters, of which 1.6B are attention, router and embeddings and each expert is 176M. The KV cache costs 128 KiB per token.

Worked example: Mixtral on two H100s

Deploy Mixtral in BF16 on two H100 SXM GPUs (80 GB and 3.35 TB/s HBM3 each, 989 TFLOPS dense BF16), with an average context of 4,096 tokens and a rental price of $3.00 per GPU-hour. The price is an assumption; substitute yours. After 93.4 GB of weights, 66.6 GB remain, enough for about 111 concurrent sequences at that context with 10 per cent headroom. Running the model gives:

BatchExperts read per layerWeight GBKV GBMemory msCompute msTokens/s$ per M tokens
12.0025.80.545.610.031789.34
45.4764.92.1514.290.112805.95
87.2084.44.2918.910.234233.94
167.9292.58.5921.550.457422.25
328.0093.417.1823.580.901,3571.23
648.0093.434.3627.241.812,3490.71
968.0093.451.5430.912.713,1060.54

Compare two dense models with the same attention shape on the same hardware. A dense 12.9B model gives 178 tokens per second at batch 1, identical to the MoE, and 4,993 at batch 64 ($0.33 per million tokens). A dense 46.7B model gives 50 tokens per second at batch 1 ($33.38 per million) and 2,349 at batch 64 ($0.71), identical to the MoE. In words: at batch 1 the MoE costs like its active size; by batch 32 or so it reads every expert, and in a bandwidth-bound decode it costs like its total size. Compute time stays a small fraction of memory time throughout, so the MoE's FLOP saving buys nothing in decode on this hardware.

Weight bytes read per decode step vs batch size (Mixtral-shaped MoE, BF16)0255075100GB1248163264concurrent sequences in the decode batch (log scale)dense 46.7B: 93.4 GB at any batchdense 12.9B: 25.8 GBMoE: 25.8 GB at batch 1, all experts by ~32
Bytes read per decode step: an MoE behaves like its active size at batch 1 and like its total size once the batch touches every expert.

Where MoE actually saves money

So where does MoE save money? Three places.

  • Prefill. Prompt processing is compute-bound, and compute scales with active parameters. At the same MFU of 0.5 across two H100s, prefill throughput is about 989e12 / (2 x 12.88e9), roughly 38,000 tokens per second, against about 10,600 for a dense 46.7B model, a 3.6x saving. Workloads dominated by long prompts and short answers, such as retrieval-augmented question answering, see most of this.
  • Very large batches. Once batch size pushes decode into the compute-bound region, active parameters again set the cost. With the H100's ratio of roughly 295 FLOPs per byte of bandwidth, a dense model reaches that region only at batches in the hundreds; KV capacity usually runs out first, which is why large-scale MoE serving spreads experts across many GPUs to free memory for KV.
  • Quality per unit of compute. The real argument for MoE is that a 46.7B-parameter model of knowledge is served with 13B worth of FLOPs per token. If the alternative is a dense model of similar quality, the comparison is against a larger dense model, not the 12.9B one.

To see how this plays out on a real request, price one with 2,000 prompt tokens and 300 output tokens on the same two GPUs at $6.00 per hour for the pair. Prefill at 38,393 tokens per second costs 86.8 millionths of a dollar; decode at the batch-64 rate of 2,349 tokens per second costs 212.9 millionths, for a total of about $0.30 per thousand requests. Decode is 71 per cent of the bill even though it handles only 13 per cent of the tokens. The dense 46.7B model pays the same 212.9 for decode but 314.8 for prefill, about $0.53 per thousand requests. So the MoE is 43 per cent cheaper on this mix, and all of the saving comes from prefill. Change the mix to 200 prompt tokens and 1,000 output tokens and the gap nearly disappears. This is why you must price your actual prompt and output length distribution rather than a single headline throughput number: the same model can look like a large saving or no saving at all depending on the traffic.

Expert parallelism: all-to-all and imbalance

With expert parallelism, each MoE layer adds two all-to-all exchanges per step: tokens are dispatched to the GPUs holding their experts, and results are combined back. Per token per layer, each exchange moves about k x d x bytes; for Mixtral that is 2 x 4096 x 2 = 16 KiB each way, or 512 KiB per token across 32 layers in each direction. Within an NVLink domain that is small; across nodes it can dominate decode latency, because all-to-all is latency-bound at small messages. Add it to the step time as a separate term, and multiply the expert compute term by an imbalance factor (the busiest GPU's load over the mean, measured on your own traffic, since it depends heavily on balancing). The mechanics are in MoE all-to-all and the deployment sizing in MoE expert-parallel deployment.

Cost levers

LeverBill it cutsCost or risk
Larger batches (continuous batching)Bandwidth per tokenLatency per token rises; KV capacity caps it
Weight quantisation (FP8, INT4 experts)Capacity and bandwidthQuality checks needed per expert
KV quantisation or pagingCapacity, so bigger batchesSmall accuracy and kernel complexity
Expert offload to CPU memoryCapacityPCIe bandwidth makes misses very slow
Fine-grained expertsBandwidth at mid batchMore routing and all-to-all overhead
Disaggregated prefill and decodeLets each phase use its best batchExtra KV transfer

Batch size is usually the biggest lever, and continuous batching is how you keep it high. KV memory decides how far you can push it; see KV cache sizing.

Failure modes in cost planning

  • Pricing by active parameters. Quoting 13B-dense costs for a 47B MoE underestimates decode cost by about 2x at production batch sizes.
  • Ignoring capacity. A plan that fits active parameters on one GPU fails at load time.
  • Benchmarking at batch 1. Single-stream tests show MoE at its best and say nothing about cost at the batch you will run.
  • Uniform-routing latency estimates. Hot experts set step time under expert parallelism; measure the per-expert token histogram on real traffic.
  • Forgetting the KV term. At long context, KV bytes can exceed weight bytes per step, and then MoE versus dense hardly matters.

What to do next

  1. Plug your model's configuration and your GPU's bandwidth, FLOPs and price into the cost model above and print the table for your target batch range.
  2. Measure real MBU and MFU at three batch sizes and replace the planning values.
  3. Log the per-layer expert histogram on production traffic and compute the actual experts touched per step and the imbalance factor.
  4. Split your cost estimate into prefill and decode using your real prompt and output length distribution.
  5. Price the levers: quantised experts, KV quantisation and larger batches, each with a quality and latency check.
  6. Re-run the model when you change hardware: the FLOPs-to-bandwidth ratio decides where MoE pays off.
Key takeaway: An MoE model pays for capacity with all its parameters, for decode bandwidth with the experts the batch touches, and for compute with active parameters only. At batch 1 it serves like its active size; once the batch touches every expert, decode costs like its total size, and the savings come from prefill, very large batches and better quality per FLOP.