Every LLM serving engine has a setting for maximum batch size, and every cost spreadsheet seems to assume the engine runs at it. It almost never does. The batch a GPU actually runs is set by how many requests are in flight on that replica, which is set by traffic, by how many replicas share it, and by how long each request lasts. That is why two teams with the same model, the same GPU and the same engine settings can pay prices per token that differ by four times.
This article is about the batch you get rather than the batch you configure. It assumes the single-GPU physics, covered in LLM Cost Analysis and GPU Throughput Math, and works out what happens at the level of a fleet and a product: how achieved batch follows from traffic, why consolidating replicas cuts cost, what a latency target costs, where static batches waste work on padding, and how to bill tenants who share a batch.
The physics in one paragraph
One paragraph of recap. In decode, each step reads all model weights from GPU memory once, however many sequences are in the batch, plus each sequence's KV cache. At small batch the weight read dominates, so adding sequences is almost free and cost per token falls nearly in proportion. At large batch the KV reads dominate and returns shrink. Memory for KV sets a hard ceiling on batch. For an 8B model on one H100 at 2,000 tokens of context, the cost-analysis article's numbers fit a simple line: step time is about t(B) = 4.85 ms + 0.078 ms x B. We use that line throughout, at an assumed $3 per GPU-hour, counting decode only. Those figures assume peak memory bandwidth, so measured costs run higher, and prefill adds to every number, but neither changes the conclusions.
Achieved batch is set by traffic
Little's law says the number of requests in a system equals arrival rate times time in the system. For one replica decoding L output tokens per request at arrival rate lambda, each request spends about L x t(B) seconds decoding, so the batch in flight is:
B = lam * L * t(B) = lam * L * (a + s * B)
=> B = lam * L * a / (1 - lam * L * s) # valid only while lam * L * s < 1Three things follow. The configured maximum matters only when this B would exceed it; below that, raising --max-num-seqs changes nothing. Achieved batch rises faster than linearly with traffic per replica, because a bigger batch makes each step slower, which keeps requests in flight longer. And when lam x L x s approaches 1, or B hits the KV ceiling, the replica is saturated: arrivals outpace completions and queues grow without bound. Cost per token is therefore not a property of the model and GPU. It is a property of traffic per replica.
Worked example: consolidating a fleet
A service receives 40 requests per second, each producing 300 output tokens, on the 8B model above. It runs on eight replicas because that is what the first load test used. Solving the equation for each fleet size:
| Replicas | Req/s per replica | Achieved batch | Time per output token | $ per M output tokens |
|---|---|---|---|---|
| 8 | 5 | 8.2 | 5.5 ms | 0.556 |
| 4 | 10 | 19.0 | 6.3 ms | 0.278 |
| 3 | 13.3 | 28.2 | 7.0 ms | 0.208 |
| 2 | 20 | 54.7 | 9.1 ms | 0.139 |
| 1 | 40 | saturated | unbounded | n/a |
Moving from eight replicas to two cuts decode cost by four times while time per output token rises from 5.5 to 9.1 ms, still faster than anyone reads. One replica fails: at a batch of about 200, where the KV ceiling of this setup sits, it completes about 32.6 requests per second, less than the 40 arriving. The lesson is that the cheapest fleet is the smallest one that stays clear of saturation with headroom for peaks and the loss of one replica. Over-provisioning does not just waste idle GPUs; it makes the busy ones less efficient too, because it divides the batch. Apply the loss-of-one rule here and two replicas fail it: if one dies, the survivor receives all 40 requests per second and saturates. Three is the answer, at about $0.208 per million output tokens and 7.0 ms per token, still 2.7 times cheaper than eight.
Routing matters for the same reason. Round-robin across many replicas spreads traffic thin. Prefer routing that fills replicas, least-outstanding-requests or prefix-aware routing, and an autoscaler that targets requests in flight per replica, not GPU utilisation. A GPU decoding a batch of four reports nearly 100 percent utilisation, because the metric records whether a kernel is running, not how much work it does.
The price of a latency target
Fix the latency target instead of the fleet, and the same line gives the largest batch the target allows and its cost. A time-per-output-token target T allows B = (T - a) / s:
| Target time per output token | Largest batch | GPU tokens/s | $ per M output tokens |
|---|---|---|---|
| 6 ms | 14 | 2,457 | 0.339 |
| 10 ms | 66 | 6,603 | 0.126 |
| 15 ms | 130 | 8,675 | 0.096 |
| 20 ms | 194 | 9,712 | 0.086 |
Tightening the target from 10 ms to 6 ms almost triples the cost per token. Relaxing it from 10 to 20 ms saves about a third. That shape, steep on one side and flat on the other, is why the latency target is a product decision with a price, not an engineering default. Many products can offer two tiers on the same GPUs: interactive traffic with a tight target and priority, and bulk traffic that tolerates slower streams and fills the rest of each batch.
Long prompts complicate the picture by stalling decode while they prefill. Chunked prefill bounds the stall at some throughput cost; Chunked Prefill in Serving sizes the chunk from the same per-token target.
Measuring your own curve
The line above is a model. Measure your own with a concurrency sweep against the real server: hold a fixed number of streaming requests in flight, record time per output token and total tokens per second, and convert to cost. This client works against any OpenAI-compatible completions endpoint; the ignore_eos field is a vLLM extension that fixes output length so sweeps are comparable, and you should drop it for a realistic run:
import asyncio, time, statistics, httpx
URL, MODEL, PRICE_PER_HOUR = "http://localhost:8000/v1/completions", "llama-8b", 3.0
async def one(client, prompt, tpots, tokens):
t0, n, first = time.perf_counter(), 0, None
async with client.stream("POST", URL, json={"model": MODEL, "prompt": prompt,
"max_tokens": 300, "stream": True, "ignore_eos": True}) as r:
async for line in r.aiter_lines():
if line.startswith("data: ") and line != "data: [DONE]":
n += 1; first = first or time.perf_counter()
if n > 1:
tpots.append((time.perf_counter() - first) / (n - 1))
tokens.append(n)
async def sweep(concurrency, prompts, seconds=120):
tpots, tokens, end = [], [], time.perf_counter() + seconds
async with httpx.AsyncClient(timeout=None) as client:
async def worker(i):
while time.perf_counter() < end:
await one(client, prompts[i % len(prompts)], tpots, tokens)
start = time.perf_counter()
await asyncio.gather(*(worker(i) for i in range(concurrency)))
tps = sum(tokens) / (time.perf_counter() - start)
p99 = statistics.quantiles(tpots, n=100)[98]
usd = PRICE_PER_HOUR / (tps * 3600) * 1e6
print(f"c={concurrency:4d} tok/s={tps:7.0f} p99 TPOT={p99*1e3:5.1f} ms $/M={usd:.3f}")
for c in (1, 8, 32, 64, 128):
asyncio.run(sweep(c, prompts=["<replace with real prompts>"] * 64))Use real prompts with real length distributions; uniform synthetic prompts flatter the result. In production, compare the sweep with the batch you actually run: vLLM exports a running-requests gauge, vllm:num_requests_running, on its Prometheus endpoint. If production sits at 8 while the sweep showed cost flattening near 64, the cheapest change is fewer replicas, not faster kernels.
Static batches and padding
Static batching, used for encoders, embedding models, classifiers and many offline jobs, pads every sequence in a batch to the longest one. Padded positions cost compute and produce nothing, so the useful fraction of a batch is mean length over maximum length. With chat-like inputs, where most texts are short and a few are long, a randomly assembled batch of 32 easily runs at 30 to 40 percent efficiency. Static Batching describes the mechanism; the fix for cost is to group sequences of similar length before batching:
def padded_efficiency(batches):
real = sum(sum(b) for b in batches)
paid = sum(len(b) * max(b) for b in batches)
return real / paid
def bucketed(lengths, batch_size):
order = sorted(lengths) # or sort within windows of 50 batches
return [order[i:i + batch_size] for i in range(0, len(order), batch_size)]
# lengths = token counts of tomorrow's embedding job
# print(padded_efficiency(random_batches), padded_efficiency(bucketed(lengths, 32)))Sorting a whole offline job is free, since order does not matter. For online static batching, sort within a short window instead, and cap the window by latency. Engines that pack variable-length sequences without padding remove the problem; check whether yours does before writing bucketing code.
Filling troughs with offline work
Work that can wait is the cheapest way to raise achieved batch. Offline jobs, such as evaluation sets, document enrichment and embedding back-fills, have no latency target, so they can run at the largest batch the memory allows. Hosted providers price this in: OpenAI's Batch API and Anthropic's Message Batches API both charge half the standard rate in exchange for results within 24 hours. On your own GPUs the equivalent is a low-priority queue that fills each replica's batch during traffic troughs and yields to interactive requests. The interactive fleet's nightly idle hours become paid-for capacity doing useful work, which improves the utilisation figure that dominates real cost.
Attributing cost inside a shared batch
Shared batches make per-tenant cost hard to see: one step serves many tenants, and its cost depends on everyone in it. A fair split follows the physics. Divide each step's weight-read cost evenly among the sequences in it, and charge each sequence's KV read in proportion to its context length:
def attribute_step(step_seconds, usd_per_s, seqs, a, s_per_ctx_token):
"""seqs: list of (tenant, context_len). Returns dollars per tenant for one step."""
fixed = a * usd_per_s / len(seqs) # weights, shared evenly
kv_total = sum(ctx for _, ctx in seqs) * s_per_ctx_token
scale = (step_seconds - a) / kv_total if kv_total else 0.0
out = {}
for tenant, ctx in seqs:
out[tenant] = out.get(tenant, 0.0) + fixed + ctx * s_per_ctx_token * scale * usd_per_s
return outTwo effects show up in the bills. Tenants who send traffic at quiet hours pay more per token, because they share the weight read with fewer neighbours; that is real and worth pricing. And long-context tenants pay for their KV reads and for the batch slots their memory takes from others, which is the main reason to price long contexts higher.
Failure modes
- Maximum batch set above what memory holds. The engine admits sequences, runs out of KV blocks and preempts some, recomputing their work. Throughput drops while the configured number looks generous.
- Autoscaling on GPU utilisation. Utilisation reads high at any batch, so the fleet never scales in. Scale on requests in flight per replica.
- Cost model built on average context. A few long conversations consume KV out of proportion and cap the batch for everyone.
- Dynamic batching delays at low load. A batching window that waits to fill adds its full delay to every request when traffic is light. Bound the wait.
- Benchmarks with ignore-EOS and uniform prompts only. Real output lengths vary, batches drain unevenly and achieved batch is lower. Validate against production.
Trade-offs
Bigger batches trade latency for cost, smoothly until saturation and then abruptly. Fewer replicas trade resilience for efficiency, so keep enough headroom to lose one replica at peak. Tiered service trades scheduling complexity for filling batches with work that pays. Bucketing trades a little latency for less padding. The choice between these strategies by workload is laid out in GPU Batching Strategies.
What to do next
- Plot requests in flight per replica from production metrics next to your configured maximum.
- Run the concurrency sweep with real prompts and find where cost per token flattens and where p99 time per output token crosses your target.
- Compute the smallest replica count that keeps peak traffic, minus one replica, below saturation, and move toward it in steps.
- Switch autoscaling and routing to requests in flight per replica.
- Measure padding efficiency on static-batch jobs and add length bucketing where it is below 70 percent.
- Move deadline-tolerant work to a batch API or a low-priority queue that fills troughs.