Small language models, roughly the 0.5 to 8 billion parameter range, are chosen because they can run on a laptop, a phone or a single modest GPU. That makes their speed a hardware budget problem. There are many published optimizations, from quantization to speculative decoding, and each has its own guide on this site. What is usually missing is the step before choosing one: working out which resource your model is actually short of, so that you apply the technique that helps instead of the one that is fashionable.
This article is about that method. It defines the two latencies users feel, explains from first principles why prompt processing and token generation are limited by different hardware resources, works the arithmetic for a 1-billion-parameter model on an illustrative device, and then maps each technique to the bottleneck it relieves. It ends with a benchmark harness, an order of operations and the failure modes that make optimizations look better in a benchmark than in the product.
Two latencies, two phases
A request to a decoder-only model runs in two phases. Prefill processes the whole prompt in one pass, computing attention keys and values for every prompt token and storing them in the KV cache. Decode then generates output one token at a time; each step runs the full model on a single new token and appends to the cache.
Users feel these as two numbers. Time to first token (TTFT) is dominated by prefill and decides whether an assistant feels responsive. Time per output token (TPOT), sometimes called inter-token latency, is set by decode and decides how fast text streams. A chat reply of 200 tokens at 40 ms per token takes eight seconds after the first token, so both matter, and optimizations usually help one much more than the other.
Why decode is memory-bound
A transformer layer is mostly matrix multiplication. During decode at batch size 1, each weight is read from memory, used for one multiply-add with the current token's activation, and not needed again until the next step. That is about 2 floating-point operations per weight, while a 16-bit weight is 2 bytes. The arithmetic intensity, operations per byte moved, is therefore around 1.
Hardware has a balance point, often called the ridge of the roofline model: peak operations per second divided by memory bandwidth. Take an illustrative device with 50 GB/s of memory bandwidth and 4 TFLOPS of usable throughput; these round numbers stand in for no particular chip, and you should measure your own. Its ridge is 4e12 / 50e9 = 80 operations per byte. Decode, at about 1, sits far below it, so the processor spends most of each step waiting for weights to arrive. The useful mental model follows directly: decode speed at batch 1 is roughly memory bandwidth divided by the bytes read per token.
Prefill is different. A 1,000-token prompt reuses each weight 1,000 times per read, so its intensity is in the hundreds and it is limited by arithmetic throughput. That is why prompt processing can run at hundreds or thousands of tokens per second on hardware that generates only tens.
The arithmetic, worked
Take a 1B-parameter model with 16 layers, 8 key-value heads of dimension 64, on the 50 GB/s, 4 TFLOPS device. The function below computes the ceilings.
def decode_ceiling(params, bits_per_weight, bandwidth_gbs, ctx=0,
layers=16, kv_heads=8, head_dim=64, kv_bytes=2):
"""Upper bound on tokens/s at batch size 1: every weight and the whole KV cache read per step."""
weight_bytes = params * bits_per_weight / 8
kv_per_token = 2 * layers * kv_heads * head_dim * kv_bytes # K and V
bytes_per_step = weight_bytes + kv_per_token * ctx
return bandwidth_gbs * 1e9 / bytes_per_step
def prefill_seconds(params, prompt_tokens, effective_tflops):
"""About 2 FLOPs per parameter per token for the matrix multiplies."""
return 2 * params * prompt_tokens / (effective_tflops * 1e12)
print(decode_ceiling(1e9, 16, 50)) # 25.0 tokens/s at fp16
print(decode_ceiling(1e9, 4.5, 50)) # ~88.9 at ~4.5 bits incl. scales
print(decode_ceiling(1e9, 4.5, 50, ctx=4096)) # ~71.8 once a 4k-token KV cache is read too
print(prefill_seconds(1e9, 1000, 4)) # 0.5 s to first token for a 1,000-token prompt| Configuration | Bytes read per step | Decode ceiling |
|---|---|---|
| fp16 weights, short context | 2.0 GB | 25 tokens/s |
| about 8.5 bits per weight | 1.06 GB | about 47 tokens/s |
| about 4.5 bits per weight | 0.56 GB | about 89 tokens/s |
| 4.5 bits, 4,096-token KV cache in fp16 | 0.70 GB | about 72 tokens/s |
Three conclusions fall out. First, weight precision is the biggest lever for decode: going from 16 bits to about 4.5 bits (4-bit values plus per-group scales) raises the ceiling roughly 3.5 times. Second, the KV cache is not free. Each token of context costs 2 x 16 x 8 x 64 x 2 = 32 KB here, so 4,096 tokens add 128 MiB to every step, and at long context the cache becomes the next thing to shrink. Third, prefill of 1,000 tokens needs about 2 x 109 x 1,000 = 2 TFLOP, which is half a second on this device: a long system prompt costs real TTFT on every request unless it is cached.
These are ceilings, not predictions. Real runtimes lose some of the bandwidth to activation traffic, dequantization work and scheduling gaps, so how close a given runtime gets varies by device and should be measured rather than assumed. A measured number far below the ceiling points at a software problem such as unfused kernels, thread contention or thermal throttling, not at the model. A measured number close to it tells you the opposite: no amount of kernel tuning will help, and only reading fewer bytes per token, or reusing each read for more tokens, can make decode faster.
The same arithmetic sizes memory. The 4.5-bit weights need about 0.56 GB, a 4,096-token cache adds 128 MiB per concurrent sequence, and the runtime needs working buffers on top. On a phone that shares memory with the operating system and other apps, that total, not the parameter count, decides whether the model can stay resident.
Match the technique to the bottleneck
| Technique | Attacks | Main effect | Main risk |
|---|---|---|---|
| Weight quantization | Decode bandwidth, memory footprint | Fewer bytes per step | Quality loss, especially below 4 bits |
| Smaller KV cache (GQA models, KV quantization, window limits) | Decode at long context, memory | Fewer bytes per step, more context fits | Recall loss at long range |
| Prefix caching | Prefill of repeated prompts | TTFT drops to the uncached suffix | Cache memory, invalidation on prompt change |
| Fused or flash-style attention | Prefill and long-context decode | Fewer passes over memory | Kernel availability per device |
| Speculative decoding | Decode latency | Several tokens per weight read | Draft overhead when acceptance is low |
| Continuous batching | Server throughput | One weight read serves many requests | Per-request latency rises with batch |
Each has a dedicated guide: quantization, KV cache architecture, prefix caching, speculative decoding and batching. The point of the table is the second column. A technique aimed at the wrong phase does nothing: weight quantization is not what sets TTFT on a compute-bound prefill, and prefix caching does nothing for streaming speed.
Speculative decoding, in numbers
Speculative decoding exploits the gap between decode's arithmetic intensity and the ridge. A cheap draft model proposes k tokens, and the target model checks all k in one forward pass, which costs about the same weight reads as generating one. Leviathan and colleagues showed in 2023 that with a per-token acceptance rate a, the expected number of tokens produced per target pass is (1 - ak+1) / (1 - a). With a = 0.7 and k = 4 that is about 2.77 tokens per pass.
The gain is smaller than that figure because the draft costs time too, and it collapses when the draft disagrees with the target, as it tends to on code, rare languages or high-temperature sampling. For a small target model the draft must be very small, or an n-gram or prompt-lookup scheme, or the overhead eats the benefit. Measure acceptance on your own traffic before you commit.
An order of operations
- Measure first with the harness below: TTFT and TPOT at p50 and p95, tokens per second, and peak memory, on prompts shaped like production.
- Compare with the ceilings. If you are far below the bandwidth ceiling, fix the software path before touching the model: use the runtime's optimized kernels, set thread counts to physical cores, and keep the model resident.
- Quantize the weights to the lowest precision that passes your quality evaluation. This is usually the largest single win for on-device decode.
- Cap and shrink the KV cache if your contexts are long: set the maximum context you actually need, prefer models with grouped-query attention, and evaluate KV quantization.
- Cache prefixes if requests share a long system prompt or document, to cut TTFT.
- Add speculative decoding for latency-sensitive single-user decode, if acceptance on real traffic is high.
- Batch on servers where throughput per device matters more than the latency of any one request.
import statistics, time
def bench(stream, prompts, warmup=2):
"""stream(prompt) must yield tokens as they are produced; supply your runtime's API."""
for pr in prompts[:warmup]:
for _ in stream(pr):
pass # load weights, fill caches, settle clocks
ttft, tpot, tokens = [], [], 0
t_all = time.perf_counter()
for pr in prompts:
t0 = time.perf_counter()
stamps = [time.perf_counter() for _ in stream(pr)]
if not stamps:
continue
ttft.append(stamps[0] - t0)
if len(stamps) > 1:
tpot.append((stamps[-1] - stamps[0]) / (len(stamps) - 1))
tokens += len(stamps)
wall = time.perf_counter() - t_all
q = lambda xs, f: sorted(xs)[min(len(xs) - 1, int(f * len(xs)))]
return {
"ttft_p50_ms": 1000 * statistics.median(ttft), "ttft_p95_ms": 1000 * q(ttft, 0.95),
"tpot_p50_ms": 1000 * statistics.median(tpot), "tpot_p95_ms": 1000 * q(tpot, 0.95),
"tokens_per_s": tokens / wall,
}The harness is deliberately provider-neutral: stream is a function you write around your runtime that yields tokens as they arrive. The warm-up loop matters, because the first request after load pays for page faults on memory-mapped weights, kernel compilation and cache allocation.
Failure modes
- Benchmarking the wrong shape. Short prompts and short outputs hide both prefill cost and KV growth. Use prompt and output lengths sampled from production.
- Thermal throttling. Phones and fanless laptops sustain their peak for seconds, not minutes. Run long benchmarks and report sustained, not burst, numbers.
- Quality regressions nobody measured. Aggressive quantization can pass a perplexity check and still break structured output or arithmetic. Evaluate on the task, including format validity.
- Memory pressure. On mobile, the operating system may kill a process that grows its KV cache without limit. Budget memory for weights plus maximum context plus the rest of the app.
- Cold starts. Loading a model from storage can take longer than a whole reply. Load once, keep it resident where the platform allows, and measure the cold path separately.
- Mismatched tokenizers in speculation. The draft and target must share a vocabulary, or tokens must be re-mapped, which adds overhead and bugs.
What to do next
- Write the two targets for your product: a TTFT and a TPOT at p95, on your target device.
- Wrap your runtime in a stream function and run the harness on production-shaped prompts, after warm-up and for long enough to reach thermal steady state.
- Compute the decode ceiling for your model, precision, context and measured memory bandwidth, and compare it with the measured TPOT.
- Close any large software gap first, then quantize weights, re-running your task evaluation at each precision.
- If TTFT misses its target, measure prompt length, cache shared prefixes and trim the system prompt.
- Only then try speculative decoding or batching, and keep each change only if the harness and the quality evaluation both improve.