Most GPU buying advice is a list of model names sorted by price, and it goes stale within a year. Two teams with the same budget can need completely different hardware: one serves a 7-billion-parameter model to thousands of users with short prompts, the other fine-tunes a 30-billion-parameter model on long documents. The first is limited by memory bandwidth and cost per token; the second is limited by memory capacity and the interconnect between cards.

This article gives you a method that outlives any generation of parts. You profile the workload, compute how much memory it needs, estimate which resource it saturates first, check whether it must be split across GPUs, and then rank the parts that pass by cost per unit of useful work. Specifications quoted here were checked against vendor datasheets in September 2026; newer parts will keep arriving, and the method applies to them unchanged.

Advertisement

The question behind every GPU choice

A GPU is a bundle of four resources: memory capacity (how many bytes fit on the card), memory bandwidth (how many bytes per second the compute units can read from that memory), compute throughput (how many floating-point operations per second the tensor cores can perform at a given precision), and interconnect (how fast the card can exchange data with its neighbours). Every workload exhausts one of these first.

Different phases of machine learning hit different limits. Training and the prefill phase of inference, where the whole prompt is processed at once, do many operations per byte loaded, so they are usually compute-bound. The decode phase of inference, which generates one token per step and must read every weight for each step, does very few operations per byte and is bandwidth-bound. Large models that do not fit on one card are capacity-bound and then often interconnect-bound, because splitting them requires constant communication. The diagram below is the order in which to check.

Choose a GPU by finding the first resource your workload runs out ofWorkload profilemodel size, context, batch, phase1. Memory capacityweights + KV cache + optimizer2. Memory bandwidthdecode: bytes read per token3. Computeprefill and training: FLOPs4. InterconnectNVLink vs PCIe when sharded5. Software and opskernels, FP8, MIG, drivers6. Availabilityquota, region, lead timedoes not fit? shardCandidate set: parts that fit and meet latencyrank by cost per unit of work: $/million tokens, $/training runBenchmark the top two on your own workload
The selection method. Capacity is a hard filter; bandwidth and compute set speed; interconnect matters once the model is sharded; software, operations and availability prune the list; cost per unit of work ranks what remains.

Step 1: memory capacity is a hard filter

Start with arithmetic you can do on paper. Model weights take parameters multiplied by bytes per parameter: 2 bytes at BF16 or FP16, 1 byte at FP8 or INT8, about half a byte at 4-bit. A 32-billion-parameter model therefore needs 64 GB at BF16 and 32 GB at FP8 before anything else is loaded. FP8 quantization is often the single largest lever on which cards are even eligible.

Inference adds the KV cache: for every token of every active sequence the model stores a key and a value vector per layer. Its size per token is 2 multiplied by layers, key-value heads, head dimension and bytes per element. For a hypothetical model with 64 layers, 8 key-value heads (grouped-query attention) and head dimension 128 at 2 bytes, that is 262,144 bytes, or 256 KiB per token. Thirty-two concurrent sequences at 8,192 tokens each need 64 GiB of KV cache, which is as much as the BF16 weights. Long context and high concurrency buy capacity, not compute.

Training needs far more. Mixed-precision training with the Adam optimizer typically holds about 16 bytes per parameter: 2 for BF16 weights, 2 for gradients, 4 for an FP32 master copy and 8 for the two FP32 optimizer moments. A 7B model needs about 112 GB for that state alone, before activations, so a full fine-tune does not fit on one 80 GB card without sharding the optimizer state (ZeRO or FSDP) or switching to parameter-efficient methods such as LoRA, which train a small adapter while the frozen base can even stay quantized.

def serving_memory_gb(params_b, weight_bytes, layers, kv_heads, head_dim,
                      kv_bytes, concurrent, context, overhead=0.10):
    """Rough device memory needed to serve one model replica."""
    weights = params_b * 1e9 * weight_bytes
    kv_per_token = 2 * layers * kv_heads * head_dim * kv_bytes
    kv = kv_per_token * concurrent * context
    return (weights + kv) * (1 + overhead) / 1e9   # activations, fragmentation, runtime

def training_memory_gb(params_b, bytes_per_param=16):
    """Weights + grads + FP32 master + Adam moments, excluding activations."""
    return params_b * 1e9 * bytes_per_param / 1e9

print(serving_memory_gb(32, 1, 64, 8, 128, 2, concurrent=32, context=8192))  # ~110.8
print(training_memory_gb(7))                                                  # 112.0
Advertisement

Step 2: bandwidth sets decode speed

During decode, each step reads all weights once (shared across the batch) plus the KV cache of every sequence in the batch, and performs only a couple of floating-point operations per weight byte. The GPU finishes the arithmetic long before the bytes arrive, so step time is roughly bytes read divided by memory bandwidth. This gives a useful upper bound: for a single sequence, tokens per second cannot exceed bandwidth divided by weight bytes. Real systems reach perhaps 60 to 80 percent of it.

Batching is how serving systems escape the bound: one weight read serves many sequences, so aggregate throughput rises almost linearly until either the KV cache fills memory or the batch becomes large enough to turn compute-bound. That is why capacity and bandwidth together decide inference economics. High-bandwidth memory is the reason data-center parts dominate here: HBM stacks deliver several times the bandwidth of the GDDR memory on PCIe and consumer cards.

PartMemoryBandwidthClassTypical fit
NVIDIA L424 GB GDDR6300 GB/sPCIe, low powersmall-model inference, video, embeddings
NVIDIA L40S48 GB GDDR6864 GB/sPCIemid-size inference, fine-tuning small models
GeForce RTX 409024 GB GDDR6X1,008 GB/sconsumerdevelopment, local inference, research
GeForce RTX 509032 GB GDDR71,792 GB/sconsumerdevelopment, local inference
NVIDIA A100 80GB (SXM)80 GB HBM2eabout 2.0 TB/sdata centertraining and inference, still widely rented
NVIDIA H100 (SXM)80 GB HBM33.35 TB/sdata center, NVLinktraining and large-model serving
NVIDIA H200141 GB HBM3e4.8 TB/sdata center, NVLinkmemory-bound serving, long context
AMD Instinct MI300X192 GB HBM35.3 TB/sdata center, Infinity Fabriclarge-model serving on ROCm
NVIDIA B200 (HGX)180 GB HBM3eabout 8 TB/sdata center, NVLink 5frontier training and serving

Capacity filters before speed ranks: a 24 GB card cannot hold a 32B model even at FP8. Among parts that fit, decode throughput tracks bandwidth, so an H200 decodes noticeably faster than an H100 despite similar tensor-core throughput.

Step 3: compute sets prefill and training time

Training cost is usually estimated with the rule that one training token costs about 6 floating-point operations per parameter (2 for the forward pass, 4 for the backward pass). A 7B model trained or continued-pretrained on 20 billion tokens needs about 6 x 7e9 x 20e9 = 8.4e20 FLOPs. Divide by delivered throughput: an H100 SXM is rated at roughly 989 dense BF16 TFLOPS (the datasheet's 1,979 assumes sparsity), and well-tuned jobs achieve 35 to 50 percent of peak (model FLOPs utilization, MFU). At 40 percent that is about 590 GPU-hours. Compute-bound workloads scale with the tensor-core rating at the precision you actually use, so check whether your framework and kernels support FP8 on the target part before counting its FP8 numbers.

Prefill behaves like training's forward pass: a 4,000-token prompt is one large matrix multiplication per layer, and time to first token is set by compute. Long-prompt, short-answer workloads such as retrieval-augmented QA are therefore more compute-sensitive than chat.

Step 4: interconnect decides whether sharding works

When a model or its training state exceeds one card, you split it. Tensor parallelism splits each layer across GPUs and exchanges activations twice per layer per step, so it needs the fastest link available; it is normally confined to GPUs inside one server connected by NVLink (900 GB/s per GPU on H100 systems), described in NVLink and NVSwitch. Pipeline and data parallelism communicate less often and can span servers over InfiniBand or RoCE.

PCIe-only cards, including L40S, L4 and consumer GeForce parts, exchange data over PCIe, which is an order of magnitude slower than NVLink. Tensor parallelism over PCIe works for small degrees but adds latency to every token. If your plan requires splitting a model four or eight ways, the practical choice is an NVLink server or a bigger-memory part that avoids the split.

Worked example: serving a 32B model

Suppose you must serve the hypothetical 32B model above to an internal assistant: up to 32 concurrent conversations, context up to 8,192 tokens, and a target of at least 25 tokens per second per user. Compute requirements before looking at prices.

  1. Capacity. At FP8 weights and a BF16 KV cache: 32 GB of weights plus 68.7 GB (64 GiB) of worst-case KV cache plus about 10 percent overhead gives roughly 111 GB. That rules out every 24 to 48 GB card for a single replica and rules out a single 80 GB card unless you also store the KV cache in FP8, which halves it to about 34 GB and brings the total to about 73 GB: tight but feasible.
  2. Bandwidth. With 32 sequences averaging 4,096 tokens, each decode step reads 32 GB of weights plus about 34 GB of KV cache. On an H100 (3.35 TB/s) the upper bound is about 50 steps per second, roughly 1,600 tokens per second aggregate; that is conservative, since the FP8 KV cache the H100 needs halves the KV read. On an H200 (4.8 TB/s) it is about 72 steps per second. Both exceed 25 tokens per second per user even at 60 percent efficiency.
  3. Candidates. One H200 or MI300X holds the model and full BF16 KV cache on one card with room to grow; one H100 works only with an FP8 KV cache and a hard cap on concurrency; two L40S cards (96 GB together) with tensor parallelism over PCIe fit only with an FP8 KV cache, and give about a third of an H200's bandwidth per replica plus PCIe communication latency on every token.
  4. Decision. Price each candidate by cost per million output tokens at your measured throughput, then benchmark the top two with your real prompts. The arithmetic only tells you which benchmarks are worth running.

Cost per unit of work, and rent versus buy

An hourly price is not a cost. What you care about is dollars per million generated tokens, dollars per training run, or dollars per thousand embeddings, and those depend on throughput. A card that costs twice as much per hour but delivers three times the throughput on your workload is the cheaper card. The formula is simple, and writing it down forces you to measure throughput honestly:

def cost_per_million_tokens(hourly_price, tokens_per_second, utilization):
    """utilization: fraction of paid hours doing useful work (0..1)."""
    tokens_per_hour = tokens_per_second * 3600 * utilization
    return hourly_price / tokens_per_hour * 1e6

# Same workload, two candidate parts, illustrative relative prices only.
a = cost_per_million_tokens(hourly_price=1.0, tokens_per_second=1000, utilization=0.6)
b = cost_per_million_tokens(hourly_price=1.8, tokens_per_second=2300, utilization=0.6)
print(round(a, 3), round(b, 3))   # b is cheaper per token despite a higher hourly rate

Utilization dominates the rent-versus-buy question. Owned hardware is paid for whether or not it runs, so it wins only when you can keep it busy for most of its multi-year life and have the power, cooling and staff to run it. Rented capacity costs more per hour but costs nothing when idle, and it lets you move to a newer generation without a write-off. Reserved or committed cloud capacity sits in between. GPU cost optimization covers spot capacity, right-sizing and scheduling, and fine-tuning costs works through training budgets in more detail.

Software, operations and availability

Hardware that your software stack cannot use well is slow hardware. Check four things before committing. Kernel support: does your serving engine or training framework have optimized kernels for this architecture and precision, including FP8 and attention kernels? CUDA support is the broadest; ROCm support for AMD Instinct parts has improved considerably but still deserves a benchmark on your exact model and engine. Partitioning: Multi-Instance GPU (MIG) on A100, H100 and newer data-center parts splits one card into isolated slices, which is valuable for many small models; Ada-generation cards such as L40S and L4 do not offer it. Licensing and placement: consumer GeForce cards are excellent for development, but check the driver licence terms before deploying them in a data center, and note that they lack NVLink.

Availability is a real constraint too: the best part on paper is useless without quota in your region. Keep a benchmarked second choice.

Failure modes in GPU selection

  • Sizing for weights only. The model fits, then production traffic with long contexts exhausts the KV cache, the engine preempts sequences and latency collapses. Size for worst-case concurrency multiplied by context.
  • Buying peak TFLOPS for a decode workload. Chat serving is bandwidth-bound; a part with more compute but the same bandwidth does not generate tokens faster.
  • Counting sparse or FP4 peaks you will not use. Datasheets headline the lowest-precision, sparsity-assisted number. Compare at the precision and density your kernels actually run.
  • Sharding over PCIe. Tensor parallelism across PCIe cards fits the model but adds communication latency to every token.
  • Ignoring utilization. An idle owned cluster is the most expensive GPU there is.

What to do next

  1. Write down your workload profile: model parameters, layers, KV heads, head dimension, precision, peak concurrency, maximum and average context, and whether you train, fine-tune or only serve.
  2. Compute weights, KV cache and, for training, optimizer state with the functions above; discard every part that cannot hold one replica.
  3. For serving, compute the decode upper bound from bandwidth; for training, compute GPU-hours from 6ND and a realistic MFU.
  4. If the model must be split, decide the parallelism and confirm the interconnect supports it.
  5. Check kernel support for your precision on each surviving part and confirm quota or lead time.
  6. Benchmark the top two candidates with replayed production prompts and rank them by cost per million tokens or per training run, not by hourly price.
Key takeaway: Choose a GPU by finding the resource your workload exhausts first. Capacity filters: weights, KV cache and optimizer state must fit. Bandwidth sets decode speed, compute sets prefill and training time, and interconnect decides whether sharding is practical. Among the parts that pass, rank by cost per unit of useful work at realistic utilization, then confirm with a benchmark on your own traffic.