People describe GPU work by what it is for: training, inference, rendering, simulation. That label is useful for budgeting and useless for engineering. Two inference services can behave in opposite ways on the same card, one pinned at the tensor-core ceiling and the other idling its math units while it waits on memory. What decides how a job runs, what it costs and what will make it faster is the resource it saturates first.

This article builds a taxonomy on that idea: the resources a GPU has, arithmetic intensity to predict which one a kernel exhausts, where common workloads land, how to classify a job from hardware counters, and what each class implies for scheduling and tuning.

Advertisement

A GPU is several resources, not one

A data-center GPU gives a job a small set of resources, and each can run out independently. Tensor throughput is the rate at which tensor cores multiply matrices, quoted in TFLOPS per precision. Memory bandwidth is how fast bytes move between HBM and the chip. Memory capacity decides what fits at all: weights, activations, optimizer state, KV cache. Interconnect, NVLink inside a node and InfiniBand or Ethernet between nodes, carries gradients and sharded tensors. Host paths, PCIe, CPU and storage, feed data in. Finally there are specialised engines that only some workloads touch: FP64 pipes for scientific code, RT cores for ray tracing, and NVENC and NVDEC for video.

The specialised engines matter more than their size suggests, because they differ by product. NVIDIA's H100 SXM datasheet lists 7 NVDEC and 7 JPEG decoders and no NVENC encoders, while the L4 lists 2 NVENC encoders, 4 NVDEC decoders and RT cores. A transcoding job on an H100 therefore cannot use hardware encoding at all, however fast the tensor cores are. A workload type is, in practice, a statement about which of these resources binds first.

The GPU selection guide uses these resources to pick hardware. Here the direction is reversed: given a job, which resource is it consuming?

Arithmetic intensity and the ridge point

For most GPU code, the deciding question is whether math or memory traffic runs out first. Arithmetic intensity answers it: the number of floating-point operations a kernel performs per byte it moves to or from HBM. The roofline model says attainable throughput is the smaller of peak compute and bandwidth multiplied by intensity. The intensity at which the two are equal is the ridge point. Below it a kernel is memory-bound and faster math units do nothing for it. Above it the kernel is compute-bound and extra bandwidth is wasted.

The code below computes intensity for the operations that dominate modern models, using the H100 SXM datasheet figures: about 989 dense BF16 tensor TFLOPS (NVIDIA quotes 1,979 with 2:4 sparsity, which ordinary dense training does not use) and 3.35 TB/s of HBM bandwidth.

# Arithmetic intensity: FLOPs per byte of HBM traffic. bf16 = 2 bytes per element.
def gemm_intensity(m, n, k, bytes_per=2):
    flops = 2 * m * n * k                              # one multiply + one add per term
    traffic = bytes_per * (m * k + k * n + m * n)      # read A and B, write C once
    return flops / traffic

def elementwise_intensity(flops_per_elem=1, reads=1, writes=1, bytes_per=2):
    return flops_per_elem / (bytes_per * (reads + writes))

PEAK_FLOPS = 989e12        # H100 SXM, dense BF16 tensor (NVIDIA datasheet)
PEAK_BW = 3.35e12          # H100 SXM, HBM3 bytes per second
RIDGE = PEAK_FLOPS / PEAK_BW

cases = {
    "training GEMM 16384x8192x8192": gemm_intensity(16384, 8192, 8192),
    "prefill GEMM 2048x8192x8192":   gemm_intensity(2048, 8192, 8192),
    "decode batch 64":               gemm_intensity(64, 8192, 8192),
    "decode batch 1 (GEMV)":         gemm_intensity(1, 8192, 8192),
    "residual add":                  elementwise_intensity(1, reads=2, writes=1),
}
print(f"ridge = {RIDGE:.0f} FLOP/byte")
for name, ai in cases.items():
    ceiling = min(PEAK_FLOPS, ai * PEAK_BW) / 1e12
    bound = "compute" if ai >= RIDGE else "memory"
    print(f"{name:32s} {ai:8.1f} FLOP/B  {bound:7s} ceiling {ceiling:6.1f} TFLOPS")

Running it prints a ridge of about 295 FLOP/byte. A training-sized GEMM reaches about 3,277 and a 2,048-token prefill GEMM about 1,365, both well above the ridge and therefore compute-bound. Decode at batch 1 is a matrix-vector product with intensity near 1.0, so its ceiling is about 3.3 TFLOPS: roughly a third of one percent of the tensor peak. Batch 64 lifts decode to about 63 FLOP/byte, still memory-bound but about sixty times more work per byte. A residual add sits near 0.17. The memory side is covered in more depth in the HBM architecture article.

Two cautions. The formula assumes each operand is read once, which only a well-tiled kernel achieves; a poor kernel moves more bytes and lands further left. And intensity belongs to a kernel, not a job: one training step mixes GEMMs far right of the ridge with elementwise work far left of it.

Arithmetic intensity (FLOPs per byte moved from HBM), log scaleTFLOPSmemory roof: bandwidth x intensitycompute roof: peak tensor FLOPSridge point (~295 on H100 BF16)elementwiseLLM decode, batch 1embedding lookupsdecode, batch 64prefill, diffusiontraining GEMMsOff the chartNVENC/NVDEC, RT cores, FP64,network and input pipeline
The roofline with common workloads placed by arithmetic intensity. Points left of the ridge are capped by bandwidth; points right of it by tensor throughput. Several workload types are bound by resources the roofline does not draw at all.
Advertisement

The taxonomy: workload types by binding resource

Combining intensity with the other resources gives a table that is more useful than the training versus inference split. Read it as a set of starting hypotheses to confirm with counters, not as a set of rules.

WorkloadDominant workBinds firstCounter to check first
LLM pretrainingLarge GEMMs, attention, all-reduce or sharded collectivesTensor throughput, then interconnect at scaleTensor active, NVLink and network bytes
Fine-tuning, LoRASame GEMMs at smaller scale; optimizer stateMemory capacity, then tensor throughputFramebuffer used, tensor active
LLM prefillLarge GEMMs over the whole promptTensor throughputTensor active
LLM decodeGEMV-like steps; KV cache readsBandwidth, then KV cache capacityDRAM active, framebuffer used
Embedding and encoder batch jobsMid-size GEMMs, tokenisation on CPUOften the host input pipelineEngine active below SM active expectations
Recommender inferenceEmbedding-table lookups, small MLPsBandwidth and capacity; often hostDRAM active, PCIe bytes
Image and video diffusionConvolutions and attention over latents, many stepsTensor throughputTensor active
Rendering, visualisationRasterisation, ray tracing via graphics APIsRT cores and graphics pipelineGraphics tools, not tensor counters
Video transcodeEncode and decodeNVENC and NVDEC enginesnvidia-smi encoder and decoder utilisation
Scientific HPCFP64 solvers, stencils, FFTsFP64 pipes or bandwidthFP64 pipe active, DRAM active

Three things stand out. First, inference is not one type: prefill and decode sit on opposite sides of the ridge, which is why modern servers batch decode aggressively and why some split the two phases onto different hardware. The inference latency article turns this into time-to-first-token and per-token models. Second, many jobs described as GPU workloads are actually bound by something off the GPU: the data loader, the tokeniser, the network. Third, rendering, video and FP64 work are limited by hardware that the tensor-core headline numbers do not describe, so a card that is excellent for training can be a poor or impossible choice for them.

Training up close: why one step contains several types

A training step is the clearest example of a job that is not one workload type. The forward and backward GEMMs are compute-bound, and their share of step time rises with model width and tokens per GPU. Layer norms, activations, dropout and the optimizer update are elementwise and memory-bound. On data-parallel or sharded runs, gradient reduction or parameter gathering uses the interconnect, and whether it overlaps with compute decides whether it costs anything. Around all of it, the data loader has to deliver the next batch before the GPU asks for it.

So the right question for training is what fraction of step time the GEMMs take and what fills the rest. Fusion and compiled graphs attack the elementwise tail, larger micro-batches and better sharding cut communication, and prefetching removes loader stalls. The training step anatomy article walks through each phase with pseudocode.

Classifying a running job from counters

You rarely know in advance what a production job binds on, especially one another team wrote. The coarse GPU utilisation figure that nvidia-smi prints cannot tell you: it reports the fraction of time any kernel was running, so one tiny kernel on one SM counts as fully busy. The datacenter profiling fields in NVIDIA DCGM read hardware counters instead and report ratios between 0 and 1. Four of them separate the main classes: graphics engine active (field 1001), SM active (1002), tensor pipe active (1004) and DRAM active (1005). The DCGM article explains the host engine, sampling costs and dcgm-exporter; the sketch below turns those four ratios into a class.

# Classify a job from DCGM profiling ratios (0..1) averaged over a steady window.
# Collect with: dcgmi dmon -e 1001,1002,1004,1005 -d 1000 -c 300 -i 0
#   1001 GR_ENGINE_ACTIVE  1002 SM_ACTIVE  1004 PIPE_TENSOR_ACTIVE  1005 DRAM_ACTIVE
# Thresholds are starting heuristics, not NVIDIA guidance: calibrate them on
# jobs whose behaviour you already know before trusting them on new ones.
def classify(s):
    if s["engine"] < 0.5:
        return "starved: waiting on the host, input pipeline or collectives"
    if s["sm"] < 0.5:
        return "under-filled: kernels too small or too few; batch, fuse or use CUDA graphs"
    if s["tensor"] >= 0.4:
        return "tensor-bound: training, prefill, diffusion"
    if s["dram"] >= 0.6:
        return "bandwidth-bound: decode, embeddings, elementwise-heavy code"
    return "non-tensor compute or latency-bound: profile kernels with Nsight Compute"

def steady_mean(rows, key, skip=30):
    vals = [r[key] for r in rows[skip:]]          # drop warm-up and compilation
    return sum(vals) / len(vals)

The order of the checks is deliberate. A starved GPU looks memory-light and tensor-light, so testing engine activity first stops it from being mistaken for a small workload. Under-filled SMs come next because a job running tiny kernels cannot reach either roof, and batching or fusion helps it more than any hardware change. Only then do tensor and DRAM activity decide between the two sides of the ridge. For video jobs, use the encoder and decoder columns of nvidia-smi dmon -s u; for FP64 code, add the FP64 pipe field. When the classifier says non-tensor compute, move to Nsight Compute, because the answer now lives in individual kernels.

Worked example: three jobs on one node

Consider a shared eight-GPU node running three jobs, with steady-state DCGM averages as below. The numbers are illustrative, chosen to show the method, not measured on a particular system.

JobEngineSMTensorDRAMClass
A: 7B fine-tune, 4 GPUs0.970.930.520.35tensor-bound
B: chat model serving, 2 GPUs0.950.880.060.74bandwidth-bound
C: embedding backfill, 2 GPUs0.410.380.180.20starved

Job A is where it should be: tensor-heavy, with a healthy SM ratio. Improvements come from reducing the non-GEMM share of the step, not from more hardware. Job B is decode-dominated. Its tensor units are almost idle and the fix is on the memory side: raise the running batch with continuous batching until latency targets bind, quantise the weights or KV cache to cut bytes per token, and check whether KV cache capacity is what stops the batch from growing. Job C is labelled as GPU work but is really a host job: the GPU is idle more than half the time. Profiling shows single-threaded tokenisation on the CPU. Moving tokenisation to parallel workers and prefetching two batches ahead lifts engine activity above 0.9, and the backfill finishes in less than half the time on the same two GPUs.

What the class means for placement and sharing

The class also tells you what can share a GPU. Two bandwidth-bound services on one card compete for the same resource, so co-locating them roughly halves each one's throughput. A starved job is the best candidate for sharing, because most of its GPU time is idle, but the first fix is to stop starving it. MIG partitions a supported GPU into isolated slices with their own memory and compute, which suits several small inference services with strict isolation needs. MPS lets processes run kernels concurrently on the same SMs without that isolation. Time-slicing only interleaves, so it adds latency without adding throughput.

For placement, tensor-bound jobs want the highest tensor throughput per dollar, and large training jobs also want the fastest interconnect. Bandwidth-bound decode wants bandwidth and memory capacity per dollar, which is why older or lower-tier cards with good HBM can be cost-effective for serving. Video wants cards with encoders. FP64 simulation wants cards with full-rate FP64, which several inference-oriented products lack. Put the class in the scheduler's metadata, and the placement decisions follow from it.

Failure modes

  • Trusting GPU-Util. A job at 100 percent utilisation can have SM activity of 0.2. Use the profiling fields.
  • Classifying from the job name. A serving job with long prompts and short answers is prefill-dominated and tensor-bound; the same model with short prompts and long answers is decode-bound.
  • Measuring during warm-up. Compilation, cache fills and autotuning distort the first minutes. Average over a steady window only.
  • Buying compute for a bandwidth problem. A faster tensor core does nothing for batch-1 decode. Check the ridge before upgrading.
  • Ignoring the host. Starved GPUs are common and the cheapest to fix; the fix is usually CPU parallelism or I/O, not a GPU change.
  • Assuming all GPUs have the same engines. Encoders, RT cores and FP64 rates differ by product. Check the datasheet for your exact model.

What to do next

  1. Pick your three most expensive GPU jobs and record steady-state averages of DCGM fields 1001, 1002, 1004 and 1005 for each.
  2. Classify each with the rules above, then confirm the class with one Nsight Systems trace per job.
  3. Compute the arithmetic intensity of the dominant kernels with the formula in this article, and compare it with your GPU's ridge point.
  4. For every starved job, fix the input path before changing anything on the GPU.
  5. For bandwidth-bound serving, raise the batch and reduce bytes per token before buying faster compute.
  6. Store each job's class in scheduler metadata and use it for placement and sharing rules, then re-measure after every major model or traffic change.
Key takeaway: A GPU workload type is best defined by the resource it exhausts first: tensor throughput, memory bandwidth, memory capacity, interconnect, the host path, or a specialised engine. Arithmetic intensity against the ridge point predicts the compute-versus-memory split, and DCGM's engine, SM, tensor and DRAM activity fields confirm it for jobs you did not write. Classify by measurement, fix starvation first, and let the class decide placement, sharing and which optimisation to try next.