The NVIDIA H200 is a Hopper GPU with a bigger, faster memory system. It runs the same Hopper tensor cores and the same CUDA software as the H100, but swaps the H100's 80 GB of HBM3 for 141 GB of HBM3e and raises memory bandwidth from 3.35 TB/s to 4.8 TB/s. That sounds like an incremental spec bump, and for compute-bound training it nearly is. For large language model inference it is not, because LLM token generation is limited by how many bytes the GPU can hold and move, not by how many multiply-adds it can do.

This article explains the H200 from that software point of view. It covers what changed and what did not, why memory capacity and bandwidth control decode speed, a worked example that sizes a 70-billion-parameter model on one H100 and one H200, how training jobs can spend the extra capacity, how to configure a serving stack to actually use it, the failure modes teams hit, and when an H200 is the wrong purchase. Figures come from NVIDIA's published H100 and H200 product pages; tensor throughput numbers there include structured sparsity, so this article halves them when it means dense math.

What changed from the H100, and what did not

Both columns below are the SXM form factor used on eight-GPU HGX baseboards; the PCIe-card NVL variant gets its own section later.

H100 SXMH200 SXM
GPU memory80 GB HBM3141 GB HBM3e
Memory bandwidth3.35 TB/s4.8 TB/s
FP8 tensor (with sparsity)3,958 TFLOPS3,958 TFLOPS
BF16 tensor, dense (half the sparse figure)about 989 TFLOPSabout 989 TFLOPS
Max powerup to 700 Wup to 700 W
NVLink per GPU900 GB/s900 GB/s
MIGup to 7 instances of 10 GBup to 7 instances of 18 GB

Three things follow. First, the compute ceiling is unchanged, so any kernel that was already saturating the tensor cores on an H100 runs at about the same speed. Second, capacity rises by about 76 percent and bandwidth by about 43 percent, so anything that was waiting on memory gets faster or fits where it did not fit before. Third, the H200 is the same architecture, so your CUDA kernels, FlashAttention builds, FP8 Transformer Engine recipes and NCCL topology all carry over without recompilation. The upgrade question is therefore never about porting; it is about whether your workload is memory-bound.

One warning before any arithmetic: the memory you can allocate is less than the spec-sheet figure. The CUDA context, the framework's reserved pools and fragmentation all take a share, so read the real number from nvidia-smi or torch.cuda.mem_get_info() on your own stack rather than planning from 141 GB.

Why memory, not FLOPs, sets decode speed

The reason memory dominates LLM inference is arithmetic intensity: the number of floating-point operations a kernel does per byte it reads from HBM. A GPU can only hit its compute peak when intensity exceeds the ratio of peak FLOPs to peak bandwidth, called the ridge point. For dense BF16 on an H200 SXM that ratio is about 989 TFLOPS divided by 4.8 TB/s, roughly 206 FLOPs per byte. On an H100 SXM it is about 295, because the same compute sits on slower memory.

Now look at token generation. Each decode step multiplies every weight matrix by a thin activation matrix with one row per sequence in the batch. A BF16 weight is two bytes and takes part in two FLOPs (a multiply and an add) per sequence, so at batch size one the intensity is about one FLOP per byte, two hundred times below the ridge. The tensor cores sit mostly idle while the memory system streams weights. Intensity grows roughly linearly with batch size, but attention also reads each sequence's key-value cache, which does not amortise across the batch at all. In practice decode stays memory-bound across the batch sizes production servers run.

Prefill, the pass that processes the prompt, is the opposite case. Hundreds or thousands of prompt tokens share each weight read, intensity is high, and prefill is compute-bound. On an H200 prefill runs at H100 speed, which is why the H200 helps long-output chat and agent workloads more than prompt-heavy classification or embedding jobs.

# Back-of-envelope decode ceiling: one full weight read per step.
def decode_steps_per_sec(weight_bytes, kv_bytes_read, hbm_bytes_per_sec):
    return hbm_bytes_per_sec / (weight_bytes + kv_bytes_read)

W = 70e9                                     # 70B params in FP8 = ~70 GB
print(decode_steps_per_sec(W, 0, 3.35e12))   # H100 SXM: ~48 steps/s
print(decode_steps_per_sec(W, 0, 4.8e12))    # H200 SXM: ~69 steps/s

Each step produces one token for every sequence in the batch, so aggregate throughput is steps per second times batch size. The 1.43x ratio between those two lines is a ceiling, not a forecast: it ignores KV reads, kernel launch gaps and communication. The bigger lever is the batch size the memory lets you run, which is the subject of the worked example.

Worked example: a 70B model on one GPU

Take a 70-billion-parameter model with the published Llama 3 70B shape: 80 transformer layers, grouped-query attention with 8 key-value heads, and a head dimension of 128. Serve it with FP8 weights, which is the common choice on Hopper because the tensor cores run FP8 natively. The weights then occupy roughly 70 GB; a few layers such as embeddings are often kept in higher precision, so treat 70 GB as a floor.

The key-value cache stores, for every token of every live sequence, one key vector and one value vector per KV head per layer. The size per token is easy to compute:

def kv_bytes_per_token(layers, kv_heads, head_dim, dtype_bytes):
    return 2 * layers * kv_heads * head_dim * dtype_bytes   # 2 = K and V

per_tok = kv_bytes_per_token(80, 8, 128, 2)   # FP16 cache
print(per_tok)                                # 327,680 bytes = 320 KiB

def kv_tokens(gpu_gb, weights_gb, runtime_gb, per_tok):
    free = (gpu_gb - weights_gb - runtime_gb) * 1e9
    return int(free // per_tok)

print(kv_tokens(80, 70, 5, per_tok))    # H100 SXM: ~15,000 tokens
print(kv_tokens(141, 70, 5, per_tok))   # H200 SXM: ~200,000 tokens

The 5 GB runtime reserve covers the CUDA context, activation workspace, communication buffers and fragmentation; measure yours, because it varies by framework and batch shape. With that reserve, a single H100 has room for about 15,000 cached tokens. That is one or two requests with an 8,000-token context, so the GPU decodes at batch size one or two and wastes most of its bandwidth on weight reads that serve almost nobody. A single H200 has room for about 200,000 tokens, or roughly 24 concurrent 8,000-token requests. Weight reads now serve 24 sequences per step.

Serving a 70B model in FP8 on one GPU: where the memory goesH100 SXM, 80 GBweights about 70 GBruntimeKV ~5 GBH200 SXM, 141 GBweights about 70 GBKV cache about 66 GBAt 320 KiB of FP16 KV per token (80 layers x 8 KV heads x 128 dims x K and V)H100: about 15,000 tokensone or two 8K-token requestsH200: about 200,000 tokensroughly 24 concurrent 8K requestsx13 capacity
The same model on both GPUs. The extra 61 GB on the H200 goes almost entirely to KV cache, so the batch the GPU can run grows by an order of magnitude even though bandwidth rises only 43 percent.

This is the core of the H200 argument. The bandwidth increase speeds each step by up to 43 percent; the capacity increase multiplies the batch, and with it the aggregate throughput, by far more. On the H100 the usual fix is to shard the model across two GPUs with tensor parallelism, which buys KV room but adds an all-reduce to every layer and doubles the GPU count. On the H200 the same model fits on one GPU with a working cache. Two further levers stack on top: an FP8 KV cache halves the 320 KiB per token, and paged allocation avoids reserving the full context length for every request.

What training gets from the extra memory

Training gains are smaller and depend on where your step time goes. Dense transformer training at sensible batch sizes is mostly compute-bound matrix multiplication, and the H200 has the same tensor cores, so a job that already fits comfortably on H100s gets little from the switch. The extra memory still has real uses:

  • Less activation recomputation. Activation checkpointing trades roughly one extra forward pass for memory. With 141 GB, you can checkpoint fewer layers or none, which recovers that compute directly.
  • Larger micro-batches. Bigger micro-batches raise arithmetic intensity in the matrix multiplies and reduce the number of gradient-accumulation steps and their synchronisation overhead.
  • Fewer model-parallel shards. A model that needed tensor parallelism of 8 might fit with 4, or fit with fully sharded data parallelism alone. Every shard you remove deletes collective communication from the critical path.
  • Longer sequences. Attention activations grow with sequence length; long-context fine-tuning runs out of memory long before it runs out of compute.
  • Memory-bound kernels. Optimizer updates, layer norms, element-wise ops and embedding lookups read and write a lot of bytes per FLOP, so the 43 percent bandwidth increase shortens them.

To decide, profile an H100 step and split it into matrix-multiply time, memory-bound kernels, communication and recomputation. Only the last three shrink on an H200, so a step dominated by matrix multiplies barely improves unless the extra memory lets you restructure parallelism.

Configuring software to use the capacity

Nothing in a serving framework knows that it should use the extra 61 GB; it has to be told. Using vLLM as the example, the relevant settings are the fraction of GPU memory the engine may claim, the maximum context length, the tensor-parallel degree and the KV cache data type:

# One H200 serving a 70B FP8 checkpoint, no tensor parallelism.
vllm serve <your-70b-fp8-checkpoint> \
  --tensor-parallel-size 1 \
  --gpu-memory-utilization 0.92 \
  --max-model-len 32768 \
  --kv-cache-dtype fp8

Check flag names against the version you run; serving frameworks rename options between releases. Memory utilisation decides how much of the card becomes KV cache after weights load. Maximum model length caps per-request context so one long request cannot consume the whole cache. Tensor parallelism of one removes the all-reduce; raise it only if per-token latency, not throughput, is the goal. On startup the engine logs how many KV blocks it allocated; compare that with the arithmetic above.

For your own code, measure rather than assume:

import torch
free, total = torch.cuda.mem_get_info()
print(f"total {total/2**30:.1f} GiB, free {free/2**30:.1f} GiB")
# Budget KV cache from 'free' after the weights load, not from the spec sheet.

H200 NVL versus H200 SXM

The H200 NVL is a dual-slot, air-cooled PCIe card with the same 141 GB and 4.8 TB/s, a lower power ceiling of up to 600 W, and correspondingly lower tensor throughput (NVIDIA lists 1,671 BF16 TFLOPS against 1,979 for SXM, both with sparsity). Instead of an NVSwitch baseboard it uses NVLink bridges connecting two or four cards at 900 GB/s per GPU, with PCIe Gen5 to the host. MIG instances are 16.5 GB rather than 18 GB.

Choose the NVL card for standard air-cooled servers and models that fit on one, two or four GPUs. Choose SXM in an eight-GPU HGX system for eight-way tensor parallelism or mixture-of-experts all-to-all traffic. A 70B FP8 model fits on one card of either kind; the NVL's lower compute costs some prefill speed but little decode speed, because decode waits on memory that is identical on both.

Failure modes

The failures teams hit on H200 rollouts are mostly configuration, not hardware:

  • Running the H100 configuration unchanged. A deployment template copied from H100 nodes keeps tensor parallelism at 2 or a low memory-utilisation fraction, so half the card sits empty and the upgrade shows no gain.
  • Benchmarking at batch size one. A single-request latency test measures only the bandwidth ceiling and reports something under the 1.43x maximum. The capacity gain only shows up under concurrent load.
  • Planning from the spec sheet. Budgeting KV cache from 141 GB instead of measured free memory leaves the cache smaller than expected and surfaces as out-of-memory errors at peak concurrency.
  • Unbounded context. Raising maximum context to the model's limit lets a handful of long requests evict everyone else's cache, and throughput collapses under preemption even though memory looks plentiful.
  • Prefill-heavy workloads. Retrieval-augmented prompts with tens of thousands of input tokens and short answers are compute-bound; they see little improvement and may be cheaper on H100s.
  • Power and cooling. SXM parts are specified up to 700 W each. Racks provisioned for lower-power GPUs throttle under sustained load, which looks like an unexplained performance regression.

When the H200 fits, and what you trade

The H200 is the right buy when your bottleneck is memory and your software is already Hopper-tuned: large-model inference, long-context serving, and training jobs that are fighting recomputation or excessive sharding. It is a weak buy for compute-bound training of models that already fit, where an H100 delivers the same matrix throughput. Newer Blackwell parts raise both compute and memory and add lower-precision formats, but they come with new software stacks and different power and cooling envelopes; see the B200 article for that comparison. Price and availability change faster than any article, so compare quotes in cost per million generated tokens at your real concurrency, not in cost per GPU-hour.

To keep learning, the H100 deep dive covers the shared Hopper architecture, HBM explained covers the memory technology itself, KV cache sizing generalises the arithmetic above, tensor parallelism explains the sharding the H200 lets you remove, and the B200 article covers the next generation.

What to do next

  1. Profile one production H100 workload and split time into prefill, decode, communication and recomputation; only the memory-bound parts benefit.
  2. Compute weights, runtime reserve and KV bytes per token for your model from measured free memory, and derive the concurrent-token budget for 80 GB and 141 GB.
  3. If the model is sharded on H100s purely for memory, plan an H200 configuration with tensor parallelism of one and measure both latency and throughput.
  4. Raise the serving engine's memory fraction, set an explicit maximum context, and try an FP8 KV cache with a quality check on your own evaluation set.
  5. Benchmark at your real concurrency and prompt-to-output ratio, and report cost per million output tokens rather than per GPU-hour.
  6. Confirm rack power and cooling for 700 W SXM parts, or choose the 600 W NVL card for air-cooled PCIe servers.
Key takeaway: The H200 is an H100 with 141 GB of HBM3e at 4.8 TB/s instead of 80 GB at 3.35 TB/s; compute is unchanged. LLM decode is memory-bound, so the bandwidth raises the per-step ceiling by up to 43 percent while the capacity lets a 70B FP8 model run on one GPU with roughly thirteen times more KV cache, which multiplies concurrency. Reconfigure serving to use the memory, benchmark at real concurrency, and expect small gains for compute-bound training.