NVIDIA announced the Blackwell architecture at GTC on 18 March 2024 with two datacenter GPUs. The B200 was the full-power part. The B100 was positioned as the same generation squeezed into the 700 W per-GPU envelope of the H100, so that an eight-GPU HGX B100 board could go into servers, racks and cooling designed for HGX H100. Its launch numbers sat a little below the B200's, and its memory capacity, memory bandwidth and NVLink did not.

That makes the B100 a good way to study a question every GPU buyer eventually faces: what do you lose when a GPU runs on less power, and what do you not lose? The answer depends on whether your work is limited by arithmetic or by memory traffic, and it changes how you should count capacity: per rack kilowatt, not per GPU. This page works through that reasoning with the B100's launch figures, then covers the software side of moving Hopper code to Blackwell and how to see power limiting in practice. The full-power part, its tensor formats and its memory hierarchy are covered in the B200 deep dive.

Advertisement

What was announced, and what happened next

NVIDIA's launch table for the eight-GPU HGX B100 listed 112 petaFLOPS of FP4, 56 of FP8, 28 of FP16 and 14 of TF32, all tensor figures quoted with structured sparsity, plus 240 teraFLOPS of FP64. The HGX B200 table listed 144, 72, 36 and 18 petaFLOPS and 320 teraFLOPS. Divided by eight GPUs, a B100 was quoted at 7 petaFLOPS of sparse FP8, which is 3.5 dense, against 9 sparse (4.5 dense) for the B200: about 78 percent of the B200's arithmetic at 70 percent of its power. Both were launched with up to 192 GB of HBM3e per GPU at 8 TB/s, and fifth-generation NVLink at 1.8 TB/s per GPU.

Then plans changed. In August 2024, analyst reports described the B100 as deprioritized or effectively cancelled, with NVIDIA's packaging capacity going to the B200, the GB200 NVL72 rack and a cut-down B200A aimed at enterprise servers. No NVIDIA statement confirming or denying a cancellation was found, so treat the B100 as something to confirm with your server vendor before you plan on buying it. Shipping HGX B200 boards list 180 GB per GPU rather than the 192 GB of the launch table, a reminder that launch numbers are a guide and the datasheet of the SKU you order is the contract.

How the B100 reaches its lower power at the silicon level has not been published, so this page reasons only from the published power and throughput figures.

Per GPU, launch figuresH100 SXMB100B200
Power envelope700 W700 W1,000 W
Dense FP8 tensorabout 1.98 PF3.5 PF4.5 PF
HBM capacity80 GB HBM3192 GB HBM3e192 GB HBM3e (180 GB on shipping boards)
HBM bandwidth3.35 TB/s8 TB/s8 TB/s
NVLink per GPU900 GB/s1.8 TB/s1.8 TB/s

Why less power means fewer FLOPS but not less bandwidth

The dynamic power of a chip is roughly proportional to the switched capacitance times the square of the supply voltage times the clock frequency. A higher clock needs a higher voltage to switch reliably, so power rises much faster than linearly with clock speed. Run the same kind of logic at a lower power target and the clock and voltage fall together; you give up some speed and save a disproportionate amount of power. That is why a 30 percent cut in power can cost only about 22 percent of peak arithmetic, as in the launch figures.

Memory bandwidth behaves differently. HBM bandwidth is set by the number of stacks, the width of their interfaces and the memory's own data rate, none of which depend on how fast the tensor cores are clocked. The launch figures show exactly this: both parts share 8 TB/s.

The roofline model makes that precise. A kernel's arithmetic intensity is the number of floating-point operations it performs per byte moved from memory. Below the ridge point, the peak FLOPS divided by the bandwidth, the kernel waits on memory; above it, on arithmetic. Large matrix multiplies in training and in the prefill phase of inference sit far above the ridge. Token-by-token decode at small batch sizes reads every weight once per token and does about two operations per weight, so it sits far below. A power cap moves the ceiling for the first group and leaves the second untouched. NVLink is also the same at launch, so a step limited by tensor-parallel all-reduces or expert all-to-alls will not notice the clock either. Mixture-of-experts serving sits in between: arithmetic intensity is low unless batching groups many tokens per expert.

One rack, one power budget: where the watts goRack GPU budget22.4 kW for accelerators4 x HGX B10032 GPUs at 700 W2 x HGX B20016 GPUs at 1000 W6.4 kW left that fits no whole boardPower to one GPUP is roughly C x V^2 x fLower V and ffewer FLOPS per secondcompute-boundbandwidth-boundGEMM-heavy: prefill, trainingtracks the clock: about 14/18 of B200 at launch peaksWeight streaming: decodetracks HBM: 8 TB/s on both at the launch specPlan in throughput per rack kilowatt, not per GPUthe answer depends on which of the two boxes above dominates your hours
A fixed rack budget fits more 700 W GPUs than 1,000 W ones. The power cut lowers the clock, which hurts GEMM-heavy work, while decode is limited by HBM bandwidth that the cap does not touch.
Advertisement

Planning in throughput per rack kilowatt

Datacenters usually run out of power and cooling before floor space, and power is provisioned per rack, so the question is how much work a rack produces within its budget. The small model below compares the three parts within a fixed GPU power budget of 22.4 kW, which is exactly four HGX boards at 8 x 700 W. It counts only GPU power, rounds down to whole eight-GPU boards, and assumes 40 percent model FLOPS utilization for training and 70 percent of peak bandwidth for decode.

# Launch-spec peaks per GPU (NVIDIA, March 2024), tensor figures with sparsity.
PARTS = {
    #          watts  FP8 sparse PF  HBM TB/s
    "H100": (700, 3.958, 3.35),
    "B100": (700, 7.0, 8.0),
    "B200": (1000, 9.0, 8.0),
}

RACK_KW_FOR_GPUS = 22.4   # GPU power budget per rack, e.g. four 8 x 700 W boards

def plan(name, mfu=0.40, decode_bw_eff=0.70):
    watts, fp8_sparse, tbps = PARTS[name]
    dense_pf = fp8_sparse / 2
    gpus = int(RACK_KW_FOR_GPUS * 1000 // watts)
    gpus -= gpus % 8                       # whole 8-GPU boards only
    train = gpus * dense_pf * mfu          # sustained dense FP8 PFLOPS
    ridge = dense_pf * 1e15 / (tbps * 1e12)  # FLOP per byte at the roofline knee
    # Decode of a 70B FP8 model: every token streams ~70 GB of weights.
    tok_s = tbps * 1e12 * decode_bw_eff / 70e9
    return gpus, train, ridge, tok_s

print(f"{'part':5} {'GPUs/rack':>9} {'train PF':>9} {'ridge F/B':>9} {'decode tok/s/GPU':>17}")
for name in PARTS:
    g, t, r, d = plan(name)
    print(f"{name:5} {g:9d} {t:9.1f} {r:9.0f} {d:17.0f}")

Running it prints:

part  GPUs/rack  train PF ridge F/B  decode tok/s/GPU
H100         32      25.3       591                34
B100         32      44.8       438                80
B200         16      28.8       562                80

Three things stand out. First, within the same envelope, the B100 rack delivers about 1.8 times the H100 rack's training throughput and more than twice its single-stream decode speed, without changing the power or cooling. Second, the B200 rack is worse than the B100 rack in this particular budget, because 22.4 kW holds only two whole B200 boards and strands 6.4 kW. Third, the ridge point falls from about 591 FLOPs per byte on H100 to about 438 on B100, so slightly more kernels become compute-bound.

The numbers rank options; they do not predict a benchmark. Replace them with your measured utilization and real per-board power, including CPUs, NICs and fans.

Porting Hopper software to compute capability 10.0

The B100, B200 and GB200 report compute capability 10.0, target sm_100. CUDA 12.8 was the first toolkit able to generate native cubin for it, and PyTorch 2.7 was the first release with prebuilt CUDA 12.8 wheels that include it. NVIDIA's Blackwell compatibility guide makes one rule central: an application runs on Blackwell only if it contains either native sm_100 cubin or PTX that the driver can compile just in time. Kernels shipped only as Hopper cubin will not run.

The trap for performance code is architecture-specific features. Hopper's fastest kernels use features compiled for sm_90a, such as warpgroup matrix-multiply instructions, and the guide states that PTX compiled for compute_90a is not supported on Blackwell. Anything built that way, including attention and GEMM kernels written specifically for Hopper, needs a Blackwell build or a Blackwell-specific replacement. Generic PTX kernels will JIT and run, often far below peak.

# Build native Blackwell cubin plus PTX (CUDA 12.8 or later), as NVIDIA's guide recommends.
nvcc -O3 my_kernels.cu -o my_kernels \
  -gencode=arch=compute_90,code=sm_90 \
  -gencode=arch=compute_100,code=sm_100 \
  -gencode=arch=compute_100,code=compute_100

# Prove the binary does not depend on a cubin you forgot: ignore all cubins, JIT the PTX.
CUDA_FORCE_PTX_JIT=1 ./my_kernels --self-test

# For PyTorch extensions, list the targets the installed wheel was built for.
python -c "import torch; print(torch.version.cuda, torch.cuda.get_arch_list())"

Run the CUDA_FORCE_PTX_JIT=1 test in CI on a Blackwell machine. It ignores every cubin, so a pass proves the binary carries PTX it can fall back on. Then check, kernel by kernel, that the fast paths you rely on actually have Blackwell implementations in the libraries you pin, because a silent fallback to a generic kernel looks exactly like a slow GPU.

Seeing power limiting in practice

At its power limit, a GPU lowers clocks to stay under it. There is no error, only an SM clock that drops in GEMM-heavy phases, so record power and clocks alongside step times.

# What range of power limits does this board accept? Read it; do not assume 700 W is allowed.
nvidia-smi -q -d POWER

# Sample once a second while a training step or benchmark runs.
nvidia-smi --query-gpu=index,power.draw,enforced.power.limit,clocks.sm,temperature.gpu,utilization.gpu \
           --format=csv -l 1 > power_trace.csv

# Field names differ between driver releases; list the ones your driver supports.
nvidia-smi --help-query-gpu | grep -i -E "power|clocks"

If you want to approximate a 700 W Blackwell on a higher-power board, you can lower the power limit with nvidia-smi -pl as root, but only within the range the first command reports, and the result will not exactly reproduce a part designed for that envelope. In a trace, power limiting looks like power flat at the enforced limit, a falling SM clock and step times that track it; bandwidth-bound decode usually draws less and runs at full clock.

Worked example: four HGX H100 racks, retrofit or rebuild

A team runs four racks, each with four HGX H100 boards in a 22.4 kW GPU budget. Their work is half fine-tuning of a 70B dense model and half serving that model to interactive users. They compare replacing the boards with HGX B100 in the same racks against rebuilding for HGX B200 at higher power per rack.

With the rack budget unchanged, the calculator gives 44.8 sustained FP8 petaFLOPS per rack for B100 against 25.3 for H100, and 80 against 34 tokens per second per GPU for single-stream decode, with 32 GPUs per rack either way. The 70B FP8 weights, about 70 GB, fit in one B100's 192 GB with room for KV cache, where on an 80 GB H100 they needed two GPUs per replica; so serving capacity rises by more than the per-GPU speedup. A B200 retrofit in the same budget fits only 16 GPUs per rack and loses on both counts.

A rebuild to about 32 kW of GPU power per rack would fit four HGX B200 boards: 57.6 sustained petaFLOPS at the same assumptions, about 29 percent more training throughput than the B100 rack, and identical decode speed per GPU. The decision turns on the cost of new power and cooling against that gain, and on what they can actually buy; given the availability reports, they ask their vendor first.

Failure modes

  • Planning on launch numbers: the shipping SKU differs, as the 180 GB HGX B200 boards show; size from the datasheet of the part you order.
  • Counting per GPU, not per kilowatt: the faster GPU loses when the rack budget strands power it cannot use.
  • Sparse peaks treated as dense: launch FLOPS were quoted with sparsity; halve them unless your model is actually 2:4 sparse.
  • Hopper-only kernels: sm_90a builds do not run on Blackwell, so a container that worked on H100 fails or silently falls back.
  • Ignoring the rest of the server: CPUs, NICs, fans and power-supply losses add to the board's GPU power, and a retrofit can exceed the rack budget that the GPUs alone fitted.
  • Reading clocks without power: a slow step with a low SM clock is power limiting, not a bad GPU; check power draw against the enforced limit before filing a hardware ticket.

Choosing between the options

The 700 W part fits when the power and cooling are fixed, the work is mostly serving, or the faster part would strand rack capacity. The 1,000 W part fits when power can be added and the work is mostly training. A rack-scale NVL72 system, covered in the GB200 deep dive, fits when models need a 72-GPU NVLink domain and you can build for liquid cooling. Hopper parts, covered in the H100 article, remain reasonable where software is not ready for compute capability 10.0. Why bandwidth is the lever for decode, and why it does not follow the clock, is covered in the HBM architecture article.

What to do next

  1. Measure what share of your fleet's GPU hours is compute-bound (training, prefill) and what share is bandwidth-bound (decode).
  2. Write down your real power budget per rack, including CPUs, NICs and fans, and compute throughput per rack kilowatt for each candidate part.
  3. Ask your vendor which Blackwell HGX SKUs are orderable, and size from that SKU's datasheet, not the launch table.
  4. Rebuild your kernels and extensions with CUDA 12.8 or later for sm_100 plus PTX, and add a CUDA_FORCE_PTX_JIT=1 test to CI.
  5. Audit every sm_90a dependency, especially attention and GEMM libraries, for a Blackwell build.
  6. Record power draw, enforced limit and SM clock with your step times, so power limiting is visible before anyone suspects the hardware.
Key takeaway: The B100 was announced as Blackwell in Hopper's 700 W envelope: about 78 percent of the B200's quoted arithmetic, with the same launch memory bandwidth and NVLink. Power caps lower clocks, so compute-bound training and prefill lose roughly the arithmetic gap, while bandwidth-bound decode loses almost nothing. Count capacity per rack kilowatt, not per GPU. Confirm availability with your vendor, since analysts reported the part deprioritized, and rebuild Hopper code for compute capability 10.0, because sm_90a kernels do not run on Blackwell.