NVIDIA's H100 was the first GPU with a confidential computing mode. When the mode is on, and the GPU is passed to a confidential virtual machine, the host operating system, the hypervisor and anyone watching the PCIe bus should see only ciphertext, while CUDA code inside the VM runs unchanged. It is the hardware piece that lets you run a model on rented machines without handing its weights or your users' prompts to the machine owner.
This page is about the GPU side and day-to-day operation: what flipping the mode changes inside the chip, how the mode is set and checked, what a CUDA training or inference job feels, how to measure the cost on your own hardware, and the mistakes that leave a deployment claiming protection it does not have. The CPU half, measured boot and composite attestation are covered in Confidential Compute for LLM; the chip itself in the H100 architecture page.
Product details here come from NVIDIA's documentation and engineering blog, its open-source admin tools, and cloud provider guides. Driver releases change behaviour, so treat command output shown below as an example to check against your own system.
Three modes and what changes in the GPU
NVIDIA describes three modes. The mode is a property of the GPU, set from the host, and takes effect after a GPU reset.
| Mode | What is enabled | Use it for |
|---|---|---|
| CC-Off | Standard operation, no confidential features | Ordinary workloads |
| CC-On | All confidential features: firewalls active, bus traffic encrypted, performance counters disabled | Production confidential workloads |
| CC-DevTools | Security features on, but the blocks that stop profiling and debugging are lifted | Developing and profiling, never production |
Underneath the modes sit a few hardware facts worth knowing precisely. Each H100 has an identity key fused in at manufacture, with a certificate chain to NVIDIA, and an on-die root of trust that measures the GPU firmware at boot. In CC-On the GPU blocks direct access to its memory from outside the trust boundary. The HBM itself is not encrypted: NVIDIA's position is that on-package memory is protected against everyday physical attack tools by being on the package, so the defence is access control rather than cryptography. Performance counters are disabled because they can leak information through side channels, which is also why profilers stop working.
Command buffers and CUDA kernels are encrypted and signed before crossing PCIe, along with data. The keys come from an SPDM session that the NVIDIA driver inside the confidential VM negotiates with the GPU when it initialises the device.
The trust boundary
The boundary only makes sense with a CPU-side trusted execution environment. Hopper cannot read the confidential VM's encrypted private memory, so data moves through bounce buffers: the driver encrypts into shared memory with AES-GCM, the GPU copies the ciphertext across PCIe and decrypts it on chip, and results return the same way. Put a CC-On GPU behind an ordinary VM and the driver, the keys and the plaintext all live in memory the host can read. The CPU side is AMD SEV-SNP or Intel TDX; see Intel TDX in depth for one of them.
Notice which party sets the mode: the host administrator, using a privileged tool. That is fine, because the guest never has to trust the host's claim. The GPU's attestation report covers its confidential-mode state, and the guest checks that report before using the GPU.
Setting the mode on the host
On the host, NVIDIA's open-source gpu-admin-tools repository provides nvidia_gpu_tools.py, run as a privileged command. It reaches the GPU through vfio-pci and reports an error if another driver still holds the device, so stop workloads first. Selecting the GPU by PCI address and resetting after the switch looks like this:
# Host side, as root, with workloads on this GPU stopped.
sudo python3 nvidia_gpu_tools.py --gpu-bdf=43:00.0 --set-cc-mode=on --reset-after-cc-mode-switch
# Leave confidential mode, for example to return the GPU to a general pool.
sudo python3 nvidia_gpu_tools.py --gpu-bdf=43:00.0 --set-cc-mode=off --reset-after-cc-mode-switchOn a public cloud you do not run this; the provider exposes confidential GPU instance types and has already set the mode. On your own servers, put the switch into provisioning automation so a GPU's mode always matches the pool it is in, and never let a CC-DevTools GPU drift into a production pool. A host that forgets to switch modes is caught later by attestation, but only if the guest actually checks.
Inside the VM: verify, then mark ready
Inside the confidential VM, the order of events is: the driver loads, opens the SPDM session, and the GPU comes up in a state where it will not accept work. Software then verifies the GPU's attestation and only afterwards sets the ready state. nvidia-smi exposes the relevant pieces under its conf-compute subcommand. The documented ones used in deployment guides are -f, which prints the CC status, -mgm for the multi-GPU mode on eight-GPU systems, and -srs 1 to set the GPU ready.
A common pattern is a systemd drop-in that marks GPUs ready right after nvidia-persistenced starts. That is convenient and also skips the point of the gate: the GPU becomes usable whether or not anyone verified it. Put a verification step first, and fail closed:
#!/usr/bin/env python3
"""Run once at boot inside the CVM, before any workload starts."""
import subprocess, sys
# Placeholder: your pinned nvtrust verifier invocation (see below).
VERIFIER_CMD = ["/opt/gpu-gate/verify-gpu-evidence"]
def sh(*args):
return subprocess.run(args, capture_output=True, text=True, check=True).stdout
def fail(msg):
print("GPU gate FAILED: " + msg, file=sys.stderr)
sys.exit(1) # unit fails; workloads ordered After= it never start
status = sh("nvidia-smi", "conf-compute", "-f")
multi = subprocess.run(["nvidia-smi", "conf-compute", "-mgm"],
capture_output=True, text=True).stdout
protected_pcie = "Protected PCIe" in multi
if not protected_pcie and "CC status: ON" not in status:
fail("single-GPU system not in CC mode: " + status.strip())
# Verify attestation evidence with NVIDIA's verifier (nvtrust). The module name and
# flags differ by release; pin the version and copy the invocation from its docs.
verify = subprocess.run(VERIFIER_CMD, capture_output=True, text=True)
if verify.returncode != 0:
fail("attestation verification failed")
sh("nvidia-smi", "conf-compute", "-srs", "1") # accept work only now
print("GPU gate passed")VERIFIER_CMD is a placeholder for you to replace: one cloud guide runs NVIDIA's local verifier as python3 -m verifier.cc_admin --user_mode for a single GPU and a separate PPCIe verifier for eight-GPU systems, but module paths have moved between nvtrust releases. Whatever you run must check the certificate chain, revocation, the firmware measurements against NVIDIA's reference values, and a fresh nonce. Local verification also protects only this boot; anyone else relying on the machine needs the evidence too, which is the subject of what the end user can verify.
What CUDA workloads feel
From the application's point of view almost nothing changes. You do not recompile kernels or change PyTorch code; CUDA support for confidential computing arrived with the CUDA 12.2 update in July 2023. What changes is the cost model. NVIDIA reports that raw GPU compute and GPU memory bandwidth are at parity with CC-Off, while CPU-to-GPU bandwidth is limited by CPU encryption speed, which it measured at roughly 4 GB/s in its 2023 engineering post. Your number depends on CPU model, vCPU count and driver, so measure it, but plan for host-to-device copies being roughly an order of magnitude slower than the PCIe link itself.
That splits workloads cleanly. Anything that keeps data on the GPU is unaffected: a training step whose parameters, gradients and optimiser state all live in HBM pays nothing per step beyond moving the batch. Anything that streams through the host pays on every byte: loading weights, CPU offload of optimiser state or KV cache, copying logits back every token, frequent checkpoints, and data loaders that send uncompressed images. Encryption also burns vCPU time, so an undersized VM throttles the GPU.
Profiling is the other change. With counters disabled in CC-On, tools such as Nsight cannot collect hardware metrics. Profile in CC-DevTools or CC-Off on identical hardware, then confirm end-to-end timings in CC-On. Independent benchmarking has also reported that some timing facilities and GPUDirect RDMA paths are unavailable in CC mode; check the release notes for your driver before designing around them.
Measuring the cost yourself
Measure host-to-device bandwidth inside the confidential VM with wall-clock timing around explicit synchronisation, which works in every mode, then run the same script with CC off on the same machine type:
import time, torch
def h2d_gbps(mb=1024, reps=10, pinned=True):
x = torch.empty(mb * 2**20, dtype=torch.uint8, pin_memory=pinned)
y = torch.empty_like(x, device="cuda")
y.copy_(x); torch.cuda.synchronize() # warm-up
t0 = time.perf_counter()
for _ in range(reps):
y.copy_(x, non_blocking=True)
torch.cuda.synchronize()
return reps * mb / 1024 / (time.perf_counter() - t0)
for pinned in (True, False):
print(f"pinned={pinned}: {h2d_gbps(pinned=pinned):.1f} GB/s")Record the result with the CPU model, vCPU count and driver version. Repeat for device-to-host. The ratio between the two modes is the number every capacity plan below should use.
Worked example: a fine-tuning job
Suppose a team fine-tunes an 8-billion-parameter model on one H100 inside a confidential VM. Weights in bf16 are about 16 GB. At the roughly 4 GB/s figure, the initial load takes about 4 seconds of copying, which is irrelevant for a job that runs for hours. With LoRA, the frozen weights stay in HBM, adapters are small, and each step moves only a batch of token IDs, a few megabytes. Overhead: negligible.
Now the same team tries full fine-tuning with optimiser state offloaded to CPU memory, because parameters plus Adam state at 16 bytes per parameter, about 128 GB, do not fit in 80 GB of HBM. Each step now ships about 16 GB of bf16 gradients down and 16 GB of updated parameters up. At 4 GB/s that is about 8 seconds of transfer per step; at the PCIe Gen5 x16 link's theoretical 64 GB/s per direction, one direction after the other, it would be about half a second. The design that was fine without confidential computing is now transfer-bound. The fix is architectural: LoRA, a smaller model, or sharding the state across GPUs in a protected multi-GPU configuration so it stays on-device.
Inference follows the same rule. Large models that compute for most of a request see little overhead; small models with long prompts feel the per-request transfer, and KV-cache offload to host memory is the most expensive habit to carry over.
Eight GPUs: protected PCIe
Single-GPU passthrough gives each GPU its own trust domain. For eight-GPU HGX boards NVIDIA added a protected PCIe mode, often written PPCIe, that passes the whole board to one confidential VM. On such a system nvidia-smi conf-compute -mgm reports the multi-GPU mode as Protected PCIe, and a cloud guide notes that -f normally reports CC status OFF there, which is why the gate script checks -mgm first. On Hopper, GPU-to-GPU traffic over NVLink is not encrypted, so the security argument depends on attesting the GPUs and NVSwitches together as one physical unit. Confidential Compute for LLM covers the consequences for tensor parallelism.
Failure modes
- Ready without verification. Setting the ready state in a boot script that never checks evidence makes the GPU usable on a host that lied about its mode.
- DevTools in production. CC-DevTools keeps the security features on but lifts the profiling blocks. Make sure your attestation policy tells the two apart and requires CC-On.
- A CC-On GPU in an ordinary VM. The GPU is protected, the VM holding its keys is not.
- Trusting the gate inside the box. A local check stops honest mistakes; a remote key broker or relying party must re-verify before releasing weights.
- Offload-heavy designs. KV-cache or optimiser offload that was cheap becomes the bottleneck.
- Profiling numbers taken in the wrong mode. Counters from DevTools runs are fine for kernels but not for end-to-end throughput, which only CC-On measurements show.
- Stale firmware. Verifiers compare measurements against reference values and revocation lists; unpatched GPU firmware should fail, so patch on a schedule.
Trade-offs
| Option | Protects against | Costs |
|---|---|---|
| CC-Off on a trusted host | Nothing beyond your own controls | None |
| Single H100 in CC-On with a CPU TEE | Host OS, hypervisor, bus snooping | Slower host transfers, no profiler counters |
| Eight-GPU protected PCIe | Same, for models that need one board | NVLink unencrypted on Hopper, board-level attestation |
| Avoid sensitive data on rented GPUs | Everything above, by not taking the risk | Capacity, or a smaller model you host |
What to do next
- Decide who you are protecting against; if it is not the host operator, CC mode may not be needed.
- Confirm your GPUs, CPU TEE, driver and CUDA versions are on NVIDIA's supported list for CC.
- Automate host mode switching in provisioning and keep separate CC-On and CC-DevTools pools.
- Replace any ready-state shortcut with a gate that verifies attestation first, and test that it fails closed.
- Measure host-to-device bandwidth in both modes and redo the transfer budget for your workload.
- Remove KV-cache and optimiser offload from confidential deployments where you can.
- Make your key broker require CC-On, current firmware and a fresh nonce before releasing weights.