AMD Secure Encrypted Virtualization (SEV) lets a virtual machine run with its memory encrypted under a key that the hypervisor never sees. Its current form, SEV-SNP (Secure Nested Paging), adds integrity protection and a signed attestation report, so a remote party can check exactly what booted inside the VM before trusting it with secrets. For AI workloads that means you can run a model server on a cloud host and release model weights or user prompts only to a guest whose measured software you approved, even though the cloud operator controls the machine.

This article explains how the hardware works in terms a platform engineer uses: what the three generations added, how the C-bit and the Reverse Map Table split memory into private and shared, why the guest kernel must change, what the attestation report contains and how to verify it, and what SEV costs a model server. Intel's counterpart is covered in Intel TDX, and the full bring-up of a confidential inference node is in confidential compute for LLMs.

What SEV protects, and from whom

SEV changes the threat model of virtualisation. Normally the hypervisor is fully trusted: it can read and write any guest page and register. Under SEV-SNP the trusted computing base shrinks to the CPU package, including the AMD Secure Processor (ASP, a dedicated security core that manages keys), its firmware, and the software inside the guest. The hypervisor, host kernel and other tenants are outside it. Someone reading DRAM directly sees only ciphertext, but active physical attacks on the memory bus are not covered: the RMP enforces integrity by access checks, not cryptographically.

What SEV does not promise matters just as much. The hypervisor still decides when the guest runs, so denial of service is out of scope. It still sees which pages are touched at page granularity, the timing of exits and the size and timing of I/O. And the guest's own software, including a vulnerable model server, is inside the boundary: SEV does nothing about an injected prompt that convinces your agent to send data out through its normal network path. Write these limits into your threat model before you rely on SEV for anything.

Three generations: SEV, SEV-ES, SEV-SNP

SEV arrived in three steps, each closing a hole the previous one left open.

GenerationFirst EPYC familyAddsGap it left
SEV7001 (Naples)Per-VM memory encryption keyed by address space IDRegisters exposed on exit; no integrity
SEV-ES7002 (Rome)Encrypted register state (VMSA) on every exitHypervisor could remap or replay pages
SEV-SNP7003 (Milan)Reverse Map Table integrity, VMPLs, signed reportsSide channels; see failure modes

The plain SEV and SEV-ES modes are of historical interest for new designs: published attacks such as SEVered showed a hypervisor could remap encrypted pages and use the guest itself as a decryption oracle. Treat SNP as the minimum for protecting model weights or user data.

Memory encryption, the C-bit and the RMP

SEV-SNP guest: who holds which key, and where the attestation report travelsHost (untrusted): hypervisor, host OS, cloud operatorHypervisorschedules, maps pagesShared pagesGHCB, bounce buffersGPU (CC mode) or NICDMA only to shared memorySNP guest (trusted): firmware, kernel, model serverPrivate pages (C-bit set)encrypted with per-guest keyModel serverweights, promptsGuest kernelrequests reportcopyAMD Secure Processorkeys, RMP updates, signs reportsMemory controller AESencrypts on write, decrypts on readRelying party / KMSVCEK chain via AMD KDS, releases keyreportsigned reportThe hypervisor still controls scheduling and I/O, but it sees ciphertext for private pages and the RMP blocks remapping.
Trust split in an SEV-SNP deployment. Private guest memory is ciphertext to everything outside the guest; I/O passes through explicitly shared pages.

Encryption happens in the memory controller. Each SEV guest is tied to an address space ID, and the ASP loads a key for that ID into the controller. Cache lines are encrypted on the way to DRAM and decrypted on the way back, so software inside the guest sees plaintext and pays no instruction-level cost. The guest chooses which pages are private through the C-bit, a physical address bit in its own page-table entries whose position is reported by CPUID leaf 0x8000001F. Pages with the C-bit set are encrypted with the guest key; pages without it are shared and readable by the host.

Integrity comes from the Reverse Map Table (RMP), a system-wide table with one entry per 4 KB host page recording which guest owns it, at which guest physical address, and whether the guest has validated it. Every write by the hypervisor and every guest access is checked against the RMP in hardware. The guest must accept each private page with the PVALIDATE instruction before use; if the hypervisor later swaps or remaps that page, the guest gets an exception instead of silently reading attacker-chosen data. This is why an SNP guest needs an SNP-aware kernel and firmware: memory acceptance, the #VC exception handler for intercepted instructions, and the Guest-Hypervisor Communication Block (GHCB), a shared page through which the guest deliberately hands the hypervisor only the register values an exit needs.

SNP also introduces Virtual Machine Privilege Levels, VMPL0 to VMPL3. A small trusted component such as a virtual TPM or secure services module can run at VMPL0 while the ordinary guest kernel runs at a lower level, which lets one VM hold a secret its own kernel cannot read.

Measurement and attestation reports

At launch, the hypervisor asks the ASP to create the guest and passes in the initial memory, usually the OVMF firmware plus, for measured direct boot, a table of kernel, initrd and command-line hashes. The ASP hashes these pages into a 48-byte SHA-384 launch MEASUREMENT. Later the guest can ask the ASP for a report containing that measurement, 64 bytes of REPORT_DATA chosen by the guest, HOST_DATA supplied at launch, the guest POLICY (debug allowed, SMT allowed, migration rules), the VMPL of the requester, the CHIP_ID and the TCB versions of the bootloader, ASP firmware, SNP firmware and microcode. The ASP signs it with ECDSA P-384 using a Versioned Chip Endorsement Key (VCEK) derived from a chip secret and the current TCB version.

On Linux the guest has two interfaces: the older /dev/sev-guest device with an ioctl, and since kernel 6.7 the vendor-neutral configfs-tsm tree, which also serves TDX.

# Guest side: request a report bound to a fresh nonce (Linux 6.7+ configfs-tsm)
R=/sys/kernel/config/tsm/report/req0
mkdir "$R"
head -c 64 /dev/urandom > /tmp/nonce          # in practice: hash(nonce || TLS key)
cat /tmp/nonce > "$R/inblob"                   # becomes REPORT_DATA
cat "$R/provider"                              # sev_guest on SNP
cat "$R/outblob" > /tmp/report.bin             # signed attestation report
cat "$R/auxblob" > /tmp/certs.bin              # optional cert table from host
rmdir "$R"

Verification happens on the relying party, typically a key management service. The certificate chain runs from the AMD Root Key (ARK) through the AMD SEV Key (ASK) to the VCEK, all served by the AMD Key Distribution Service at kdsintf.amd.com, with paths of the form vcek/v1/Milan/cert_chain for the intermediate and root. On some clouds the ASP instead signs with a Versioned Loaded Endorsement Key (VLEK), which AMD issues to the cloud provider and the provider loads into the ASP; it has its own chain under vlek/v1. The logic, written as pseudocode so it does not depend on one library's API:

def verify_snp(report, nonce, policy):
    chain = kds_cert_chain(report.product)        # ARK, ASK from AMD KDS (cache it)
    vcek  = kds_vcek(report.chip_id, report.reported_tcb)
    check(chain.ark.self_signed() and chain.ark.fingerprint in AMD_ROOT_PINS)
    check(chain.ask.signed_by(chain.ark) and vcek.signed_by(chain.ask))
    check(vcek.verify_ecdsa_p384(report.body, report.signature))
    check(report.report_data == expected_binding(nonce, report.tls_pubkey))
    check(report.measurement in policy.allowed_measurements)
    check(not report.policy.debug_allowed)
    check(report.reported_tcb >= policy.minimum_tcb)    # per component
    check(report.vmpl == policy.expected_vmpl)
    return True

The measurement only covers what the ASP hashed at launch. If your firmware then boots an unmeasured disk image, the report says nothing about the model server. Either use measured direct boot so the kernel, initrd and command line are in the launch digest, with an integrity-protected root filesystem whose hash is on the command line, or extend runtime measurements through a vTPM at VMPL0. The virtee project's sev-snp-measure tool computes expected measurements from the same inputs, so your build pipeline can publish the allow-list alongside the image.

Worked example: releasing model weights to a guest

Consider a team serving a proprietary model on a cloud SNP instance. Weights are stored encrypted in object storage under a data key wrapped by their own KMS. At boot, an agent in the guest generates a TLS key pair, requests a report with REPORT_DATA set to SHA-512 of the KMS nonce concatenated with the public key, and sends report plus key to the KMS. The KMS runs the verification above, checks the measurement against the allow-list published by the image build, and wraps the data key to the attested public key. The guest decrypts weights into private memory only.

Binding the TLS key into REPORT_DATA is what stops relay attacks: a valid report from some other guest cannot be paired with an attacker's key. Requiring debug disabled stops the host from launching an approved image in a mode where it can read memory. Requiring a minimum TCB means that after AMD ships a firmware fix, hosts that have not applied it cannot obtain keys. For end-user-verifiable serving, where clients rather than your KMS check the node, see confidential computing for LLM inference.

What SEV costs a model server

Compute inside the guest runs at near-native speed because encryption is in the memory path. The costs come from crossings. Every intercepted instruction becomes a #VC exception plus a GHCB round trip. Devices cannot DMA into private memory, so the guest kernel bounces I/O through shared buffers (SWIOTLB) and copies each block, which costs CPU and bandwidth on storage-heavy and network-heavy paths such as loading tens of gigabytes of weights or streaming tokens. Page acceptance adds boot latency proportional to memory size unless the kernel accepts memory lazily.

GPUs are the big question for AI. A GPU outside the guest would see plaintext over PCIe, so confidential GPU computing, as on NVIDIA H100 in CC mode, establishes an encrypted session between the guest driver and the GPU and passes data through encrypted bounce buffers in shared memory, with separate GPU attestation. Transfer-heavy workloads feel this most; long compute-bound kernels much less. Measure your own model with CC on and off before committing to a capacity plan.

Failure modes

Research has repeatedly found weaknesses in SNP's edges. CipherLeaks showed that deterministic encryption lets a hypervisor observe when a ciphertext block repeats and infer secrets from register save areas. CacheWarp (CVE-2023-20592) abused cache invalidation to roll back guest writes and was fixed by microcode. BadRAM used a modified DIMM SPD chip to create aliased addresses and bypass RMP protections, mitigated by firmware that checks memory configuration. Interrupt-injection attacks such as Heckler target how guest kernels handle hypervisor-delivered interrupts. The pattern: fixes arrive as firmware and microcode, which raise the TCB version, which only helps if your verifier enforces a minimum TCB.

Operational failures are more common than exotic ones. Verifiers that check the signature but not REPORT_DATA accept replayed reports. Allow-lists that hold a measurement of firmware only approve any kernel. Debug policy left permitted in a staging template reaches production. VCEK fetches from KDS are rate limited and can fail at scale, so cache certificates per chip and TCB. And image updates break key release when the build pipeline forgets to publish the new expected measurement.

Trade-offs

SEV-SNP protects whole VMs, so existing model servers run unmodified, at the price of a large trusted guest: every library in the image is inside the boundary. Process-level enclaves have a smaller trusted base but force application rewrites. Compared with Intel TDX, SNP offers a similar VM model with VMPLs for in-guest privilege separation and a chip-specific VCEK chain instead of a quoting enclave; pick by which CPUs your provider offers with confidential GPUs. Against no confidential computing at all, SNP adds operational work, measured builds, a verifier and TCB tracking, in exchange for removing the cloud operator from the list of parties who can read your weights.

What to do next

  1. Write down which parties SEV removes from your trust boundary and which it leaves, including metadata leaks.
  2. Launch an SNP guest on your provider and read a report through configfs-tsm with a random nonce.
  3. Verify it offline: ARK, ASK, VCEK chain, signature, REPORT_DATA, measurement, debug bit, TCB.
  4. Switch to measured direct boot and publish expected measurements from your image pipeline.
  5. Gate model-key release in your KMS on the verified report, bound to a TLS key.
  6. Benchmark weight loading and token throughput with confidential mode on and off.
  7. Subscribe to AMD security bulletins and raise your minimum TCB when fixes ship.
Key takeaway: SEV-SNP encrypts each guest's memory with a key the hypervisor never holds and uses the Reverse Map Table to stop the host remapping or replaying pages. Its value for AI comes from attestation: verify the VCEK chain, the nonce binding, the measurement, the debug bit and a minimum TCB before releasing model keys, measure the kernel and root filesystem, and budget for bounce-buffer I/O and GPU confidential mode costs.