Intel Software Guard Extensions lets an ordinary process carve out an enclave: a region of its address space whose memory the operating system, the hypervisor and other processes cannot read or modify, and whose contents a remote party can verify by attestation before trusting it with a secret. It shaped the vocabulary of confidential computing and has a long, instructive list of attacks.

This article explains how SGX actually works, from the memory model and enclave lifecycle through measurement, sealing and DCAP remote attestation, and what has changed: EPID attestation is gone, client CPUs dropped SGX, server enclaves grew to hundreds of gigabytes, and 2025 research showed physical memory-bus attacks on DDR4 Scalable SGX. In AI systems it mostly fits small trusted services such as key brokers, not GPU inference; for VM-level isolation see Intel TDX.

The threat model

SGX assumes the attacker controls all software on the machine except the enclave: the kernel, the hypervisor, the BIOS and firmware outside the CPU, device drivers and DMA-capable devices. The trusted computing base is the CPU package with its microcode, plus the code you put in the enclave. Everything the enclave receives from outside, including system call results, file contents, time and randomness from the OS, is untrusted input.

That TCB is far smaller than a confidential VM's, and the programming model is harder: the enclave must exit to the untrusted host for all I/O, the host decides whether to run it, and it sees which pages the enclave touches. So availability is never protected, and access patterns leak.

Enclave memory: ELRANGE, EPC and EPCM

Intel SGX: an enclave inside an untrusted process, with memory the OS cannot readUntrusted host processApp codenetworking, I/OEnclaveELRANGE, EPC pagesECALL in, OCALL outOS, hypervisor, BIOS, devicesuntrusted: schedule, page, observeCPU package + microcodeEPCM checks, memory encryption, keysMeasurementMRENCLAVE = hash of build; MRSIGNER = signer keyQuoting EnclaveECDSA quote (DCAP)PCCS cachePCK certs, TCB info, CRLsREPORTRelying party / key brokerverify quote, check TCB status, release secretquotecollateralsecret over attested channelOut of scope or weak: side channels, physical bus interposers on DDR4 Scalable SGX, rollback of sealed stateTrust shrinks to the CPU and the enclave code; everything around them is assumed hostile.
Enclave code runs in the host process but its pages live in the EPC, checked by the CPU on every access. Attestation turns a hardware report into a quote a remote key broker can verify.

An enclave occupies a virtual address range in its host process called ELRANGE. The physical pages behind it come from the Enclave Page Cache, a region of RAM reserved at boot by the BIOS. The CPU keeps a private table, the EPCM, recording for every EPC page which enclave owns it, its type (regular, thread control structure, version array and others), its permissions and the virtual address it was mapped at. On every access the CPU checks the EPCM, so the OS can still map and unmap pages but cannot redirect an enclave to a page that is not its own, or read one from outside.

Data leaving the CPU package is encrypted. Early client and Xeon E3 parts used a Memory Encryption Engine with an integrity tree, giving replay protection but a small EPC, commonly 128 MB. Scalable SGX on Xeon Scalable dropped the integrity tree for multi-key total memory encryption, trading physical replay protection for size: Intel documents 512 GB of EPC per socket on Xeon 6. Beyond the EPC, the kernel pages enclave memory out with EWB and back with ELDU, slowly.

Building and entering an enclave

An enclave is built page by page. ECREATE allocates the SGX Enclave Control Structure. EADD copies each initial page into the EPC and EEXTEND hashes its contents into a running measurement. EINIT checks the signature structure, SIGSTRUCT, supplied by the enclave author, and finalises the measurement; after EINIT the enclave can run and its initial contents can no longer change. EENTER moves a thread in, EEXIT moves it out, and an interrupt causes an asynchronous exit that saves the register state inside the enclave before the OS sees anything. SGX2 added dynamic memory management, with instructions such as EAUG, EMODPR, EMODT and EACCEPT, so an enclave can grow its heap and change page permissions after initialisation.

With the Intel SGX SDK you describe the boundary in an EDL file: ECALLs the host may invoke and OCALLs the enclave may call out to. Generated code copies buffers according to the direction attributes:

enclave {
    trusted {
        /* host -> enclave: wrapped key in, nothing secret out */
        public sgx_status_t unwrap_and_store([in, size=len] const uint8_t* blob, size_t len);
        public sgx_status_t sign_request([in, size=n] const uint8_t* msg, size_t n,
                                         [out, size=64] uint8_t* sig);
    };
    untrusted {
        /* enclave -> host: I/O must leave the enclave */
        void ocall_log([in, string] const char* line);
    };
};

Every [in] and [out] is a security decision. [user_check] pointers let the host aim into enclave memory or change buffers mid-read. Keep the EDL small, copy everything in and validate lengths inside.

Identity: MRENCLAVE, MRSIGNER and the debug flag

SGX gives an enclave two identities. MRENCLAVE is the measurement, a SHA-256 hash of the exact sequence of pages and their layout as built; change one byte of code or initial data and it changes. MRSIGNER is the hash of the public key that signed SIGSTRUCT, identifying the author. The signature structure also carries a product ID and a security version number, ISVSVN, and the enclave attributes, including the DEBUG flag.

Debug enclaves can be inspected by a debugger, so every production verifier must reject the DEBUG attribute; accepting it is a common review finding.

Sealing and the rollback problem

Enclaves lose their memory when they exit or the machine reboots. To persist secrets, an enclave asks the CPU for a sealing key with EGETKEY. The key is derived from a fused per-CPU secret and a policy: bound to MRENCLAVE, only the identical enclave build on the same CPU can unseal; bound to MRSIGNER, any enclave from the same author and product ID with an equal or higher SVN can, which is how you migrate data across upgrades. In the SDK, sgx_seal_data uses the signer policy by default and sgx_seal_data_ex lets you choose.

Sealing gives confidentiality and integrity but not freshness. The host stores the sealed blob, so it can hand back an older version, for example a counter before it was incremented or a key before it was revoked. Treat rollback as your problem: keep a version number in an external service that the enclave contacts over an attested channel. Sealed data is also tied to one physical CPU, so migration needs a key broker.

Attestation after EPID: DCAP

Local attestation uses EREPORT: a report MACed with a key only the target enclave on the same CPU can derive, carrying the reporter's identities and 64 bytes of user data, usually a public-key hash that binds a channel.

Remote attestation wraps that report in a quote signed by a Quoting Enclave. The original scheme, EPID with the Intel Attestation Service, is gone: Intel discontinued it on 2 April 2025. The replacement is DCAP, the Data Center Attestation Primitives. Each platform has a Provisioning Certification Key certificate chained to an Intel root, and the Quoting Enclave signs quotes with an ECDSA key certified by that PCK. Verifiers fetch collateral (certificates, TCB information, revocation lists) through a local cache called PCCS, or use Intel Tiber Trust Authority as a hosted verifier. Checking the signature is only step one:

def verify_enclave(quote, collateral, policy, expected_nonce):
    # 1. Signature chain: quote -> attestation key -> PCK cert -> Intel root,
    #    with revocation lists checked (DCAP's quote verification library does this).
    result, tcb_status, advisories = verify_quote(quote, collateral)
    if result != "OK":
        raise Reject("bad quote")
    # 2. Platform patch level. Decide explicitly which statuses you accept.
    if tcb_status not in policy.accepted_tcb_statuses:      # e.g. {"UpToDate"}
        raise Reject(f"TCB {tcb_status}, advisories {advisories}")
    body = quote.report_body
    # 3. Enclave identity and attributes.
    if body.attributes.debug:
        raise Reject("debug enclave")
    if body.mrsigner != policy.mrsigner or body.isv_prod_id != policy.prod_id:
        raise Reject("unknown author or product")
    if body.isv_svn < policy.min_svn:
        raise Reject("enclave version too old")
    # 4. Freshness and channel binding: report_data = H(nonce || enclave public key).
    if body.report_data[:32] != sha256(expected_nonce + quote.enclave_pubkey):
        raise Reject("stale or unbound quote")
    return quote.enclave_pubkey

TCB status is where policy lives. After a microcode fix, unpatched platforms report OutOfDate and ones needing software mitigation report SWHardeningNeeded. Accepting everything defeats the point; accepting only UpToDate can take a fleet offline on advisory day. Decide in advance and use a time-boxed grace period.

The attack history and what it teaches

SGX has been attacked more than any other TEE, and the attack families tell you what your own code must defend.

Controlled channels. The OS manages page tables, so it can unmap enclave pages and watch the page faults, recovering secret-dependent access patterns at page granularity.

Microarchitectural leaks. Cache timing, branch prediction and transient execution have all been used against enclaves. Foreshadow extracted enclave memory and attestation keys through the L1 cache, SGAxe extended cache-based leakage to the quoting enclave, LVI injected values into enclave execution, ÆPIC Leak read stale data through an architectural bug in the APIC, and Downfall leaked through vector gather instructions.

Fault injection. Plundervolt undervolted the CPU through software-accessible interfaces to flip bits in enclave computations.

Physical bus attacks. In 2025, WireTap and Battering RAM used cheap interposers on the DDR4 bus against Scalable SGX. Because the memory encryption is deterministic, equal plaintext at the same address gives equal ciphertext, and WireTap reported recovering the DCAP quoting enclave's attestation key in under 45 minutes, letting an attacker forge quotes. Battering RAM targeted integrity. Physical attackers were outside the Scalable SGX threat model, which is the point: SGX on that hardware does not stop an insider with hands on the server.

Fixes arrive as microcode plus TCB recovery. Your part is strict attestation policy, constant-time code for secret-dependent work and small enclaves.

Running real software in an enclave

The SDK route, with Intel's SDK or Open Enclave, partitions your program at an EDL boundary: smallest TCB, most work. The library OS route runs mostly unmodified Linux binaries under a LibOS such as Gramine or Occlum, which forwards system calls to the host and checks results. In a Gramine manifest every file the enclave reads is hashed into the measurement as trusted, or mounted encrypted:

# keybroker.manifest.template (Gramine)
libos.entrypoint = "/usr/bin/python3"
loader.argv = ["python3", "/app/broker.py"]

sgx.debug = false
sgx.enclave_size = "1G"
sgx.max_threads = 16
sgx.remote_attestation = "dcap"

sgx.trusted_files = [
  "file:/usr/bin/python3",
  "file:/usr/lib/python3/",
  "file:/app/broker.py",
]
fs.mounts = [
  { type = "encrypted", path = "/data/", uri = "file:/data/", key_name = "_sgx_mrenclave" },
]

Check key names against your Gramine version; the syntax has changed between releases. Mainline Linux has carried the SGX driver since 5.11, exposing /dev/sgx_enclave; SGX must be enabled in firmware and DCAP needs a reachable PCCS.

Where SGX fits in AI systems

An SGX enclave cannot use a GPU, so it is the wrong tool for confidential LLM inference at scale; that is the job of confidential VMs with GPU confidential computing, covered in the confidential compute and end-user verifiable inference articles. SGX fits the small, sharp pieces around them:

  • Key brokers. A service that holds model-weight decryption keys and releases them only to inference nodes whose attestation passes policy. Its own logic is small and auditable, which is where a tiny TCB pays off.
  • Policy and redaction gateways. Prompt filtering, PII tokenisation or signing of audit records, where the operator should not be able to read the plaintext or forge the log.
  • Data clean rooms. Joining two parties' data inside an enclave both have attested.
  • Small-model CPU inference. Classifiers and embedding models that fit in EPC, accepting the CPU-only speed and the side-channel caveats.

Worked example: a model-weight key broker

A company ships encrypted model weights to inference nodes in several clouds. The weights key must never sit in plaintext on a disk an operator can read. The design puts a key broker in an SGX enclave. At start-up the broker unseals its master key, sealed with MRSIGNER policy so upgrades work. An inference node, itself a confidential VM, connects and presents its attestation evidence plus a fresh public key. The broker verifies the evidence against an allow-list of image measurements, then wraps the weights key to the node's public key.

Clients verify the broker with the function above before trusting the TLS key bound into its report data. Each release is logged through an OCALL, signed inside the enclave. Rollback appears at once: the host could replay an old sealed allow-list, so the allow-list comes from a signed, versioned source and the broker refuses versions older than the highest in the external log.

Failure modes

  • Accepting debug enclaves. Reject the DEBUG attribute in every production verifier.
  • Unbound quotes. A quote without a nonce and key hash in report data can be replayed or relayed.
  • Loose EDL. [user_check] pointers and missing length checks turn the boundary into a read or write primitive.
  • Sealed state rollback. The host can return an old blob; add external freshness.
  • Secret-dependent memory access. Lookup tables indexed by key bytes leak through page faults and caches.
  • Stale collateral. A PCCS that stopped refreshing keeps accepting platforms whose TCB status has changed.

Trade-offs

SGX offers the smallest TCB, at the cost of porting work, CPU-only execution, costly ECALL and OCALL transitions and a weaker physical story on Scalable SGX. TDX and AMD SEV-SNP protect whole VMs with less porting, a larger TCB and GPU support. Use SGX for small high-value logic and confidential VMs for the rest.

What to do next

  1. Write down your attacker: remote software, malicious operator, or physical insider. If physical, SGX alone is not enough.
  2. Inventory what must stay secret and how small the code that touches it can be.
  3. Prototype the trusted component with Gramine or an SDK, and keep the interface to a handful of calls.
  4. Implement the verifier checks above, including DEBUG, SVN, TCB status and nonce binding.
  5. Design rollback protection for every piece of sealed state.
  6. Run PCCS with monitoring and an explicit TCB grace-period policy.
  7. Continue with Intel TDX and model extraction defences for the rest of the picture.
Key takeaway: SGX shrinks trust to the CPU and a small enclave, proves what is running through DCAP attestation, and persists secrets by sealing. It does not protect availability, freshness of sealed state, secret-dependent access patterns or, on DDR4 Scalable SGX, against physical bus interposers. Use it for small high-value services such as key brokers, verify quotes strictly, and use confidential VMs for GPU inference.