Ordinary LLM serving protects data in transit with TLS and data at rest with disk encryption, but during inference the prompt, the intermediate activations, the key-value cache and the model weights all sit in plaintext in CPU and GPU memory. Anyone who controls that machine can read them: the cloud operator, a compromised hypervisor, a malicious insider with host access, or another tenant who finds a way across. Secure inference is the set of techniques that keep those things confidential while the computation runs.
There are three families, and they are not interchangeable. Trusted execution environments (TEEs) run the ordinary model inside hardware-isolated, encrypted memory and prove what is running through attestation. Homomorphic encryption (HE) computes directly on ciphertext. Secure multi-party computation (MPC) splits the computation between parties so no single party sees the data. For LLMs today, TEEs are the only family that runs large models at near-native speed; the cryptographic families give stronger guarantees at costs that are still orders of magnitude too high for billion-parameter generation. This article explains why, and how to build the TEE path properly.
Start with the threat model
Name three parties before choosing a technique: the data owner (who sends prompts and receives outputs), the model owner (who owns the weights) and the operator (who runs the hardware). In a public API all three may be different; in an on-premises deployment they may be one organisation. Then list the assets: prompts, outputs, weights, the KV cache and logs. Secure inference answers the question of which of these the operator must not see, and whether the model owner may see the data or the data owner may see the weights.
It does not answer everything. A model that faithfully processes a prompt inside an enclave can still leak training data through its outputs, which is the problem covered in membership inference. An authorised client can still steal a model through its API, covered in model extraction. Secure inference protects the computation from the platform; it does not make the model's behaviour safe. Keep the hardening in LLM deployment hardening in place regardless.
| Approach | Protects against | Trust assumption | Practical for LLMs today |
|---|---|---|---|
| TEE (CPU CVM + GPU CC) | Operator, hypervisor, host OS, other tenants | Hardware vendor, attested code image | Yes, near-native speed |
| Homomorphic encryption | Server sees only ciphertext | Mathematics; client holds key | Small models and linear parts only |
| MPC | No single party sees data or weights | Parties do not collude | Research scale; heavy traffic |
The TEE architecture for LLM serving
A confidential GPU deployment has two isolation layers that must both be present. The CPU side is a confidential virtual machine built on AMD SEV-SNP or Intel TDX: the processor encrypts the VM's memory with keys the hypervisor never sees and protects its integrity, so the host can schedule the VM but not read it. The GPU side is a GPU running in confidential computing mode; NVIDIA's Hopper generation, the H100, was the first data-centre GPU with this capability (see the H100 deep dive for the rest of the chip).
The two sides are joined by the driver inside the CVM, which establishes an authenticated, encrypted session with the GPU. Because the GPU cannot read the CVM's encrypted private memory directly, data moves through bounce buffers: the driver encrypts a copy into shared memory, the GPU pulls it across PCIe, and decrypts it inside the GPU. On H100, GPU memory itself is protected by hardware access controls on a protected region rather than by memory encryption, on the reasoning that HBM sits on the same package as the GPU; the host's attempts to read that region are blocked.
Attestation and key release
Isolation without attestation is just a promise. Attestation produces signed evidence of what is running: the CPU produces a report containing the measurement of the CVM's initial image, and the GPU produces a report signed by a device key chained to the vendor, covering its firmware and confidential-mode state. A verifier checks the signatures, the certificate chains and revocation status, and compares the measurements against expected values.
The design point that makes this useful is to never ship secrets to an unverified environment. The model weights are stored encrypted, and a key broker releases the decryption key only to a CVM whose evidence passes a policy. Clients can do the same: verify evidence bound to the TLS key before sending a prompt, so that the TLS session is known to terminate inside the measured image.
def release_model_key(evidence, policy, kms):
cpu = verify_cpu_report(evidence.cpu_report) # signature, chain, TCB version
gpu = verify_gpu_report(evidence.gpu_report) # device cert chain, firmware, CC mode on
if cpu.debug_enabled or not gpu.cc_mode_enabled:
raise Denied("debug or non-confidential mode")
if cpu.launch_measurement not in policy.allowed_image_measurements:
raise Denied("unknown image")
if cpu.tcb_version < policy.min_tcb or gpu.firmware_version < policy.min_gpu_fw:
raise Denied("stale firmware")
if evidence.nonce != policy.expected_nonce: # freshness, prevents replay
raise Denied("stale evidence")
if cpu.report_data != sha256(evidence.tls_public_key):
raise Denied("evidence not bound to this TLS key")
return kms.wrap_for(evidence.tls_public_key, key_id=policy.model_key_id)The function names here are placeholders for your vendor SDK and key management service; the checks are what matters. Missing any one of them is a real bypass: accepting debug-mode VMs, accepting any image rather than a pinned measurement, skipping freshness, or failing to bind evidence to the channel so an attacker can relay a genuine report.
What it costs
The overhead is smaller than many teams expect, and it is concentrated in data movement rather than computation. The Hopper benchmark study on arXiv (2409.03992) measured LLM inference with confidential mode on and off and reported overhead below 7% for most typical queries, falling to nearly zero for larger models and longer sequences, because the GPU computes at full speed and the cost is encrypting transfers between CPU and GPU. The corollary is that workloads with a lot of CPU-GPU traffic relative to compute (small models, very short requests, frequent host-side sampling or swapping of KV cache to host memory) pay more. Measure with your own model, batch sizes and serving stack before quoting a number.
Operational costs are larger than performance costs: fewer instance types and regions, pinned driver and firmware versions, attestation services on the request path at startup, and images you must build reproducibly so their measurements are stable.
What TEEs do not protect
- The code inside. The enclave protects whatever image you measured. A logging statement that writes prompts to an external sink sends plaintext out of the enclave. Audit the image, disable debug endpoints, and treat the measurement as a statement about that code, not about safety.
- Side channels. TEEs have a long history of microarchitectural and timing attacks, and vendors patch them through firmware and TCB versions, which is why the policy checks minimum versions. Timing across tenants also matters at the application level: shared prefix caching can reveal whether another user sent the same prefix. Partition caches per tenant, as in tenant isolation.
- Availability. The operator can still stop, throttle or delete your VM. Confidentiality is protected, not uptime.
- Metadata. Request sizes, timing and token counts are visible outside the enclave unless you pad them.
- Trust in the vendor. You are trusting the CPU and GPU manufacturers' hardware, firmware and key infrastructure.
Homomorphic encryption: computing on ciphertext
HE lets a server compute on encrypted inputs and return an encrypted result that only the client can decrypt. For neural networks the usual scheme is CKKS, which encrypts vectors of approximate real numbers and supports addition and multiplication on them. Each multiplication adds noise and consumes a level of the ciphertext's budget; when the budget runs out the ciphertext must be refreshed by bootstrapping, an expensive operation. Linear layers, the matrix multiplications that dominate transformer FLOPs, map to HE reasonably well, though still far slower than plaintext.
The problem is the nonlinear parts. Softmax needs exponentials and division, GELU and layer normalisation need functions HE cannot compute exactly, so they are replaced with polynomial approximations that increase depth, error and cost, or the model is retrained with HE-friendly substitutes. Autoregressive generation then repeats the whole forward pass per token. The result is that HE is practical today for narrow jobs: scoring an encrypted embedding against a model, small classifiers, or a single linear layer, not for serving a chat model.
MPC: splitting trust between parties
In secure multi-party computation, inputs are split into random-looking shares held by two or more parties, who run a protocol that computes the model on the shares; no party learns the input or the weights, provided they do not collude. Private transformer inference systems typically combine techniques: HE or shared-secret arithmetic for the linear layers and garbled circuits or oblivious transfer for comparisons and nonlinear functions.
The bottleneck is communication. Published systems for BERT-base, far smaller than an LLM, report per-inference traffic in the tens of gigabytes for recent designs such as BOLT, and hundreds of gigabytes for earlier ones such as Iron. Those results were measured under different network settings, so compare orders of magnitude, not exact figures. Multiply by model size and by the number of generated tokens and it is clear why MPC for generative LLMs remains research. It is worth watching for small models, for private retrieval and for settings where no single hardware vendor is trusted.
Worked design: clinical notes to a hosted model
A hospital wants summaries of clinical notes from a 70-billion-parameter open-weight model hosted by a cloud provider, and its policy says the provider must not be able to read notes or summaries. The data owner is the hospital, the model owner is effectively the hospital (it licences open weights), and the operator is the cloud. HE and MPC are out on cost, so the design is a TEE.
Steps: build a minimal, reproducible serving image with no outbound logging of request bodies, and record its measurement. Deploy it on confidential VMs with GPUs in confidential mode. Keep the weights encrypted in object storage under a key held by the hospital's key broker, with the release policy from the code above. Have the hospital's gateway attest each serving instance before sending traffic, and pin sessions to the attested TLS key. Log request identifiers and timings, not text, and store any necessary text logs encrypted to a hospital key. Re-attest on restart, rotate model keys when the image changes, and alert on any attestation failure.
One caveat changes the plan: at 16-bit precision a 70-billion-parameter model needs about 140 GB for weights alone, more than one H100 holds, so it spans several GPUs. Multi-GPU confidential configurations have their own hardware and driver requirements, and GPU-to-GPU traffic must be protected too; check NVIDIA's current documentation for your exact topology. The single-GPU overhead figures from the Hopper study may not carry over, so benchmark the real multi-GPU deployment before committing capacity.
Operating secure inference
- Build images reproducibly and publish their measurements; a measurement nobody can recompute is not auditable.
- Treat attestation policy as code: version it, review changes, and update minimum firmware versions when vendors publish security advisories.
- Fail closed: an instance that cannot attest receives no keys and no traffic.
- Keep the host-visible surface minimal: no SSH into the CVM, no debug consoles, and metrics without payloads.
- Partition KV and prefix caches per tenant, and pad or bucket response sizes if metadata leakage matters.
- Layer it with the controls in defence in depth, since a TEE does nothing about prompt injection or unsafe outputs.
Failure modes
| Failure | Consequence | Prevention |
|---|---|---|
| Keys released without attestation | Weights readable by the operator | Key broker enforces the policy; no manual override |
| Debug mode accepted | Host can inspect memory | Reject debug in policy |
| Evidence not bound to TLS | Relay attack terminates TLS elsewhere | Hash TLS key into report data |
| Plaintext logging inside image | Prompts leave the enclave | Code review of image; egress allow-lists |
| Stale firmware accepted | Known side-channel fixes missing | Minimum TCB and firmware versions |
| Shared prefix cache | Cross-tenant timing leak | Per-tenant cache partitions |
Trade-offs
TEEs give near-native performance and run any model, at the price of trusting hardware vendors and accepting side-channel risk managed through patching. HE removes trust in the server entirely, at a cost that limits it to small or linear computations. MPC removes trust in any single party, at a communication cost that rules out large generative models today. Most teams should deploy TEEs with rigorous attestation now, and revisit the cryptographic options for narrow high-sensitivity tasks where a small model suffices.
What to do next
- Write down the parties, assets and which party must not see which asset.
- If the operator is the adversary and the model is large, plan a confidential VM with a confidential-mode GPU.
- Encrypt weights at rest and put key release behind an attestation policy that checks measurement, debug mode, firmware versions, freshness and channel binding.
- Build the serving image reproducibly, remove payload logging, and publish its measurement.
- Benchmark your model with confidential mode on and off at production batch sizes.
- Partition caches per tenant and decide whether metadata needs padding.
- Keep output-side controls, since secure inference protects the computation, not the model's behaviour.