Security teams describe data in three states. At rest it sits on disk or in object storage, and AES with a managed key is routine. In transit it crosses a network, and TLS is routine. In use it is being processed, which means it is sitting in RAM, in CPU caches, in GPU memory, in a KV cache, in a log line or in a crash dump, and for most of computing history it had to be plaintext there because processors cannot add encrypted numbers. Encryption in use is the set of techniques that narrow that gap: hardware that encrypts memory so the host cannot read it, keys that are only released to a measured workload, and software designs that keep the most sensitive values out of the processing path entirely.
LLM systems make this state unusually important. A prompt can contain an entire medical record, a contract or a customer's account history, and the model must see the content to be useful. This article maps where that plaintext actually lives in an LLM pipeline, explains what each in-use control does and does not cover, and works through a design you can build. The confidential-serving architecture itself, along with homomorphic encryption and multi-party computation, is covered in secure inference for LLMs; here the focus is the data lifecycle around it.
Why in use is the hard state
Encryption at rest and in transit share a property: the data is not being computed on, so it can stay ciphertext until the moment it is needed. In use, the computation needs the values. There are three ways out. You can compute on ciphertext directly (homomorphic encryption, practical today for narrow operations and far too slow for a full transformer forward pass). You can split the computation between parties who each see only shares (MPC, with heavy communication costs). Or you can let the computation see plaintext but put it inside a boundary the attacker cannot reach, which is what trusted execution environments do. For production LLM serving, the third option plus aggressive minimisation of what enters the boundary is what actually ships.
That framing gives you the right question. Not "is the data encrypted" but "in which places does plaintext exist, for how long, and who can read each place". The answer is always a list, and every item on it is either protected by a boundary, minimised, or an accepted risk you have written down.
Where plaintext lives in an LLM pipeline
Follow one request through a typical hosted deployment and record every place its content exists unencrypted:
| Location | What is there | Who can normally read it | Primary control |
|---|---|---|---|
| Gateway or load balancer | Full request after TLS termination | Gateway operators, its memory and logs | Terminate TLS inside the boundary, or log nothing |
| Application server RAM | Prompt, retrieved documents, output | Host root, hypervisor, memory dumps | Confidential VM; minimise content |
| Swap and core dumps | Pages of process memory | Anyone with disk access | Disable or encrypt; LimitCORE=0 |
| Logs and traces | Prompts copied for debugging | Everyone with log access, often for months | Structured logging without content fields |
| GPU memory | Weights, activations, KV cache | Host drivers, other tenants on shared GPUs | GPU confidential mode; no unsafe sharing |
| PCIe transfers | Inputs and outputs moving CPU to GPU | Anyone who can observe the bus or DMA | Encrypted bounce buffers in CC mode |
| Prompt or prefix caches | Reusable prefixes, sometimes across users | Other requests via timing or reuse | Per-tenant cache keys; TTLs |
| Vector store and RAG | Chunks and embeddings | Database operators and backups | Encrypt at rest, filter by ACL at query time |
What memory encryption actually does
Confidential VMs on current server CPUs encrypt guest memory with keys the hypervisor never sees. On AMD SEV-SNP, the AMD Secure Processor generates a per-VM key and loads it into an AES-128 engine in the memory controller, so DRAM contents are ciphertext to the host; SNP adds the Reverse Map Table, which records which guest owns each page and stops the hypervisor from remapping or replaying guest memory. Intel TDX encrypts trust-domain memory with AES-XTS through the multi-key memory encryption engine and adds integrity protection against software and, in its cryptographic mode, against tampered memory. Both measure the initial guest image and sign that measurement in an attestation report.
For GPUs, NVIDIA's confidential computing mode on H100 and later puts the GPU inside the trust boundary: the driver in the confidential VM establishes a session with the GPU, and data moving across PCIe goes through bounce buffers in shared memory, encrypted with AES-GCM. The GPU produces its own attestation report. Treat the GPU package as the boundary and do not build your threat model on finer claims about how on-package memory is protected; descriptions differ, so read the current documentation for your generation.
What none of this covers is equally important. The code inside the boundary sees plaintext and can log it, cache it or send it anywhere. A bug that writes prompts to a log inside a confidential VM ships the log out as plaintext. Side channels remain a research area. And the boundary only helps if something checks the attestation before handing over secrets, which is the next section.
Keys that only an attested workload can get
Memory encryption protects data that is already inside the VM. The data usually arrives encrypted, under a data key from a key management service, and the whole design rests on the KMS releasing that key only to the right workload. The pattern is envelope encryption with an attestation condition: the workload produces a signed report of its measurement, the KMS or an attestation service verifies it, and the key policy permits decryption only when the measurement matches an approved build.
AWS expresses this directly for Nitro Enclaves. The enclave passes a signed attestation document as the Recipient of a KMS call, and the key policy can require that the enclave image digest, which corresponds to PCR0, matches an approved value:
{
"Sid": "DecryptOnlyFromMeasuredEnclave",
"Effect": "Allow",
"Principal": {"AWS": "arn:aws:iam::111122223333:role/inference-enclave-parent"},
"Action": ["kms:Decrypt", "kms:GenerateDataKey"],
"Resource": "*",
"Condition": {
"StringEqualsIgnoreCase": {
"kms:RecipientAttestation:ImageSha384": "<sha384 of the approved enclave image (PCR0)>"
}
}
}With the condition in place, KMS returns the plaintext key encrypted to the enclave's public key, so even the parent instance that relays the call cannot read it. Azure and Google offer comparable flows for confidential VMs through their attestation services and key release policies; the names differ but the shape is the same. Three rules make it hold up: pin measurements, not instance identities; keep the list of approved measurements in version control with the build that produced each one; and alert when a decrypt is refused, because that is either an attack or a deploy that forgot to register its new measurement.
Shrinking plaintext in software
The strongest in-use control is not to have the value in use at all. Most LLM tasks do not need a customer's real account number, email address or national ID; they need to know that one exists and to refer to it consistently. Replace identifiers with stable tokens before the text reaches the model, keep the mapping outside the model's reach, and restore the real values in the response only where the caller is entitled to see them.
import hmac, hashlib, re
EMAIL = re.compile(r"[\w.+-]+@[\w-]+\.[\w.-]+")
ACCT = re.compile(r"\b\d{10,12}\b")
class Pseudonymiser:
"""Replace identifiers with stable tokens before text reaches the model.
The HMAC key comes from the KMS at startup and never leaves this process.
The reverse map lives only for the request and is never logged."""
def __init__(self, hmac_key: bytes):
self.key = hmac_key
def _token(self, kind: str, value: str) -> str:
digest = hmac.new(self.key, f"{kind}:{value}".encode(), hashlib.sha256)
return f"[{kind}_{digest.hexdigest()[:10]}]"
def protect(self, text: str):
reverse = {}
def sub(kind):
def repl(m):
tok = self._token(kind, m.group(0))
reverse[tok] = m.group(0)
return tok
return repl
text = EMAIL.sub(sub("EMAIL"), text)
text = ACCT.sub(sub("ACCT"), text)
return text, reverse
@staticmethod
def restore(text: str, reverse: dict) -> str:
for tok, value in reverse.items():
text = text.replace(tok, value)
return text
# p = Pseudonymiser(key_from_kms)
# safe, rev = p.protect("Refund 4711002233 for ana@example.com")
# -> "Refund [ACCT_...] for [EMAIL_...]"; the model never sees the raw valuesKeyed HMAC tokens are deterministic, so the same email gets the same token across a conversation and the model can reason about "the same customer", but nobody without the key can reverse or dictionary-attack them. Regexes catch structured identifiers; names and free-text addresses need a named-entity detector, and PII leakage in LLM systems covers how to evaluate detectors and what they miss. Where analytics need equality joins on protected fields, deterministic encryption or tokenisation in the data store gives you that without exposing values, at the cost of revealing which records share a value.
Pseudonymisation also shrinks every other row of the plaintext table at once: the logs, the KV cache, the prompt cache and the vendor's systems now hold tokens, not identifiers.
Process hygiene inside the boundary
Memory encryption stops the host from reading RAM. It does nothing about the copies your own process makes. Core dumps write the whole address space to disk, swap writes pages, debug logging copies prompts into a pipeline with long retention, and exception trackers capture local variables including the request body. Turn these off deliberately:
# systemd unit for the inference service: no core files, no swap-backed secrets
[Service]
LimitCORE=0
MemorySwapMax=0
Environment=PYTHONFAULTHANDLER=1
# kernel side, on the host image
# kernel.core_pattern=|/bin/false (or route to an encrypted, access-controlled store)
# vm.swappiness=0 and encrypted swap if swap exists at allTwo further habits matter. First, define the log schema so content fields do not exist: request ID, tenant, token counts, latency, model version and a hash of the prompt for deduplication, but never the prompt. Second, accept that managed runtimes cannot reliably zeroise memory. Python strings are immutable and copied freely, so "wipe the secret after use" is not achievable; the realistic goal is short-lived processes, no persistence paths, and secrets fetched from a proper secrets manager rather than baked into images or environment dumps.
On GPUs, avoid sharing a device between tenants unless the sharing mode gives memory isolation, and remember that caches keyed only on prefix text can leak across users through timing. Key prefix caches by tenant.
Worked example: an insurance claims assistant
An insurer wants an LLM to summarise claim files that include policyholder names, policy numbers, medical notes and bank details, served from a GPU cluster the insurer does not physically control. Walking the plaintext table produces this design.
- Claim documents are stored encrypted with per-tenant data keys. Their key policy releases them only to inference VMs whose measurement matches the approved image.
- A pseudonymiser runs inside the confidential VM before prompt assembly. It replaces policy numbers, bank details and names with HMAC tokens. Medical content stays, because the summary needs it.
- Inference runs on GPUs in confidential mode, attached to the same VM. The VM checks the GPU's attestation before loading any data.
- Logs carry claim ID, token counts and latency. Prompts and outputs are never logged, and the prompt cache is keyed by tenant.
- The summary leaves the VM with tokens in place. The claims application, which is authorised to see identities, restores them for the adjuster.
What remains is a short, explicit list. Medical text is in plaintext inside an attested boundary, and the insurer accepts that. Identifiers never enter the model. The risk of a compromised inference image is handled by build provenance and measurement pinning. Each line on that list has an owner.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Prompts found in a log index | Debug logging or an exception tracker captured request bodies | Content-free log schema; scrub tracker payloads; test with canary strings |
| Key release works for any VM | Policy checks the role, not the measurement | Add the attestation condition; test that an unapproved image is refused |
| Deploy fails with decrypt denied | New image measurement not registered | Register measurements in the release pipeline before rollout |
| Identifiers still reach the vendor | Detector missed free-text names or new formats | Measure recall on labelled samples; add NER; alert on regex hits in outbound text |
| Crash leaks memory to disk | Core dumps enabled on the host or container | LimitCORE=0, encrypted dump store if dumps are needed |
| Confidential mode silently off | Instance type or driver fell back to normal mode | Verify attestation at startup and refuse to serve without it |
Trade-offs
Confidential VMs and GPUs cost some throughput, mostly from encrypting CPU-GPU transfers, and narrow your choice of instance types and drivers. Pseudonymisation costs some answer quality when the model would have used the real value, and needs a detector you maintain. Attestation-gated keys add a release step to every deploy. Against these, the alternative is trusting every operator, hypervisor, log pipeline and backup that touches the data, and that list is long. Start with the cheap controls (content-free logs, pseudonymisation, no dumps), then add hardware boundaries where the data or the hosting arrangement demands them. LLM infrastructure security covers the rest of the node and control-plane hardening that these controls depend on.