Encryption at rest means data is stored as ciphertext and is only readable by a process that can get the key. Every major cloud turns it on by default for disks and object storage, which is why teams often tick the box and move on. For an LLM system that is not enough. The sensitive data is not in one database. It is spread across prompt logs, conversation history, retrieval corpora, vector indexes, caches, fine-tuning sets and model checkpoints. Default encryption protects against a narrow set of threats, and it does nothing about the ones that usually cause incidents.

This article sets out what each layer of at-rest encryption actually defends against, where LLM data comes to rest, and how to build envelope encryption with per-tenant keys, including runnable code. It then covers what crypto-shredding can and cannot delete. Protecting data while it is being computed on is a different problem, covered in encryption in use for LLM workloads.

What each layer defends against

Start with the threat model, because the three layers of at-rest encryption answer different questions.

LayerWho holds the keyDefends againstDoes not defend against
Disk or volume encryptionCloud provider or hostStolen or discarded physical mediaAnyone who can read the mounted volume
Storage-service encryption (default or customer-managed key)Provider KMS; you may control the keyMedia theft; with your own key, revocable access and audit of key useAny caller the service authorises, such as an over-broad role or a public bucket
Application or field-level encryptionYour service, via KMSLeaked backups, mis-scoped storage roles, operators browsing tablesA compromised application that holds decrypt rights

The key point is that the first two layers decrypt transparently. If a bucket of prompt logs is readable by an analytics role, default encryption hands that role plaintext. Only the third layer makes possession of the stored bytes insufficient, because a separate permission, decrypt on a specific key, is also required. That separate permission is what you can scope per tenant, audit per call and revoke.

Where LLM data comes to rest

Before choosing mechanisms, list every place LLM data rests. Most teams find more copies than they expected.

StoreWhat it containsTypical lifetime
Prompt and completion logsFull user input, retrieved context, model outputDays to years, often forgotten
Conversation historyMulti-turn state per userSession to indefinite
Retrieval corpus and chunksSource documents and their split textLife of the document
Vector indexEmbeddings plus chunk text or IDs as payloadRebuilt periodically; old snapshots linger
Response and semantic cachesPrompts and answers keyed by similarityMinutes to days
Fine-tuning and evaluation setsCurated real conversationsLong; copied between teams
Checkpoints and adaptersWeights that may have memorised training textLong; replicated to serving
Traces, queues, backups, replicasCopies of all of the aboveOften longer than the source

Two rows deserve emphasis. Embeddings are derived from text, but they are not anonymous. Published research on embedding inversion (for example the 2023 vec2text work by Morris and colleagues) recovered a substantial share of short input texts from their embeddings when the attacker could query the same embedding model. Treat vectors as sensitive as the text they encode. Model weights trained on customer data can also reproduce fragments of it, so a fine-tuned checkpoint belongs to the same data class as its training set.

Envelope encryption from first principles

Calling a key management service for every record would be slow and expensive, and KMS services cap how much data a single call may encrypt. Envelope encryption solves both. You generate a random data encryption key (DEK) locally and encrypt the data with it using an authenticated cipher such as AES-256-GCM. You then ask the KMS to encrypt (wrap) the DEK under a key-encryption key (KEK) that never leaves the KMS. Store the wrapped DEK next to the ciphertext and discard the plaintext DEK. To read, unwrap the DEK through the KMS, which checks permission and writes an audit record, then decrypt locally.

AES-GCM has two rules you must not break. Each encryption under one key needs a unique 96-bit nonce, because reusing a nonce with the same key leaks the XOR of the plaintexts and allows forgery. And with random nonces, NIST guidance limits one key to about 2^32 encryptions. Per-object or per-batch DEKs keep you far below that. GCM also takes associated data (AAD): bytes that are authenticated but not encrypted. Put the tenant ID and record ID there. A ciphertext copied into another tenant's row then fails to decrypt instead of quietly leaking.

Envelope encryption: the KMS holds the key-encryption key; storage holds ciphertext plus a wrapped data keyApplicationLLM gateway, RAG indexerKMSper-tenant KEK, never exportedwrap(DEK)wrapped DEKAES-256-GCMDEK in memory onlyencryptStored recordnonce | ciphertext | tagStored alongsidewrapped DEK, key version, AAD idsDisk / snapshot thiefsees ciphertext onlyBackup copied outuseless without KMSTenant deleteddestroy KEK: data unreadableStorage-level encryption alone decrypts for any authorised caller; the application layer is what binds data to a tenant key.
Envelope encryption with a per-tenant KEK: stored bytes, backups and snapshots are useless without a KMS call that is scoped, audited and revocable.

Code: AES-GCM envelopes with per-tenant keys

The sketch below uses the cryptography package for AES-GCM. The KMS is an interface with wrap and unwrap methods, so you can back it with your cloud's KMS client. Those clients differ in names and request shapes, so check yours rather than copying one from memory.

import os, json, base64, time
from cryptography.hazmat.primitives.ciphers.aead import AESGCM

class Kms:                                   # adapter over your cloud KMS client
    def wrap(self, key_name: str, dek: bytes) -> bytes: ...
    def unwrap(self, key_name: str, wrapped: bytes) -> bytes: ...

def tenant_key(tenant_id: str) -> str:
    return f"llm-data/tenant-{tenant_id}"    # one KEK per tenant, created at onboarding

def encrypt_record(kms: Kms, tenant_id: str, record_id: str, plaintext: bytes) -> dict:
    dek = AESGCM.generate_key(bit_length=256)
    nonce = os.urandom(12)                   # 96-bit nonce, never reused with this DEK
    aad = f"{tenant_id}|{record_id}".encode()
    ct = AESGCM(dek).encrypt(nonce, plaintext, aad)   # ciphertext with the 16-byte tag appended
    return {
        "v": 1, "tenant": tenant_id, "id": record_id, "kek": tenant_key(tenant_id),
        "wrapped_dek": base64.b64encode(kms.wrap(tenant_key(tenant_id), dek)).decode(),
        "nonce": base64.b64encode(nonce).decode(),
        "ct": base64.b64encode(ct).decode(),
    }

class DekCache:                              # bound KMS traffic: keep unwrapped DEKs briefly
    def __init__(self, kms, ttl=300, max_items=10_000):
        self.kms, self.ttl, self.max, self.items = kms, ttl, max_items, {}
    def get(self, kek, wrapped):
        hit = self.items.get(wrapped)
        if hit and time.monotonic() - hit[1] < self.ttl:
            return hit[0]
        dek = self.kms.unwrap(kek, wrapped)  # permission check and audit happen here
        if len(self.items) >= self.max:
            self.items.clear()
        self.items[wrapped] = (dek, time.monotonic())
        return dek

def decrypt_record(cache: DekCache, rec: dict) -> bytes:
    dek = cache.get(rec["kek"], base64.b64decode(rec["wrapped_dek"]))
    aad = f"{rec['tenant']}|{rec['id']}".encode()
    return AESGCM(dek).decrypt(base64.b64decode(rec["nonce"]), base64.b64decode(rec["ct"]), aad)

The v field and the stored KEK name let you change the format or the key later without guessing. The cache TTL is a deliberate trade-off. A longer TTL means fewer KMS calls but a slower revocation, because a destroyed or disabled KEK only takes effect once cached DEKs expire.

Key hierarchy, rotation and separation of duties

Granularity. One KEK per tenant is the usual unit, because it is the unit of contract, deletion and audit. Within a tenant, use one DEK per object for large blobs (a document, a checkpoint) and one DEK per batch or per time window for high-volume small records such as log lines. Unwrap cost then stays proportional to batches, not rows.

Rotation. Rotating a KEK in a cloud KMS creates a new key version. New wraps use it, and old versions stay available for unwrapping. That alone does not touch your data. To retire an old version, re-wrap the stored DEKs: unwrap with the old version and wrap with the new one. That is a small metadata job, not a re-encryption of the payloads. Re-encrypting payloads under fresh DEKs is only needed if a DEK itself may have leaked, or a DEK has approached its usage limit.

Separation of duties. The people who administer keys should not be the people who can read the data, and the application identity should hold encrypt and decrypt rights on tenant keys but never permission to destroy them. Key destruction should go through a separate, logged workflow with a delay, because it is irreversible.

Vector stores, embeddings and caches

Vector search needs plaintext vectors in memory to compute distances, so encrypting each vector individually breaks approximate nearest-neighbour search. Schemes that search over encrypted vectors exist in research but carry large performance or leakage costs, so treat them as a specialised choice. The practical pattern has three parts:

  1. Isolate indexes per tenant (or per tenant group), and encrypt each index's storage with that tenant's key, using the vector database's customer-managed key feature if it has one, or an encrypted volume per index otherwise.
  2. Encrypt the payload, not the vector. Store chunk text in the index only as an envelope-encrypted blob, or store just an ID that points to an encrypted chunk store. Decrypt after retrieval, once the caller's tenant has been checked.
  3. Cover the snapshots. Index rebuilds leave old segment files and snapshots behind. Put them under the same key and the same lifecycle as the live index.

Caches need the same treatment. A semantic cache keyed by embedding similarity across tenants is a cross-tenant leak, whatever its encryption. Scope cache keys by tenant first, and encrypt cached values under the tenant key if the cache persists to disk. More on boundaries in tenant isolation for LLM systems.

Worked example: deleting a tenant

Take a support assistant serving 400 business tenants. Each tenant uploads knowledge-base articles, and every conversation is logged for 90 days for quality review. A tenant cancels and invokes their contract's deletion clause. Here is how a per-tenant key design handles it.

  1. The tenant's KEK is scheduled for destruction after the contractual grace period. In the meantime it is disabled, so unwraps fail at once (after cached DEKs expire).
  2. Every envelope-encrypted record wrapped under that KEK becomes unreadable: logs, conversation history, chunk payloads, cached answers on disk, and evaluation samples drawn from that tenant. This works only because they were all wrapped under the tenant's key.
  3. The tenant's vector index and its snapshots, encrypted under the same key, are dropped. Destroying the key also covers any copies a delete job misses.
  4. Backups taken before the deletion still contain wrapped DEKs, but without the KEK they cannot be decrypted, so there is no need to rewrite backup archives.

That is crypto-shredding: deleting data by destroying the only key that can decrypt it. The next section is about what it misses.

What crypto-shredding cannot reach

  • Data under a different key. Anything encrypted only with storage-level default keys, or under a shared platform key, survives. One forgotten export job breaks the guarantee.
  • Plaintext copies. Debug logs, tracing spans with prompt attributes, analytics extracts, spreadsheets used for labelling. Encryption cannot shred what was never encrypted.
  • Derived artefacts. A shared model fine-tuned on several tenants' conversations keeps whatever it memorised. Destroying keys does not remove it, so the only real options are not training shared models on tenant data, or retraining without it. Per-tenant adapters encrypted under the tenant key can be shredded.
  • Aggregates and embeddings in shared indexes. Vectors stored in a cross-tenant index under a platform key are not covered.
  • Third parties. Model providers or labelling vendors that kept copies are outside your key hierarchy. Their retention terms are part of your deletion story.

Failure modes

  • Nonce reuse from a counter that resets on restart, or a nonce derived from the record ID. Use random 96-bit nonces with fresh DEKs.
  • Missing AAD, so a ciphertext moved between tenants decrypts cleanly. Bind tenant and record IDs.
  • KMS as a hard dependency on the request path. A KMS outage or quota limit becomes an outage for the assistant. Cache DEKs, batch unwraps, and alert on KMS error rate.
  • Over-broad decrypt grants. A wildcard permission on all tenant keys turns per-tenant keys into theatre. Grant per key ring, and review the grants.
  • Accidental key destruction. Irreversible. Use scheduled destruction with a waiting period and require two people.
  • Unaudited reads. If decrypt calls are not logged and reviewed, you cannot answer who read a tenant's data. Feed KMS audit logs into the same pipeline as your LLM audit logging.

Trade-offs

DecisionGainsCosts
Default storage encryption onlyZero effortNo protection against authorised-but-wrong readers; no per-tenant deletion
Customer-managed keys at the storage layerRevocation and key-use audit per bucket or databaseStill transparent to any authorised caller
Application-level envelope encryptionPer-tenant scoping, crypto-shredding, leaked backups uselessCode to write, KMS dependency, harder ad-hoc querying
Per-tenant vector indexesClean isolation and shreddingMore indexes to operate; less sharing of capacity
Long DEK cache TTLFewer KMS calls, lower latencySlower revocation

Most systems settle on customer-managed keys at the storage layer everywhere, plus application-level envelope encryption for the stores that hold raw prompts, chunk text and training data. Classification decides which stores qualify, see data governance for LLM systems. Key handling in code follows secrets management practice.

What to do next

  1. Inventory every store in the table above for your system, including traces, queues and backups, and note which key protects each.
  2. Find any store whose plaintext is readable by a role broader than the service that writes it. Those are your first application-level targets.
  3. Create one KEK per tenant at onboarding, and give the application encrypt and decrypt rights only, never destroy rights.
  4. Adopt the envelope format with a version field, the KEK name, a random nonce and tenant-bound AAD.
  5. Move chunk text out of shared vector payloads into encrypted blobs, or into per-tenant indexes.
  6. Write and rehearse the tenant-deletion runbook end to end on a test tenant, including backups and caches.
  7. Ban prompt bodies from debug logs and trace attributes, and scan for violations.
  8. Alert on KMS error rate and on decrypt volume per tenant that deviates from its baseline.
Key takeaway: Default disk and storage encryption only stops media theft; any authorised caller still reads plaintext. For prompts, chunks, embedding payloads and training data, add envelope encryption with a per-tenant KMS key, random nonces and tenant-bound associated data. Isolate vector indexes and caches per tenant, and know exactly what crypto-shredding misses: plaintext copies, shared models and third parties.