Homomorphic encryption (HE) lets a server compute on data it cannot read. A client encrypts its input, the server runs arithmetic on the ciphertext, and only the client can decrypt the result. For LLM systems the attraction is obvious: send a prompt, a document or an embedding to a hosted service without the service ever seeing it. The catch is that every operation is orders of magnitude more expensive than the plaintext version, and the nonlinear parts of a transformer are exactly the parts HE handles worst.

This article is the mechanics page. It explains the CKKS scheme from first principles (slots, scale, levels, rescaling, rotations), shows why full encrypted transformer inference remains research-grade, and builds the workload that does fit today, encrypted embedding scoring, in real TenSEAL code. It ends with the parameter rules, the CKKS-specific attack that catches people out, and a checklist. For how HE compares with confidential GPUs and MPC at the system level, read secure inference for LLMs.

CKKS from first principles

CKKS is the HE scheme designed for approximate arithmetic on real numbers, which is what neural networks use. Its pieces, in the order you meet them:

  • Ring dimension N (poly_modulus_degree). Ciphertexts are pairs of polynomials with N coefficients. Larger N means more security headroom and more slots, and slower everything. Common values are 8192, 16384 and 32768.
  • Slots. One CKKS ciphertext packs N/2 real numbers, so N = 8192 gives 4096 slots. Additions and multiplications act on all slots at once, like SIMD lanes. Packing well is most of the performance work.
  • Scale. Reals are multiplied by a scale such as 2 to the 40 and rounded to integers before encryption. The scale sets precision: roughly, the bits of the scale minus the bits of accumulated noise are your significant bits.
  • Levels and rescaling. Multiplying two scaled values multiplies their scales, so after each multiplication the ciphertext is rescaled by dividing by one prime from the coefficient modulus chain. Each prime you drop is one level of multiplicative depth spent. When the chain runs out, you can compute nothing more without bootstrapping, an expensive refresh operation.
  • Relinearization and rotations. Ciphertext-ciphertext multiplication grows the ciphertext and needs relinearization keys to shrink it back. Moving values between slots (needed for sums and matrix products) needs rotations, which need Galois keys; these keys can be large.

The coefficient modulus chain is written as a list of prime bit sizes. TenSEAL's README uses [60, 40, 40, 60] with N = 8192: the outer 60-bit primes hold the final result and the special key-switching prime, the two 40-bit primes match a scale of 2 to the 40, and so the chain supports two rescales, meaning two sequential multiplications. Security bounds the total: the HomomorphicEncryption.org standard allows at most 218 bits of total modulus for N = 8192 at 128-bit classical security, and 438 bits for N = 16384. 60+40+40+60 is 200 bits, inside the bound. Add one more 40-bit prime and you exceed it, so deeper circuits force N = 16384, which doubles slots and roughly doubles or more the cost of each operation.

Why a transformer is deep

Count depth for one transformer layer. The linear parts are friendly: a matrix product with plaintext weights costs one multiplication level plus rotations. The rest is not. Softmax needs exponentials and a division, GELU needs an erf or tanh, and LayerNorm needs an inverse square root. HE evaluates only additions and multiplications, so each becomes a polynomial approximation, and a polynomial of degree d costs about log2(d) levels. Accurate approximations over the value ranges real activations reach need high degree. Several such functions per layer, times dozens of layers, needs bootstrapping many times per token.

Research systems have shown encrypted inference on small transformers by replacing nonlinearities with HE-friendly approximations, retraining for them, and accepting large latency. As of this writing that is not a production pattern for chat-scale models, and you should be suspicious of any vendor claim that it is without measured latency at a stated model size. The practical design rule is the opposite: keep the encrypted part shallow and linear, and push nonlinear steps to the party holding the key.

What fits today

Workloads that fit in a depth of one or two levels:

  • Encrypted retrieval scoring. The client embeds its query with a local model, encrypts the vector, and the server computes dot products against its plaintext document embeddings. One plaintext multiplication: one level.
  • Encrypted linear heads. A classifier or router on top of a client-side embedding, such as a sensitivity or topic classifier, is a vector-matrix product plus a bias. The client applies the softmax or sigmoid after decryption.
  • Encrypted aggregation. Summing encrypted usage statistics or gradients from many clients, where only the total is decrypted by a key holder. Additions are nearly free in depth.

Each of these keeps the expensive model (the embedding model) in plaintext on the client, so the client must be able to run it. That is the real requirement to check before choosing HE.

A linear head is the retrieval code below with a smaller matrix: the server multiplies the encrypted embedding by its plaintext weight matrix, adds a plaintext bias vector (an addition costs no level), and returns encrypted logits. Applying the sigmoid on the server with a degree-3 polynomial would need at least two more levels after the product, at least three in total, which the 200-bit chain cannot provide and which would force N = 16384. Returning logits avoids that cost entirely.

Encrypted retrieval scoring in code

The code below implements encrypted retrieval scoring with TenSEAL, a Python library over Microsoft SEAL. The client creates a context, keeps the secret key, and sends the server a public copy of the context; the server never holds the secret key.

Encrypted retrieval scoring: the server multiplies a ciphertext by plaintext embeddingsClient (holds the secret key)Local embedding modelquery text never leavesCKKS encrypt768 floats in 4096 slotsDecrypt scores, top-knonlinear steps hereServer (public context only)Plaintext doc embeddingsblocks of 1,000 columnsEncrypted vector x matrix1 level, rotationsEncrypted scoresone ciphertext per blockciphertext + public keysencrypted scoresStill visible to the server: query count, timing, which documents are fetched afterwardsNot provided by HE: integrity of the result (a malicious server can return garbage)Never send decrypted CKKS results back to anyone who saw the ciphertexts
Encrypted retrieval scoring. The client embeds and encrypts locally; the server computes encrypted scores against plaintext embeddings; the client decrypts and ranks. Access patterns and integrity remain outside what HE protects.
import numpy as np
import tenseal as ts

# ---- client ----
ctx = ts.context(ts.SCHEME_TYPE.CKKS, poly_modulus_degree=8192,
                 coeff_mod_bit_sizes=[60, 40, 40, 60])
ctx.global_scale = 2 ** 40
ctx.generate_galois_keys()                     # rotations for vector x matrix

public_ctx = ctx.copy()
public_ctx.make_context_public()               # drops the secret key
ctx_bytes = public_ctx.serialize()             # save_secret_key defaults to False

q = embed_locally(query_text)                  # 768 floats, L2-normalized
enc_q = ts.ckks_vector(ctx, q.tolist())
query_bytes = enc_q.serialize()

# ---- server ----
srv_ctx = ts.context_from(ctx_bytes)
assert not srv_ctx.is_private()
enc_q = ts.ckks_vector_from(srv_ctx, query_bytes)
blocks = []
for start in range(0, doc_emb.shape[0], 1000):            # doc_emb: (n_docs, 768)
    block = doc_emb[start:start + 1000].T.tolist()        # 768 x 1000
    blocks.append(enc_q.matmul(block).serialize())        # 1000 encrypted scores

# ---- client ----
scores = np.concatenate([ts.ckks_vector_from(ctx, b).decrypt() for b in blocks])
top_k = np.argsort(-scores)[:10]

The embed_locally function and doc_emb array are yours. The library calls shown match TenSEAL's README and tutorials as of its v0.3.18 release in September 2026; check the API against the version you install. Before production, compare decrypted scores with plaintext scores on a test set; with a 2 to the 40 scale the differences should be tiny relative to the gaps between top results, and if they are not, your parameters are wrong.

Worked example: sizing 50,000 documents

Size the example for 50,000 documents with 768-dimensional embeddings. The query, 768 values, fits in one ciphertext of 4096 slots. The corpus is split into 50 blocks of 1,000 columns, so the server runs 50 vector-matrix products and returns 50 ciphertexts. Each product spends one level, leaving one spare, which is why this fits the README chain.

Estimate sizes from the parameters rather than trusting a number from a blog. A ciphertext is two polynomials of N coefficients per remaining prime. Fresh, with the 60, 40 and 40-bit data primes present, that is 2 times 8192 times 3 primes, stored as 64-bit words: about 393 KB before compression, and less after a rescale drops a prime. Fifty result ciphertexts are therefore on the order of 10 to 20 MB per query, plus a one-time context with Galois keys that can be larger still. Measure with len(...serialize()) on your build; serialization compression changes the constants, not the shape.

TenSEAL computes each product with the diagonal method: for a 768-row block it performs up to 768 plaintext multiplications and 768 rotations, so compute grows with embedding size times number of blocks. Keep rows plus columns per block within the 4096 slots, which 768 plus 1,000 does. Measure the time on your hardware rather than estimating it. If it is too slow, the levers are fewer documents per query (a coarse plaintext prefilter on non-sensitive metadata), smaller embeddings, or batching several queries into one ciphertext's slots.

The CKKS decryption attack

CKKS has a trap that exact schemes do not. Because decryption returns an approximate value that includes noise, a decrypted result leaks information about the secret key. Li and Micciancio showed at Eurocrypt 2021 that an adversary who sees ciphertexts and the corresponding decrypted outputs can recover the key in practice, against the major libraries of the time. Standard IND-CPA security does not cover this; they defined the stronger notion IND-CPA-D.

The operational rule: never send raw CKKS decryptions back to the server or to anyone who saw the ciphertexts. In the retrieval design, the client keeps scores local and requests only document IDs. If a protocol needs to share decrypted values, round them or add noise first (some libraries offer this as a decryption option; check yours), or use an exact scheme such as BFV or BGV for that step.

Failure modes

  • Scale or level exhaustion. One multiplication too many and the library raises a scale-out-of-bounds error, or worse, results lose precision silently. Count depth before choosing parameters, and test decrypted against plaintext results.
  • Access pattern leak. The server cannot read scores, but sees which documents the client fetches next. Fetch a fixed-size padded set, or use private information retrieval if that matters.
  • No integrity. HE hides data; it does not prove the server computed correctly. A malicious server can return any ciphertext. Spot-check with known queries or combine with attestation.
  • Model leakage to the client. Encrypted scores against plaintext embeddings still reveal those embeddings to a client that queries enough. Rate-limit and treat embeddings as sensitive.
  • Key handling. Shipping a context with the secret key by mistake defeats everything. Assert is_private() is false on the server and store the client key like any other secret, for example with KMS envelope encryption.

Trade-offs

ApproachTrust assumptionCostFits
HE (CKKS)Only the math; server sees nothingHigh compute, large ciphertextsShallow linear steps on client data
Confidential VMs and GPUsHardware vendor and attestationLow overheadFull LLM inference
MPCParties do not colludeNetwork rounds between partiesTwo-party or multi-party settings
Local inferenceClient deviceClient must run the modelSmall models, sensitive input

For full LLM serving, memory-encrypted GPUs with attestation are the deployable option today; encryption in use for LLM systems covers that path. HE earns its place where you cannot trust any hardware vendor or operator and the computation is shallow.

What to do next

  1. Write the threat model: who must not see what, and whether integrity matters as much as confidentiality.
  2. Draw the computation and count multiplicative depth; move every nonlinear step to the key holder.
  3. Pick N and the modulus chain from that depth, and check the total bits against the HE standard bound for your N.
  4. Prototype with TenSEAL or OpenFHE, and compare decrypted results with plaintext on a test set.
  5. Measure ciphertext sizes and latency on real hardware before promising anything.
  6. Make sure no decrypted CKKS value is ever returned to a party that saw the ciphertexts.
Key takeaway: Homomorphic encryption lets a server compute on data it cannot read, but each multiplication spends a level and nonlinear functions cost many. Use CKKS for shallow, linear steps such as encrypted retrieval scoring or linear heads over client-side embeddings, push nonlinear steps to the key holder, size parameters from depth and the HE standard, and never return raw CKKS decryptions to anyone who saw the ciphertexts.