An LLM API call carries the most sensitive data most companies send over a network: customer messages, source code, contracts, medical notes and a bearer API key that is worth money on its own. Every vendor says the API uses TLS, and that is true for the first hop. It says little about the six or seven other hops a production request takes before the answer comes back.
This page treats encryption in transit for LLM systems as a hop-by-hop engineering problem. It maps every network hop in a typical deployment, separates what TLS protects from what it does not, gives working client and server configuration, explains the streaming side channels that leak content from encrypted traffic, covers hybrid post-quantum key exchange, and ends with an audit you can run this week. The TLS handshake itself is covered in TLS 1.3 internals; here we assume it and focus on where LLM deployments get it wrong.
The hops of an LLM request
Start by drawing the path. A typical request crosses these hops:
| Hop | From, to | Typical state | What crosses it |
|---|---|---|---|
| 1 | client to edge | TLS 1.3, public CA | prompt, API key or session cookie |
| 2 | edge to gateway | often plain HTTP inside the VPC | same, after edge decryption |
| 3 | gateway to hosted provider | TLS, but proxies and SDK flags vary | prompt plus the provider key |
| 4 | gateway to inference pods | frequently plaintext gRPC or HTTP | prompt, completion, LoRA IDs |
| 5 | app to vector DB, cache | often plaintext, default ports | embeddings, retrieved chunks, cached answers |
| 6 | agent to tools and MCP servers | mixed, sometimes localhost HTTP | tool arguments, tokens, results |
| 7 | everything to logging | agents ship over TLS, but sidecars and debug dumps do not | full prompts and completions |
Hop 1 is the only one most teams check, and it is usually fine. The leaks happen on hops 2, 4 and 5, which are treated as internal and therefore trusted. They are inside a network that also runs CI runners, notebooks, third-party agents and, after any compromise, the attacker. Embeddings need the same care as text: published inversion attacks recover a large part of the original text from sentence embeddings, so a plaintext vector DB connection can leak the documents you indexed.
What TLS protects, and what it does not
TLS gives three guarantees on one connection: confidentiality of the bytes, integrity (they cannot be changed undetected) and authentication of the server, plus the client too when you use mutual TLS. Everything else is outside its scope, and four gaps matter for LLM traffic:
- Termination points see plaintext. CDNs, WAFs, API gateways, service-mesh sidecars and TLS-inspecting egress proxies all decrypt and re-encrypt. Each one is a place where prompts can be logged, cached or stolen, and it should be listed and owned.
- Metadata is visible. The SNI hostname (unless Encrypted Client Hello is in use), the server IP, the timing and the sizes of records. For streamed completions, sizes and timing are enough to infer a lot, as the side-channel section shows.
- Authentication is only as strong as verification. A client with verification disabled has encryption but no idea who it is talking to, so any machine-in-the-middle with a self-signed certificate reads everything.
- It protects data in motion, not at rest. Prompts in logs, caches and provider retention are a separate problem; see LLM API key handling for the credential side.
A client that cannot be quietly weakened
Most client-side failures are a single line: verify=False added to get past a corporate proxy, an old CA bundle baked into a container image, or SSLKEYLOGFILE left set from a debugging session. That variable makes many TLS libraries write session secrets to a file, which turns any packet capture into plaintext. A defensive client pins the minimum version, uses an explicit CA bundle and refuses to start if anyone weakened it:
import os, ssl, httpx
def llm_http_client(ca_bundle: str | None = None) -> httpx.Client:
if os.environ.get("SSLKEYLOGFILE"):
raise RuntimeError("SSLKEYLOGFILE is set: TLS secrets would be written to disk")
ctx = ssl.create_default_context(cafile=ca_bundle) # verification and hostname check on
ctx.minimum_version = ssl.TLSVersion.TLSv1_2
assert ctx.verify_mode == ssl.CERT_REQUIRED and ctx.check_hostname
return httpx.Client(
verify=ctx,
timeout=httpx.Timeout(60.0, connect=5.0),
http2=False, # set True only if the h2 extra is installed
follow_redirects=False, # a redirect must never move the bearer key to another host
)
# The OpenAI and Anthropic Python SDKs both accept a custom httpx client:
# client = OpenAI(http_client=llm_http_client())follow_redirects=False stops a hostile endpoint from bouncing your Authorization header to another origin. Building the context in code, rather than through REQUESTS_CA_BUNDLE or SSL_CERT_FILE, makes the trust decision visible in review. Add a CI check that fails on verify=False, CERT_NONE and check_hostname = False anywhere in the repository; these lines get added during an incident and never removed.
Internal hops: mTLS with authorization
Internal hops need mutual TLS, because both sides need to know who the other is. The inference pod should accept prompts only from the gateway, and the gateway should send them only to a real inference pod. A service mesh (Istio, Linkerd) can do this with sidecars and short-lived workload certificates. Without a mesh, the server side in Python looks like this:
import ssl
def inference_server_context(cert, key, client_ca):
ctx = ssl.SSLContext(ssl.PROTOCOL_TLS_SERVER)
ctx.minimum_version = ssl.TLSVersion.TLSv1_3
ctx.load_cert_chain(cert, key)
ctx.load_verify_locations(client_ca) # private CA that issues gateway certs only
ctx.verify_mode = ssl.CERT_REQUIRED # no client cert, no handshake
return ctx
# After the handshake, authorize the peer, not just authenticate it:
def peer_allowed(sslsock, allowed=("spiffe://prod/ns/llm/sa/gateway",)):
cert = sslsock.getpeercert()
uris = [v for k, v in cert.get("subjectAltName", ()) if k == "URI"]
return any(u in allowed for u in uris)The second function is the step people skip. A valid certificate from your private CA proves the peer is some workload you issued a certificate to. It does not prove it is the gateway. Check the SPIFFE ID or SAN against an allow list, as described in mTLS for internal services. Short certificate lifetimes help, but only if rotation is automated; certificate rotation without an outage covers that. Do the same for the vector DB and the cache: most of them support TLS and client certificates, but their defaults leave both off.
Streaming side channels and padding
Encrypting a stream does not hide its shape. Streaming APIs send each token, or a few tokens, as its own server-sent event, and TLS 1.3 records add a fixed overhead, so the size of each record reveals the length of each token. Weiss, Ayzenshteyn, Amit and Mirsky (USENIX Security 2024, "What Was Your Prompt? A Remote Keylogging Attack on AI Assistants") turned those length sequences back into text with a fine-tuned model. They reconstructed about 27% of responses accurately and inferred the topic of about 53% from passively captured traffic to major assistants. In November 2025, Microsoft researchers published Whisper Leak, which classified the topic of prompts from packet sizes and timing across 28 models. Several providers, including OpenAI, Microsoft, Mistral and xAI, then added mitigations such as a random obfuscation field in streamed chunks.
If you run a gateway that re-streams completions to users, you own this mitigation for hop 1. Two techniques work together: padding each event to hide token length, and batching tokens on a timer so the event count and timing reveal less:
import json, secrets, time
def padded_events(token_iter, flush_ms=50, bucket=64):
"""Re-chunk a token stream: flush every flush_ms, pad each event to a multiple of bucket bytes."""
buf, last = [], time.monotonic()
for tok in token_iter:
buf.append(tok)
if (time.monotonic() - last) * 1000 >= flush_ms:
yield _event("".join(buf), bucket); buf, last = [], time.monotonic()
if buf:
yield _event("".join(buf), bucket)
def _event(text, bucket):
body = {"delta": text, "p": ""}
raw = len(json.dumps(body).encode())
target = -(-raw // bucket) * bucket + bucket * secrets.randbelow(2)
body["p"] = "x" * (target - raw) # client ignores "p"
return "data: " + json.dumps(body) + "\n\n"Padding to a bucket removes the exact token length. The random extra bucket blurs the totals. Timer-based flushing hides token-by-token timing. All three cost bandwidth and a little latency, and Microsoft's own evaluation found no single mitigation removes the leak completely, so treat this as reducing the risk, not eliminating it. Turning streaming off for the most sensitive workflows is the strongest option.
Hybrid post-quantum key exchange
Recorded traffic can be decrypted later if the key exchange is ever broken, and a large quantum computer would break the elliptic-curve exchange TLS uses today. For prompts that stay sensitive for years, such as health records, legal strategy or source code, that "harvest now, decrypt later" risk is the main argument for hybrid key exchange. The standard hybrid group is X25519MLKEM768: X25519 combined with ML-KEM-768, the NIST-standardized Kyber. It is the default in OpenSSL 3.5 and later and in current Chrome, Edge and Firefox, and Cloudflare negotiates it with clients. Check what your own endpoints negotiate:
# OpenSSL 3.5+: offer only the hybrid group; failure means the server does not support it
openssl s_client -connect api.example.com:443 -groups X25519MLKEM768 -brief </dev/null
# look for a "Negotiated TLS1.3 group: X25519MLKEM768" style line in the outputThe cost is size. An ML-KEM-768 public key is 1,184 bytes against 32 for X25519, so the ClientHello no longer fits in one packet, and old middleboxes that mishandle a fragmented ClientHello will drop the connection. Roll it out per hop, starting with the hops you control end to end (gateway to inference), and watch handshake failure rates. Signatures and certificates remain classical for now; hybrid key exchange protects confidentiality, not authentication.
TLS inspection, proxies and pinning
Many enterprises route outbound traffic through a TLS-inspecting proxy that terminates the connection to the LLM provider with a certificate from a private CA installed on every laptop. That is a deliberate trade: the security team gains data-loss prevention on prompts, and the proxy becomes a store of every prompt in the company, so restrict its logs and retention and keep its upstream connection fully verified. If you build a client that may run behind one, accept a configured CA bundle instead of disabling verification. Certificate pinning will break under inspection, so most API clients should not pin; pin only in mobile apps you fully control, and only to a key you can rotate. The same reasoning applies to agents: route their outbound calls through a controlled egress point, as described in egress control for LLM workloads.
Worked example: auditing a RAG chatbot
Take a support chatbot: a React front end, an ingress controller, a Python gateway, a vLLM deployment, a Postgres pgvector store, Redis for semantic caching and an OpenTelemetry collector. An afternoon audit runs like this:
- Enumerate connections from the hosts, not the diagram. Run
ss -tnpon each pod, or read flow logs. You find six destinations. The diagram showed four; the collector and an old Prometheus exporter that scrapes the gateway's debug endpoint were missing. - Probe each listener.
openssl s_client -connect host:portagainst each one. Ingress: TLS 1.3. vLLM on 8000: plain HTTP. Redis on 6379: plain. Postgres: TLS available butsslmode=preferin the connection string, which silently falls back to plaintext and never verifies the server. - Check what crosses the plaintext hops. Capturing on the vLLM hop shows full prompts with customer emails. The Redis cache stores prompt-completion pairs in clear.
- Fix in order of exposure. Mesh mTLS for gateway to vLLM with an allow list of one SPIFFE ID. Redis TLS with
tls-auth-clients yes. Postgressslmode=verify-full. Drop the debug endpoint. Enable padding on the gateway's SSE stream. - Prove it. Re-run the probes in CI, and add a NetworkPolicy so only the gateway can reach vLLM at all. Encryption and reachability are separate controls, and you want both.
Failure modes
| Failure | Symptom | Fix |
|---|---|---|
verify=False in a client | works everywhere, including under attack | explicit CA bundle; CI grep |
sslmode=prefer and similar "opportunistic" modes | plaintext if TLS fails, silently | require and verify the server name |
| mTLS without authorization | any workload with a cert can call inference | check the peer identity against an allow list |
| expired internal certificate | outage at 3 a.m., then someone disables TLS | short lifetimes with automated rotation and expiry alerts |
SSLKEYLOGFILE left set | captures become plaintext | refuse to start when it is set in production |
| redirects followed with auth | API key sent to another host | disable redirects for authenticated clients |
| streaming without padding | topic and text inference from encrypted traffic | pad and batch events, or disable streaming |
Trade-offs
Mutual TLS on every internal hop adds handshake latency (small with session resumption and long-lived connections), certificate operations and debugging effort. Sidecar meshes move that cost into the platform at the price of extra hops and memory per pod. Library-level TLS needs every team to configure it correctly. Padding costs bandwidth on small streamed events. Hybrid post-quantum exchange costs a larger handshake and possible middlebox breakage. TLS inspection buys visibility and creates a high-value target. Pick deliberately, write the choice down per hop, and revisit it when the deployment changes.
What to do next
- Draw your request path from flow logs, not memory, and list every hop and every TLS terminator.
- Probe every listener with
openssl s_clientand record the version, cipher and group. - Add a CI check for disabled verification, opportunistic TLS modes and
SSLKEYLOGFILE. - Put mTLS with identity allow lists on gateway-to-inference, vector DB and cache hops.
- Turn on padding and batching for streamed completions, or disable streaming for sensitive flows.
- Test
X25519MLKEM768on the hops you own and track handshake failures as you roll it out. - Keep learning: TLS 1.3 internals, mTLS for internal services, certificate rotation, API key handling and egress control.