Transport Layer Security gives a connection three properties: confidentiality (eavesdroppers see ciphertext), integrity (tampering is detected) and server authentication (the client knows it reached the holder of a private key that a trusted authority vouched for under that hostname). The protocol that delivers them is well understood and, in TLS 1.3, hard to misconfigure. What goes wrong in practice is everything around it: terminating in the wrong place, letting a certificate expire, shipping a container with no trust store, or disabling verification to make an error go away.
This article is about deploying and operating TLS. The handshake, key schedule and chain validation are explained in TLS/SSL architecture, and latency tuning in TLS handshake optimisation; here we assume the protocol works and focus on the decisions, automation and debugging that service owners own. Certificate-validity dates and the Let's Encrypt OCSP change were checked on 2026-10-01.
What you are actually deploying
On the server side, TLS needs four artefacts: a private key, which must never leave the machines that terminate TLS; a leaf certificate binding the public key to one or more hostnames; the intermediate certificates linking the leaf to a root; and a configuration saying which protocol versions and algorithms are allowed. On the client side it needs one artefact: a trust store, the set of root certificates the client accepts, plus the discipline to verify the server's chain and hostname against it.
Most outages trace to one of those five things: a key that leaked or was lost, a leaf that expired or does not list the requested hostname, a missing intermediate, a configuration that a client cannot negotiate, or a trust store that is missing, stale or bypassed. The rest of this article takes them in turn.
Where to terminate
| Topology | L7 routing and WAF | Encrypted inside? | Where keys live | Typical use |
|---|---|---|---|---|
| Edge termination | Yes | No | Load balancer only | Simple internal apps on trusted networks |
| Terminate and re-encrypt | Yes | Yes, second session | Load balancer and every backend | Most production web services; compliance regimes |
| Passthrough by SNI | No (hostname only) | Yes, end to end | Backends only | Tenants who must hold their own keys; non-HTTP TLS |
| Service mesh sidecars | Yes, at the sidecar | Yes, mutual | Every pod, issued automatically | Many internal services needing identity |
Edge termination is the cheapest and is increasingly hard to justify: anyone with access to the internal network sees passwords, tokens and personal data in clear, and zero-trust designs assume that network is hostile. Re-encryption costs one more handshake per backend connection, which connection pooling makes negligible. Passthrough gives true end-to-end encryption but blinds the router: no path-based routing, no header injection, and the backend sees the router's address unless you add the PROXY protocol. Mesh sidecars add mutual authentication, covered in mTLS.
Two handshake extensions make shared front ends possible. SNI (Server Name Indication) puts the requested hostname in the ClientHello so one IP address can serve many certificates and an L4 router can pick a backend without decrypting. SNI is sent in clear; Encrypted Client Hello aims to hide it but is still being rolled out, so treat hostnames as visible to the network. ALPN negotiates the application protocol, typically h2 or http/1.1, inside the handshake. A proxy that terminates TLS but forgets to advertise h2 silently downgrades every client to HTTP/1.1.
Certificates for operators
Hostname matching uses the Subject Alternative Name extension; modern clients ignore the subject's common name for this purpose. A certificate for api.example.com does not cover example.com, and a wildcard *.example.com covers one label only, so it matches api.example.com but not v2.api.example.com.
Choose ECDSA P-256 keys where clients support them: smaller certificates and faster signing than RSA 2048, which you may still need for old clients. Many servers can hold both and pick per handshake.
Serve the leaf plus intermediates, never the root. A server that omits an intermediate often works in browsers, which cache intermediates or fetch them, and fails in curl, Java and Go clients, the classic "works in the browser, not in the service" ticket.
# what does the server actually send?
openssl s_client -connect api.example.com:443 -servername api.example.com -showcerts </dev/null
# inspect a certificate file
openssl x509 -in fullchain.pem -noout -subject -issuer -dates -ext subjectAltName
# does the chain verify against a given root?
openssl verify -CAfile root.pem -untrusted intermediate.pem leaf.pem
Lifecycle: issue, renew, reload
Certificate lifetimes are shrinking by rule. Under CA/Browser Forum ballot SC-081v3, the maximum validity of publicly trusted TLS certificates is 200 days from 15 March 2026, 100 days from 15 March 2027 and 47 days from 15 March 2029. Let's Encrypt already issues 90-day certificates. The consequence is simple: if any certificate in your estate is renewed by a human, it will eventually expire in production. Automate all of them.
ACME (RFC 8555) is the standard automation protocol. A client creates an account key, places an order for hostnames, proves control of each one through a challenge, submits a certificate signing request and downloads the certificate. The challenge types matter for architecture: HTTP-01 serves a token on port 80 and needs every load-balanced node to answer; DNS-01 publishes a TXT record, works for wildcards and for hosts not reachable from the internet, but needs DNS API credentials, which are powerful secrets; TLS-ALPN-01 answers on port 443 with a special certificate.
Renew at about two thirds of the lifetime, so a 90-day certificate renews around day 60, giving a month to notice failures. Then make the new certificate take effect without dropping connections. Most proxies reload configuration gracefully; in your own code, load certificates through a callback rather than once at start-up:
// Go: serve whichever certificate was loaded most recently
type certStore struct {
mu sync.RWMutex
cert *tls.Certificate
}
func (s *certStore) Reload(certFile, keyFile string) error {
c, err := tls.LoadX509KeyPair(certFile, keyFile) // validates that key matches cert
if err != nil {
return err // keep serving the old certificate
}
s.mu.Lock()
s.cert = &c
s.mu.Unlock()
return nil
}
func (s *certStore) Get(*tls.ClientHelloInfo) (*tls.Certificate, error) {
s.mu.RLock()
defer s.mu.RUnlock()
return s.cert, nil
}
cfg := &tls.Config{MinVersion: tls.VersionTLS12, GetCertificate: store.Get}Revocation is the weakest part of the system. Let's Encrypt completed its move away from OCSP in 2025, browsers rely on their own revocation lists, and many non-browser clients check nothing. Short lifetimes are the practical answer: a stolen key is useful only until the certificate expires, which is one more reason to welcome the shrinking schedule.
Clients that verify
Most TLS vulnerabilities in application code are not cryptographic; they are clients that do not verify. The safe defaults exist, and the job is to not turn them off:
import ssl, socket
ctx = ssl.create_default_context() # CERT_REQUIRED, check_hostname=True, system roots
ctx.minimum_version = ssl.TLSVersion.TLSv1_2
# private CA? add it, do not disable checks:
# ctx.load_verify_locations("/etc/pki/internal-root.pem")
with socket.create_connection(("api.example.com", 443), timeout=5) as raw:
with ctx.wrap_socket(raw, server_hostname="api.example.com") as tls:
print(tls.version(), tls.selected_alpn_protocol(), tls.getpeercert()["notAfter"])The dangerous switches have recognisable names: verify=False in Python requests, InsecureSkipVerify: true in Go, rejectUnauthorized: false in Node.js, a trust-all TrustManager in Java, curl -k. Search your code base for them; each one turns TLS into encryption to an unknown party.
Trust stores differ by runtime, which is why a call works on a laptop and fails in a container. Slim images may ship without a CA bundle, so install the distribution's ca-certificates package. Python's requests uses the certifi bundle unless REQUESTS_CA_BUNDLE says otherwise, and OpenSSL honours SSL_CERT_FILE. Java uses its own cacerts keystore, managed with keytool. Node.js adds extra roots from NODE_EXTRA_CA_CERTS. Decide which bundle each service uses, and when you rotate a private root, ship the new root to every client before any server presents a certificate chained to it. Pinning, which narrows trust further, is covered in certificate pinning.
A configuration baseline
Allow TLS 1.3 and 1.2; disable everything older. For TLS 1.2 allow only ECDHE key exchange with AEAD ciphers (AES-GCM, ChaCha20-Poly1305); TLS 1.3's suites are all acceptable. Send HSTS on HTTPS responses once every subdomain you include serves HTTPS correctly. Rotate session-ticket keys. In nginx:
server {
listen 443 ssl;
http2 on; # nginx 1.25.1+; older: listen 443 ssl http2;
server_name api.example.com;
ssl_certificate /etc/tls/api/fullchain.pem; # leaf + intermediates
ssl_certificate_key /etc/tls/api/privkey.pem;
ssl_protocols TLSv1.2 TLSv1.3;
add_header Strict-Transport-Security "max-age=31536000" always;
}
Debugging playbook
| Error you see | Most likely cause | Check |
|---|---|---|
| unable to get local issuer certificate | Missing intermediate on the server, or missing root in the client store | s_client -showcerts and count the certificates sent |
| hostname mismatch / certificate is not valid for | SAN does not list the name, or the client connected by IP | openssl x509 -ext subjectAltName |
| certificate has expired | Renewal failed or reload never happened | Compare file dates with what the server presents |
| handshake failure / no shared cipher | Version or cipher policy mismatch | s_client -tls1_2 and -tls1_3 separately |
| Wrong certificate served | Client sent no SNI, so the default certificate came back | Add -servername; check the client library |
| Works in browser, not in service | Intermediate missing, or container lacks a CA bundle | Test from inside the container |
Worked example: from edge termination to re-encryption
A team runs 20 services behind a cloud load balancer that terminates TLS and forwards plaintext. An audit asks for encryption in transit everywhere. The plan: stand up an internal CA (or use the platform's private CA) and an automated issuer that gives each service a certificate with a short lifetime, for instance 30 days, renewed at 20. Add the internal root to the load balancer's backend trust settings. Change one low-risk service to listen on HTTPS with the reloading pattern above, point the load balancer's backend protocol at HTTPS, and verify end to end: the load balancer must reject a backend certificate from the wrong CA, which you prove by deliberately serving one in staging. Watch p99 latency; with pooled connections the extra handshake should be invisible. Then roll through the remaining services, add certificate expiry to monitoring for every endpoint, and delete the plaintext listener so nobody can fall back to it.
import datetime, socket, ssl
def days_left(host, port=443):
ctx = ssl.create_default_context()
with socket.create_connection((host, port), timeout=5) as s:
with ctx.wrap_socket(s, server_hostname=host) as t:
not_after = t.getpeercert()["notAfter"]
expires = datetime.datetime.fromtimestamp(ssl.cert_time_to_seconds(not_after), datetime.timezone.utc)
return (expires - datetime.datetime.now(datetime.timezone.utc)).days
# alert when days_left(host) falls below a third of the certificate lifetime
Trade-offs to decide explicitly
Visibility versus confidentiality: every hop that decrypts can inspect, route and log, and every such hop is a place keys live and data is exposed. Compatibility versus strictness: dropping TLS 1.2 or RSA certificates removes options from attackers and from old clients alike, so measure your client mix before tightening. Public versus private CAs: public certificates work everywhere but expose hostnames in Certificate Transparency logs; private CAs keep names private but put root distribution on you. Short lifetimes versus operational risk: shorter is safer once automation is reliable, and dangerous until it is. For transport choices underneath TLS, such as QUIC, see HTTP/3.
What to do next
- Inventory every TLS endpoint, its certificate expiry, issuer and renewal mechanism; anything renewed by hand is a future outage.
- Move all public certificates to ACME automation before the 100-day limit arrives in March 2027.
- Choose a termination topology per service and remove plaintext hops on networks you do not fully trust.
- Grep for disabled verification in every language you ship, and replace each case with an explicit trust anchor.
- Pin each service to a known CA bundle in its container image, and test calls from inside the image in CI.
- Add expiry monitoring and alert at one third of remaining lifetime, and rehearse a key compromise: revoke, reissue, reload.