Every new TLS connection pays for a handshake before any application data moves: round trips to agree on keys, bytes to send a certificate chain, and CPU for signatures and key exchange. On a 10 ms data-centre link this is noise. On a 150 ms mobile link it can be the largest part of time to first byte, and on a busy edge server handshakes can be the largest CPU cost. Optimizing them is about two things: making each full handshake cheaper and making full handshakes rarer.

How the TLS 1.3 handshake and key schedule work is covered in TLS/SSL architecture. This article assumes that and is about performance: counting round trips, measuring them, shrinking what the server sends, resumption and its key management, using 0-RTT without opening a replay hole, revocation, and CPU. Protocol facts follow RFC 8446 (TLS 1.3), RFC 8470 (HTTP early data), RFC 8879 (certificate compression) and RFC 9000 (QUIC).

Advertisement

Where the time goes

Count round trips first, because on long links they dominate everything else. TCP needs one round trip before the client can send anything. A full TLS 1.2 handshake then needs two more; TLS 1.3 needs one, because the client sends its key share in the first message and the server can finish its side in one flight. Resumption in TLS 1.3 also takes one round trip but skips the certificate and signature. 0-RTT lets the client send its first request together with its first handshake message. QUIC merges transport and TLS so a new connection costs one round trip and a resumed one can send data immediately.

Round trips before the first response byte, at 100 ms RTTTCP + TLS 1.2 full1 + 2 RTT, then request = 400 msTCP + TLS 1.3 full1 + 1 RTT, then request = 300 msTCP + TLS 1.3 resumed (PSK)1 + 1 RTT, no cert = 300 msTCP + TLS 1.3 0-RTT1 RTT + request = 200 msQUIC 0-RTT / reused connectionrequest only = 100 msMake full handshakes cheaperTLS 1.3 everywhere; avoid HelloRetryRequestsmall first flight: ECDSA chain, compressionfits in initcwnd / QUIC 3x amplificationno blocking revocation fetchesMake full handshakes rarerconnection reuse: keep-alive, pooling, HTTP/2resumption: shared, rotated ticket keys0-RTT only for idempotent requestsmeasure resumption rate per edge nodeFigures add one RTT for the HTTP request itself. DNS and server processing time are excluded.
Handshake round trips by protocol and mode, and the two families of optimization. Times assume a 100 ms round trip.

The cheapest handshake is the one you skip. A reused connection pays nothing, so before tuning cryptography check how often clients and proxies open new connections at all. Keep-alive timeouts, connection pools and HTTP/2 multiplexing usually give larger gains than any handshake tweak; see connection pooling.

Measure before changing anything

curl breaks a request into phases. time_connect is when TCP finished and time_appconnect is when TLS finished, so their difference is handshake time:

curl -so /dev/null https://api.example.com/health -w \
  'tcp=%{time_connect} tls=%{time_appconnect} ttfb=%{time_starttransfer}\n'

To test resumption, save a session with OpenSSL and offer it on a second connection. The session summary line reports New or Reused:

openssl s_client -connect api.example.com:443 -tls1_3 -sess_out sess.pem < /dev/null
openssl s_client -connect api.example.com:443 -tls1_3 -sess_in sess.pem < /dev/null | grep -E '^(New|Reused)'

In production, log what the server sees. nginx exposes $ssl_protocol, $ssl_session_reused and $ssl_early_data as variables. Add them to the access log and you can chart protocol mix, resumption rate and 0-RTT use per edge node, which are the numbers every later section tries to move. Go clients can record the same from the client side with net/http/httptrace: the TLSHandshakeStart and TLSHandshakeDone hooks give timing, and ConnectionState.DidResume says whether resumption worked.

Advertisement

Get to TLS 1.3 and avoid the extra round trip

Enabling TLS 1.3 removes a round trip from every full handshake for every client that supports it. Keep TLS 1.2 for older clients if you must, with ECDHE and AEAD cipher suites only. Then watch for HelloRetryRequest. In TLS 1.3 the client guesses which key exchange groups the server will accept and sends key shares for them. If the server selects a group the client did not send a share for, it answers with a HelloRetryRequest and the client starts again, costing a full extra round trip. It is invisible unless you look for it.

Post-quantum key exchange makes the guess matter more. Hybrid groups such as X25519MLKEM768 combine X25519 with ML-KEM-768, whose encapsulation key alone is 1,184 bytes, so the ClientHello grows past one typical TCP segment. Clients commonly send both a hybrid share and a plain X25519 share. Order server group preference so that whatever you prefer is something your main clients already send, and test that middleboxes on your path tolerate a ClientHello split across packets. OpenSSL added ML-KEM in its 3.5 release; confirm what your build supports with openssl list -kem-algorithms before configuring it.

Make the first flight small

In TLS 1.3 the server's first flight carries its certificate chain and a signature. If that flight is larger than the sender may transmit before hearing back, the handshake costs an extra round trip. Over TCP, the initial congestion window is commonly ten segments (RFC 6928), roughly 14 KB. Over QUIC the limit is tighter: until the client's address is validated, a server may send at most three times the bytes it has received (RFC 9000), and the client's first datagram is padded to at least 1,200 bytes, so the budget is around 3,600 bytes unless the client sends more.

Chains add up quickly. Each certificate carries a public key, a signature and extensions, and an RSA 2048 public key and signature are 256 bytes each where ECDSA P-256 uses 65 and about 72. Practical steps:

  • Serve an ECDSA P-256 certificate to clients that support it, with an RSA certificate as fallback if you need one. nginx accepts two ssl_certificate directives for this.
  • Send only leaf and intermediates; never include the root, which clients already have.
  • Choose an issuing chain with a short path and ECDSA intermediates where your CA offers one.
  • Enable certificate compression (RFC 8879) where both your TLS library and clients support it; chains compress well because certificates share structure.
  • Keep the SAN list reasonable. A certificate naming hundreds of hosts is large on every handshake.

Resumption and the ticket key problem

Resumption lets a returning client skip certificate verification and the server's signature. In TLS 1.3 the server sends a NewSessionTicket after the handshake; the client presents it on its next connection as a pre-shared key. With the psk_dhe_ke mode, the resumed handshake still runs a fresh key exchange, so traffic remains forward secret.

Tickets are encrypted with a server-side ticket key, and that key is where deployments go wrong. Behind a load balancer every server must be able to decrypt tickets issued by any other, or resumption succeeds only when a client happens to land on the same machine. And the key must rotate: anyone who obtains it can decrypt tickets, which in TLS 1.2 means recovering the session keys of recorded traffic. Distribute keys from a secret store to all nodes, encrypt with the newest, accept the previous one or two for decryption, and rotate on a schedule shorter than your ticket lifetime. In nginx the first ssl_session_ticket_key file encrypts and later ones only decrypt:

ssl_protocols       TLSv1.2 TLSv1.3;
ssl_certificate     /etc/nginx/tls/ecdsa-fullchain.pem;
ssl_certificate_key /etc/nginx/tls/ecdsa.key;
ssl_certificate     /etc/nginx/tls/rsa-fullchain.pem;
ssl_certificate_key /etc/nginx/tls/rsa.key;

ssl_session_cache   shared:SSL:50m;
ssl_session_timeout 1d;
ssl_session_tickets on;
ssl_session_ticket_key /etc/nginx/tls/ticket.current;
ssl_session_ticket_key /etc/nginx/tls/ticket.previous;

keepalive_timeout   75s;
# log_format belongs in the http {} context
log_format tls '$remote_addr $ssl_protocol $ssl_session_reused $ssl_early_data $request_time';

Clients need configuration too, especially service-to-service clients that are rebuilt per request. In Go, keep one transport and give it a session cache:

tr := &http.Transport{
    TLSClientConfig: &tls.Config{
        MinVersion:         tls.VersionTLS12,
        ClientSessionCache: tls.NewLRUClientSessionCache(1024),
    },
    MaxIdleConnsPerHost: 64,
    IdleConnTimeout:     90 * time.Second,
}
client := &http.Client{Transport: tr} // create once, share everywhere

0-RTT without a replay hole

With 0-RTT the client sends application data encrypted under keys derived from the resumption secret, before the handshake finishes. That saves a full round trip but has two costs defined in RFC 8446. Early data is not forward secret with respect to the ticket key. And it can be replayed: an attacker who captures the first flight can send it again, and the server may process the request twice. TLS servers can limit replay with single-use tickets or by recording ClientHellos within a time window, but across a fleet of servers neither is perfect.

The practical rule is to accept early data only for requests that are safe to repeat, such as GET requests without side effects, and to let the application decide. RFC 8470 defines how: a front end that forwards a request received in early data adds the header Early-Data: 1, and an origin that will not risk a replay answers 425 Too Early, which makes the client retry after the handshake completes. In nginx this is ssl_early_data on; plus proxy_set_header Early-Data $ssl_early_data;. If you cannot audit which endpoints are idempotent, leave 0-RTT off; resumption alone keeps most of the benefit on reused tickets.

Revocation without blocking

Revocation checking used to add latency when a client fetched OCSP status during the handshake; OCSP stapling moved that fetch to the server. The landscape has changed. Let's Encrypt removed OCSP URLs from its certificates in May 2025 and turned off its OCSP responders in August 2025, citing the privacy cost of telling a CA which sites each visitor uses, and relies on CRLs instead. Major browsers already used their own aggregated revocation data rather than live OCSP fetches.

So for certificates without an OCSP URL there is nothing to staple, and stapling configuration for them is dead weight. For CAs that still run OCSP, stapling remains worth enabling, because a client that does check will otherwise make its own request. Remove any Must-Staple requirement from certificate orders against CAs that no longer support it.

CPU cost on the server

A full handshake costs one key exchange and one signature on the server. An RSA 2048 signature is far more expensive than an ECDSA P-256 signature, so serving ECDSA to most clients cuts handshake CPU substantially. Measure on your own hardware with openssl speed ecdsap256 rsa2048 rather than trusting published ratios. Resumed handshakes skip the signature. That gives a simple capacity model: handshake CPU per second is roughly new connections per second times the cost of a full handshake, plus resumed connections times the much smaller key-exchange cost.

When handshake CPU is the bottleneck, the order of fixes is: increase connection reuse, raise the resumption rate, move to ECDSA, then consider hardware acceleration or terminating TLS on a tier built for it, such as a CDN edge. Post-quantum hybrid key exchange adds computation and bytes; measure it before and after enabling, but ML-KEM operations are fast and the byte cost usually matters more than the CPU.

Worked example: a far-away API

An API served from one region has mobile clients with a 180 ms round trip. Measurements show median TLS handshake time of 370 ms and a resumption rate of 9%. Investigation finds four things: TLS 1.2 still negotiated by an old load balancer; an RSA certificate, so every full handshake pays the costlier signature; a 5.8 KB chain with the root bundled; and ticket keys generated locally on each of twelve servers, so a returning client resumes only if it lands on the same server.

The fixes: enable TLS 1.3 on the load balancer, which removes a round trip; switch to an ECDSA certificate with a chain of about 2.5 KB and no root, which cuts signing cost and would also fit QUIC's pre-validation budget if they add HTTP/3; distribute shared ticket keys rotated every twelve hours with the previous key accepted. Afterwards the median handshake is about 190 ms for new sessions, resumption rises to over 60% (clients reconnect within ticket lifetime), and server CPU spent on handshakes drops by more than half. 0-RTT is enabled only for two read endpoints after confirming they are idempotent. These figures are illustrative of the mechanism; your own logs give the real ones.

Failure modes

  • Resumption that never happens: per-node ticket keys behind a load balancer, or clients that create a new HTTP client per request and lose their session cache.
  • Ticket keys that never rotate: a stolen key exposes every ticket; in TLS 1.2 it exposes recorded traffic.
  • Rotation without overlap: dropping the old key at the moment of rotation invalidates every outstanding ticket and causes a spike of full handshakes.
  • Replayed early data: a POST accepted in 0-RTT and executed twice. Restrict early data to idempotent requests and honour 425.
  • Hidden HelloRetryRequest: a group preference mismatch silently adds a round trip to every new connection.
  • Large ClientHello breakage: some middleboxes mishandle a ClientHello split over packets once post-quantum shares are added.
  • Oversized chains: an extra certificate or a root in the bundle pushes the first flight over the window.

Trade-offs

Every optimization here trades something. Longer ticket lifetimes raise resumption rates but widen the window an exposed ticket key can be used. 0-RTT saves a round trip but adds a replay risk you must manage per endpoint. Dual certificates save CPU and bytes but double renewal and monitoring work. Long keep-alive timeouts save handshakes but hold memory and file descriptors. Choose from measurement: optimize the path that dominates your latency or CPU, not every path. For transport-level gains beyond TLS, see TCP Fast Open and QUIC.

What to do next

  1. Log protocol, resumption and early-data variables at the TLS terminator and chart them per node.
  2. Measure handshake time from a client with long round-trip time using curl's connect and appconnect timings.
  3. Enable TLS 1.3, then check for HelloRetryRequest by comparing your group preference with what clients send.
  4. Measure your certificate chain size; serve ECDSA with RSA fallback and remove the root from the bundle.
  5. Distribute ticket keys from a secret store to every node, rotate them on a schedule and keep the previous key for decryption.
  6. Fix clients that do not reuse connections or session caches before tuning anything else.
  7. Enable 0-RTT only for audited idempotent endpoints, forward Early-Data and return 425 elsewhere.
Key takeaway: TLS handshake optimization comes down to fewer round trips, fewer bytes and fewer full handshakes. Use TLS 1.3 and avoid HelloRetryRequest. Keep the server's first flight small enough to fit in the initial window. Reuse connections and make resumption work across the fleet with shared, rotated ticket keys. Accept 0-RTT only where a replay is harmless. Measure resumption rate and handshake time from logs, so every change is judged by numbers rather than assumptions.