Ordinary TLS answers one question: is the server I connected to the one I meant? Mutual TLS adds the reverse: the server asks the client for a certificate and checks it, so both ends finish the handshake knowing a cryptographically verified identity for the other. Between services, that replaces IP allowlists and shared API keys with something that cannot be replayed or sniffed off the network.
This site already covers naming workloads and rolling mTLS out in mTLS for internal services, and rotation in mTLS certificate rotation. This article covers the parts in between that cause outages and silent security gaps: what actually happens on the wire, what the verifier checks and in what order, working Go code, and above all what happens to the client's identity when a proxy, load balancer or connection pool sits in the path.
The handshake, flight by flight
In TLS 1.3 (RFC 8446) the server opts into client authentication by sending a CertificateRequest message in its first flight, after ServerHello. By then the handshake keys exist, so the request and everything after it are encrypted. The request lists the signature algorithms the server accepts and can include a certificate_authorities extension naming acceptable CAs, which helps a client holding several certificates pick one.
The client verifies the server first: chain, name and the server's CertificateVerify, a signature over the handshake transcript proving possession of the private key. Only then does it send its own Certificate, CertificateVerify and Finished. Two consequences: the client's certificate is never visible to a passive observer, unlike TLS 1.2 where it crossed the wire in clear; and a client that has no acceptable certificate can send an empty Certificate, leaving the server to decide whether that is fatal. With RequireAndVerifyClientCert it is, and the server answers with a certificate_required alert.
TLS 1.3 also defines post-handshake client authentication, but RFC 8740 forbids it over HTTP/2, so for gRPC and most service traffic the client certificate is presented once, during the handshake, and applies to every request on that connection. Keep that fact in mind; it drives the connection-pool section below.
What the verifier checks, and what it does not
Verification is two separate jobs, and most incidents come from doing only the first.
| Step | Question | Typical failure text (Go) |
|---|---|---|
| Chain building | Does the leaf chain to a root in my trust bundle? | x509: certificate signed by unknown authority |
| Validity window | Is now between NotBefore and NotAfter? | x509: certificate has expired or is not yet valid |
| Key usage | Does the leaf allow clientAuth (or serverAuth)? | x509: certificate specifies an incompatible key usage |
| Proof of possession | Does CertificateVerify check against the leaf's public key? | handshake failure alert |
| Authorization | Is this identity allowed to call this service? | your own error, from your own code |
The first four are authentication, and a TLS library does them. The last is authorization and is your job. A server that trusts your internal CA and stops there accepts every workload that CA ever issued to, including the test job somebody left running. Authorize on a specific field, normally the SPIFFE ID in the URI SAN such as spiffe://prod.example/ns/payments/sa/payments-api, against an explicit allowlist per service or per route.
A working server and client in Go
Here is a minimal Go server and client. Certificates come from a store that a rotation agent updates in place; the callbacks read the current one per handshake, so rotation needs no restart.
func peerID(cs tls.ConnectionState) (string, error) {
if len(cs.PeerCertificates) == 0 {
return "", errors.New("no peer certificate")
}
leaf := cs.PeerCertificates[0]
if len(leaf.URIs) != 1 {
return "", fmt.Errorf("want exactly one URI SAN, got %d", len(leaf.URIs))
}
return leaf.URIs[0].String(), nil
}
var allowed = map[string]bool{
"spiffe://prod.example/ns/payments/sa/payments-api": true,
}
serverTLS := &tls.Config{
MinVersion: tls.VersionTLS13,
ClientAuth: tls.RequireAndVerifyClientCert,
ClientCAs: bundle, // *x509.CertPool holding only your internal roots
GetCertificate: func(*tls.ClientHelloInfo) (*tls.Certificate, error) {
return store.Current(), nil
},
VerifyConnection: func(cs tls.ConnectionState) error {
id, err := peerID(cs)
if err != nil {
return err
}
if !allowed[id] {
return fmt.Errorf("peer %s not authorized", id)
}
return nil
},
}
clientTLS := &tls.Config{
MinVersion: tls.VersionTLS13,
RootCAs: bundle,
ServerName: "ledger.prod.svc", // must match a DNS SAN on the ledger certificate
GetClientCertificate: func(*tls.CertificateRequestInfo) (*tls.Certificate, error) {
return store.Current(), nil
},
}Three details matter. ClientCAs holds only your internal roots, never the system pool. Authorization lives in VerifyConnection rather than VerifyPeerCertificate, because Go runs VerifyConnection on every connection including resumed sessions, while VerifyPeerCertificate is skipped on resumption. And if your certificates carry only a SPIFFE URI SAN with no DNS name, standard hostname checks on the client will fail; use go-spiffe's tlsconfig.MTLSClientConfig with tlsconfig.AuthorizeID, which verifies the chain and matches the ID instead.
Identity across proxies and load balancers
mTLS authenticates a TCP connection, and a connection has exactly two ends. The moment an intermediary terminates TLS, the server's peer is the intermediary. There are four arrangements.
- L4 passthrough. The load balancer forwards encrypted bytes; the handshake is end to end and the server sees the real client certificate. You lose L7 routing and per-request metrics at that hop.
- Terminate and forward identity. The proxy verifies the client, opens its own mTLS connection upstream and passes the client's identity in a header. Envoy does this with
x-forwarded-client-cert(XFCC), controlled byforward_client_cert_details. AWS Application Load Balancer mTLS has a passthrough mode that forwards the chain inX-Amzn-Mtls-Clientcertand a verify mode that validates against a trust store and forwards fields such asX-Amzn-Mtls-Clientcert-Subject. - Terminate without forwarding. Identity is simply lost; every request appears to come from the proxy. This is the default in many setups.
- Sidecars. In a mesh the sidecar terminates mTLS and talks plaintext over loopback to the application, passing identity in a header. Anything else in the pod that can reach that port inherits the sidecar's trust.
Forwarded identity headers are only as trustworthy as the hop that sets them. The proxy must strip any incoming copy from the client before adding its own (Envoy's default SANITIZE mode removes it; SANITIZE_SET removes and sets), and the backend must accept the header only on connections whose TLS peer is that proxy. Otherwise any caller can type x-forwarded-client-cert: URI=spiffe://.../admin and become admin.
Connection pools, long-lived connections and resumption
Identity is per connection, and connections are shared. Three consequences catch teams out.
The confused deputy. A gateway that pools upstream connections presents its own certificate for every request it forwards, on behalf of every caller. If the backend authorizes on the TLS peer alone, it authorizes the gateway, and the gateway can be tricked into making calls the original caller could not. Authorize on the forwarded caller identity plus the gateway's identity as the hop that vouches for it, or pass an end-user token alongside.
Connections outlive certificates. An HTTP/2 connection can stay open for days. The identity was checked once at the handshake; revoking or rotating the certificate does nothing to an established connection. Cap connection age at the server (gRPC servers support a maximum connection age setting) so every client re-handshakes within a bounded time.
Resumption carries identity forward. TLS 1.3 session tickets let a client resume without sending its certificate again; the server restores the original peer certificates from the ticket. RFC 8446 caps ticket lifetime at seven days. Set shorter ticket lifetimes than your certificate lifetimes and rotate ticket keys, and avoid 0-RTT early data for service calls because it can be replayed.
Worked example: payments to ledger through a gateway
Worked example. A payments-api deployment calls ledger through an Envoy gateway. Requirements: only payments may post entries; refunds, which also call ledger through the same gateway, may only read.
- Issue workload certificates from a private CA with
clientAuthandserverAuthkey usage, one SPIFFE ID per service account, 24-hour lifetime. - Gateway listener:
RequireAndVerifyClientCertsemantics against the internal bundle,forward_client_cert_details: SANITIZE_SETso XFCC always reflects the verified peer. - Gateway to ledger: its own mTLS connection with the gateway's certificate.
- Ledger: accept XFCC only when the TLS peer is
spiffe://prod.example/ns/edge/sa/gateway; extract the caller URI from it; allowPOST /entriesonly for payments andGETfor refunds. - Ledger server: maximum connection age of one hour, so a disabled service account loses access within an hour without waiting for certificate expiry.
Test it by calling ledger directly with a refunds certificate and a forged XFCC header claiming payments. The call must fail because the TLS peer is not the gateway.
Operating it: cost, CA choice and failure modes
Performance. A full handshake costs one round trip plus a signature and a few verifications on each side; with ECDSA P-256 keys the server-side signing cost is far below RSA-2048. In practice the cost is dominated by how often you handshake, so reuse connections and watch handshake rate per second as a metric. A sudden jump usually means a client stopped pooling.
Public CAs are leaving client authentication. Chrome's root program now requires publicly trusted TLS hierarchies to become dedicated to server authentication, and public CAs are removing the clientAuth key usage from the certificates they issue; Let's Encrypt has already dropped it from its default profile. If any service-to-service mTLS relies on public certificates as client certificates, it will break at renewal. Move it to a private CA.
Failure modes to rehearse. An expired intermediate (every client fails at once); a new root added to clients but not servers (one direction fails); a missing clientAuth usage after a CA template change; clock skew on a fresh node making new certificates "not yet valid"; a proxy silently falling back to plaintext upstream; and an allowlist that still contains a renamed service's old ID. Log the peer ID, the issuer and the failure reason on every rejected handshake, and alert on certificate time-to-expiry, not on errors.
Trade-offs
| Choice | Gains | Costs |
|---|---|---|
| Library mTLS in the app | Identity end to end, no extra hop | Every language stack must get it right |
| Sidecar or node proxy | Uniform policy, rotation handled for you | Plaintext local hop, identity via header |
| L4 passthrough at the LB | True peer identity at the server | No L7 routing or inspection at the LB |
| Terminate and forward identity | L7 features plus caller identity | Header trust must be enforced exactly |
| Short certificates, long connections | Small revocation window on paper | Real window is connection age; cap it |
mTLS is one layer of a zero-trust design: it proves which workload is talking, not which user it acts for and not whether the request is sensible. For the TLS fundamentals underneath, see TLS in networking.
What to do next
- List every hop between your two most critical services and mark where TLS terminates; that list is where identity can be lost.
- Set
ClientCAs(or the equivalent) to internal roots only and verify no service trusts the system pool for client certificates. - Add explicit peer-ID authorization in the connection or request path; delete any rule that accepts "any certificate from our CA".
- At every proxy, sanitize identity headers from clients, set them from the verified peer, and make backends trust them only from that proxy's identity.
- Cap server-side connection age and session-ticket lifetime below certificate lifetime.
- Find any mTLS that depends on publicly issued certificates and move it to a private CA.
- Run a game day: expire an intermediate in staging and forge an identity header, and confirm both fail loudly.