Most certificate outages are not attacks but an expiry nobody acted on, or a rotation that replaced one file and forgot a second file that depended on it. Mutual TLS doubles the exposure, because both ends of every connection present a certificate and both ends validate the other against a trust bundle. A rotation is safe only when every pair of peers, old and new, can still complete a handshake at every moment of the change.
This article is about that change. It assumes you know how a TLS handshake and a certificate chain work, as covered in TLS, in depth, and what mTLS gives you, as covered in mTLS: both sides prove who they are. Here the focus is the operations: when to renew, how to load new material without restarting, how to change a certificate authority with zero failed handshakes, and how to know it worked.
What actually rotates
Each workload in an mTLS system holds two different things. The first is its identity: a leaf certificate, the private key that matches it, and the intermediate certificates that link the leaf to a root. The second is its trust bundle: the set of root certificates it accepts when the other side presents a chain. A server checks client chains against its bundle and a client checks server chains against its bundle, usually the same bundle.
That gives four rotation events with very different risk. Leaf rotation happens constantly and is local: one workload swaps its own certificate and key. Key rotation should happen with it; cert-manager, for example, generates a fresh private key on every renewal by default from version 1.18. Intermediate rotation changes what chain the leaves carry. Root rotation changes what everyone trusts, and it is the only one that needs coordination across the whole fleet, because a peer that does not yet trust the new root will reject every leaf issued under it.
One rule makes all four safe: at every moment, the chains being presented must be a subset of the chains every peer can verify.
Leaf timing: the arithmetic
A leaf has a lifetime L and is renewed at some fraction f of that lifetime. The time between renewal and expiry, (1 - f) times L, is your budget for everything that can go wrong with renewal: the issuer is down, the reload failed, the new file has the wrong permissions. If the issuer can be unavailable for up to two hours, the remaining window must be comfortably longer than two hours, plus retries.
Defaults in common tools reflect this. cert-manager issues 90-day certificates by default and renews two thirds of the way through, leaving about 30 days. SPIRE's default X.509 SVID lifetime is one hour, with agents renewing well before expiry, so the issuer must be highly available. Istio's workload certificates default to 24 hours and the sidecar requests a new one roughly halfway through; the ratio is configurable.
Two details matter at short lifetimes. Clock skew: if the issuer stamps notBefore as the current second and a peer's clock runs a minute behind, the new certificate is not yet valid for that peer. Issuers usually backdate notBefore by a little; check that yours does. Jitter: workloads that started together renew together, so add a random offset of a few percent of L to each renewal time.
| Lifetime | Renew at | Outage budget | Suits |
|---|---|---|---|
| 90 days | 2/3 (day 60) | about 30 days | edge certificates, slow-changing services |
| 24 hours | about 1/2 | about 12 hours | service mesh workloads |
| 1 hour | about 1/2 | about 30 minutes | SPIFFE-style identity with a highly available issuer |
Hot reload on both sides
Renewal is useless if the process keeps serving the old certificate from memory. Restarting on every rotation works for 90-day certificates but not for hourly ones. The robust pattern is to read the certificate, key and bundle at handshake time from an in-memory snapshot that a watcher replaces atomically. Go's standard library exposes the hooks for the server side: GetConfigForClient runs for every incoming handshake and can return a config with the current certificate and the current client CA pool. On the client side, GetClientCertificate can supply the current leaf, but there is no callback that reloads RootCAs, so the simplest correct approach is to build a fresh config per dial from the snapshot.
type Material struct {
mu sync.RWMutex
cert *tls.Certificate
pool *x509.CertPool
}
// Load replaces the snapshot only if every file parses; on error the old material stays live.
func (m *Material) Load(certFile, keyFile, bundleFile string) error {
c, err := tls.LoadX509KeyPair(certFile, keyFile)
if err != nil {
return err
}
pemBytes, err := os.ReadFile(bundleFile)
if err != nil {
return err
}
pool := x509.NewCertPool()
if !pool.AppendCertsFromPEM(pemBytes) {
return errors.New("trust bundle has no certificates")
}
m.mu.Lock()
m.cert, m.pool = &c, pool
m.mu.Unlock()
return nil
}
func (m *Material) snapshot() (*tls.Certificate, *x509.CertPool) {
m.mu.RLock()
defer m.mu.RUnlock()
return m.cert, m.pool
}
// Server: leaf and client-CA pool are read on every handshake.
// Clone the base so ALPN (h2 for gRPC) and other settings survive.
func (m *Material) ServerConfig() *tls.Config {
base := &tls.Config{MinVersion: tls.VersionTLS13, NextProtos: []string{"h2", "http/1.1"}}
base.GetConfigForClient = func(*tls.ClientHelloInfo) (*tls.Config, error) {
cert, pool := m.snapshot()
c := base.Clone()
c.Certificates = []tls.Certificate{*cert}
c.ClientAuth = tls.RequireAndVerifyClientCert
c.ClientCAs = pool
return c, nil
}
return base
}
// Client: no RootCAs callback exists, so build the config per dial.
func (m *Material) ClientConfig(serverName string) *tls.Config {
cert, pool := m.snapshot()
return &tls.Config{
MinVersion: tls.VersionTLS13,
ServerName: serverName,
RootCAs: pool,
Certificates: []tls.Certificate{*cert},
}
}Call Load from a file watcher and also on a timer, because watchers miss events when a tool replaces a directory through a symlink swap, as Kubernetes does for mounted secrets. Write the certificate and key together, atomically: a reader that sees a new certificate with the old key fails to load, and the guard above keeps the old pair rather than half of each. Sidecar proxies such as Envoy solve the same problem with the secret discovery service, which pushes new material over a stream; Service Mesh Architecture covers that control plane.
Rotating the CA in four phases
Root rotation is where outages happen, because trust must change everywhere before issuance changes anywhere. Run it as four phases, each with a gate you can measure.
- Distribute the union bundle. Push a bundle containing both the old root and the new root to every workload. Nothing changes on the wire yet. Gate: every workload reports the hash of the new bundle. A workload that has not picked it up is the one that will fail in phase 2.
- Switch issuance. Point the issuer at an intermediate under the new root. Leaves now renew into the new hierarchy at their normal pace. Because every peer trusts both roots, old and new leaves verify everywhere. Gate: a canary with a new leaf completes handshakes against peers that still hold old leaves, in both directions.
- Drain. Wait until no leaf chained to the old root can still be in use: the longest leaf lifetime, plus the longest connection age, plus a margin. Gate: your inventory shows zero live leaves under the old root.
- Remove the old root. Push a bundle with only the new root. Gate: no increase in handshake failures for one full leaf lifetime.
Each phase is reversible until the next one starts, which is the point. If phase 2 shows failures, point the issuer back; the union bundle still accepts both. Cross-signing, where the new root is also signed by the old one, can shorten phase 1 for peers you cannot update quickly, at the cost of a more complex chain to reason about.
Intermediates rotate the same way but more easily, because peers trust roots, not intermediates. The only requirement is that every leaf is served with its full chain up to, but not including, the root. A server that sends only its leaf works by accident while peers have the old intermediate cached and fails when the intermediate changes.
Long-lived connections
TLS validates certificates during the handshake and never again. An HTTP/2 or gRPC connection opened at 09:00 with a certificate that expires at 10:00 keeps working at 15:00, and a connection opened before phase 4 keeps trusting the old root afterward. Rotation cannot cut off a compromised identity whose connection is already open.
Bound connection age so rotation means what you think it means. gRPC servers in several languages support a maximum connection age setting that sends a graceful close after a configured period; proxies have equivalents. Set the maximum age below the leaf lifetime, and include it in the phase 3 wait. Clients must then handle a reconnect gracefully, which they should already do.
Short lifetimes versus revocation
Revocation lists and OCSP exist so a certificate can be withdrawn before it expires. Inside a private PKI they are rarely run well: clients must fetch them, decide what to do when the fetch fails, and most choose to ignore failures, which means revocation silently does nothing. Short lifetimes replace that machinery with arithmetic. A stolen 24-hour certificate is useless tomorrow; a stolen 90-day one needs a working revocation path.
The cost moves to availability. With one-hour leaves, an issuer outage longer than the renewal budget takes down every workload whose certificate expires during it. Run the issuer redundantly, alert on renewal failures long before expiry, and make sure a workload that cannot renew keeps its current valid certificate rather than discarding it. Short lifetimes are a strong default for service identity; pair them with authorization policy as described in Cloud Zero Trust in Depth.
Keep client certificates on a private CA. Public TLS hierarchies are moving to server-authentication-only certificates under browser root program policy, so a public certificate is the wrong tool for a client identity even where one still carries the client-authentication usage.
Worked example: moving a mesh to a new root
A platform runs 3,000 workloads with 24-hour leaves renewed at about 12 hours, and gRPC servers configured with a four-hour maximum connection age. The old root is due to expire in eight months, and the team wants the new root in place well before then.
Monday: the new root and an intermediate are created offline, and the union bundle is pushed. By Tuesday morning the bundle hash metric shows 2,996 of 3,000 workloads. The four stragglers are batch jobs that read the bundle only at start; they are restarted. Wednesday: issuance switches to the new intermediate. Within 24 hours every leaf has renewed into the new hierarchy. Handshake failure rates stay flat. Thursday: phase 3 needs 24 hours plus four hours plus margin, so the team waits until Friday and confirms from the inventory that no leaf under the old root remains. The following Monday the old root is removed from the bundle, and every step until then could have been reversed by changing one setting.
Monitoring that catches it early
- Remaining lifetime as a fraction. Alert when any live leaf has less than, say, 30 percent of its lifetime left. A fixed threshold such as seven days is meaningless for one-hour certificates.
- Renewal failures per issuer and per workload, alerting on the first failure rather than on approaching expiry.
- Bundle hash per workload, so phase 1 has a measurable gate.
- Handshake failures by reason: unknown authority, expired, bad certificate. A spike in unknown authority during a CA rotation means a phase ran early.
- Inventory of live leaves by issuer, built from issuance logs or from scraping endpoints.
# Inspect what a server presents to an authenticated client, and when it expires.
openssl s_client -connect payments.internal:8443 -servername payments.internal \
-cert client.pem -key client.key -showcerts </dev/null 2>/dev/null \
| openssl x509 -noout -subject -issuer -startdate -enddate
Failure modes
- Issuance switched before trust. New leaves are rejected by peers with the old bundle. Fix the order: union bundle first, and gate on it.
- Process never reloads. Files on disk are fresh, the process serves yesterday's certificate, and it expires in memory. Test reload by rotating in staging and checking what the endpoint presents.
- Missing intermediate. The server sends only its leaf; it works until the intermediate changes.
- Mismatched pair. A non-atomic write leaves a new certificate beside the old key. Write both together and validate before swapping.
- Herd renewal. Synchronised fleets overload the issuer at every cycle. Add jitter.
- Old root removed too early. Long-lived connections or batch jobs still hold old leaves. Include connection age and job duration in the drain.
Trade-offs
| Choice | Gains | Costs |
|---|---|---|
| Short leaves (hours) | no revocation needed, small blast radius | issuer must be highly available |
| Long leaves (months) | tolerates issuer outages | needs working revocation, rare rotations get forgotten |
| Union bundle rotation | every phase reversible | takes days, needs bundle telemetry |
| Max connection age | rotation actually takes effect | more reconnects and handshakes |
What to do next
- List every place a certificate, key or trust bundle lives, on both client and server side, and who updates each one.
- Compute the renewal budget for each lifetime you use and compare it with your issuer's realistic outage length.
- Verify every service reloads certificate and bundle without a restart by rotating in staging and inspecting what the endpoint presents.
- Set a maximum connection age below the leaf lifetime on long-lived gRPC and HTTP/2 servers.
- Add the five monitoring signals above, with lifetime-fraction alerts rather than fixed day counts.
- Write the four-phase CA rotation runbook now, and rehearse it in staging before the root is close to expiry.