Inside a data centre or a cluster, most service calls used to be plaintext HTTP protected only by the network: if you could reach the port, you could call the API. Mutual TLS replaces that with a cryptographic check in both directions. The server proves who it is, as in ordinary TLS, and the client also presents a certificate that the server verifies. Done well, every connection carries an authenticated identity for both ends, encrypted, and the service can decide what each caller is allowed to do.
The general mechanism, with the handshake messages and certificate fields, is explained in mTLS: both sides prove who they are, and rotating leaves and roots without an outage has its own page in mTLS certificate rotation. This page is about running mTLS between your own services: naming workloads, authorizing on identity, choosing where TLS terminates, moving a live system from plaintext to strict, and debugging the handshakes that fail along the way. The SPIFFE and Istio rules quoted were checked against their specifications on 2026-10-02.
What mTLS gives you, and what it does not
mTLS gives three things: confidentiality and integrity on the wire, proof of the server's identity to the client, and proof of the client's identity to the server. The third is the one plaintext networks lack, and it is the foundation for the zero trust model described in zero trust security: a request is trusted because of who sent it, not where it came from.
It does not give authorization. A server that accepts any certificate signed by the internal CA has replaced 'anyone on the network' with 'any workload in the company', which is better but still lets a compromised build agent call the payments ledger. It does not carry end-user identity: the certificate says the caller is the API gateway, not which customer the gateway is acting for, so user identity still travels in a token at the application layer. And it does not help when the workload itself is compromised, because the attacker holds a valid key; short certificate lifetimes limit how long a stolen key is useful, but only authorization limits what it can reach.
Naming workloads: SPIFFE IDs in the URI SAN
A certificate is only as useful as the identity in it. Common Name fields and hostnames are poor workload names: a hostname identifies an address, many workloads can share one, and Common Name matching is deprecated for server identity in most TLS stacks. The SPIFFE standard gives a better convention, used by SPIRE, Istio and several other meshes. Each workload gets a SPIFFE ID such as spiffe://prod.example.internal/ns/payments/sa/api: a trust domain followed by a path the organization defines.
The X.509 form, called an X.509-SVID, puts that ID in the certificate's Subject Alternative Name as a URI. The standard says an X.509-SVID MUST contain exactly one URI SAN, and so exactly one SPIFFE ID; leaf SVIDs must set the digitalSignature key usage and must not be able to sign certificates, and when the extended key usage extension is present it must include both serverAuth and clientAuth, because the same certificate is used in both directions. DNS SANs may appear alongside the URI, which is how one certificate satisfies both SPIFFE-aware peers and clients doing ordinary hostname checks.
Choose the path scheme once and keep it stable, because authorization rules will be written against it. Kubernetes meshes conventionally use namespace and service account, which ties identity to something deployers already manage; outside Kubernetes, use a team-and-service scheme and keep environment in the trust domain, so a staging workload can never be mistaken for production.
Authorizing the peer identity
Verification happens in two steps that must not be confused. Chain verification asks whether the certificate was issued by a CA in the trust bundle and is within its validity period. Authorization asks whether this particular identity may call this particular endpoint. Most TLS libraries do the first for you once you configure a CA pool; the second is your code or your mesh's policy. In Go, tls.RequireAndVerifyClientCert performs chain verification, and VerifyConnection runs after it with the verified peer chain available:
// Server side: require a client certificate, then authorize its SPIFFE ID.
package mtls
import (
"crypto/tls"
"crypto/x509"
"fmt"
)
var allowed = map[string]bool{
"spiffe://prod.example.internal/ns/payments/sa/api": true,
"spiffe://prod.example.internal/ns/payments/sa/reconciler": true,
}
func ServerConfig(bundle *x509.CertPool, getCert func(*tls.ClientHelloInfo) (*tls.Certificate, error)) *tls.Config {
return &tls.Config{
MinVersion: tls.VersionTLS13,
ClientAuth: tls.RequireAndVerifyClientCert, // chain checked against ClientCAs first
ClientCAs: bundle,
GetCertificate: getCert, // reloads the rotated leaf without a restart
VerifyConnection: func(cs tls.ConnectionState) error {
leaf := cs.PeerCertificates[0]
if len(leaf.URIs) != 1 {
return fmt.Errorf("peer has %d URI SANs, want exactly 1", len(leaf.URIs))
}
id := leaf.URIs[0].String()
if !allowed[id] {
return fmt.Errorf("peer %s not allowed", id)
}
return nil
},
}
}The client should check the server's identity as well. Standard verification compares ServerName with the certificate's DNS SANs; adding a URI check means a workload that somehow answers on the ledger's address with a valid but different identity is still refused.
// Client side: present our own SVID and check the server's identity too.
func ClientConfig(bundle *x509.CertPool, getCert func(*tls.CertificateRequestInfo) (*tls.Certificate, error), want string) *tls.Config {
return &tls.Config{
MinVersion: tls.VersionTLS13,
RootCAs: bundle,
GetClientCertificate: getCert,
ServerName: "ledger.payments.svc", // DNS SAN, checked by the standard verifier
VerifyConnection: func(cs tls.ConnectionState) error {
uris := cs.PeerCertificates[0].URIs
if len(uris) != 1 || uris[0].String() != want {
return fmt.Errorf("server identity mismatch")
}
return nil
},
}
}Handshake-level allowlists are coarse: they decide who may connect at all. Per-route decisions, such as letting the reconciler call only the read endpoints, belong in the request handler or a policy engine, using the identity the TLS layer extracted. Pass that identity to handlers explicitly, for example from r.TLS.PeerCertificates in Go's HTTP server, rather than trusting a header a proxy may or may not have set.
Issuing and rotating certificates
Internal mTLS only stays up if certificates are issued and renewed by machines. The usual shape is a workload issuer, such as SPIRE, a mesh's built-in CA or cert-manager backed by an internal CA, that attests each workload, issues a certificate valid for hours to a day and renews it well before expiry. Services must load the renewed certificate without restarting, which is what the GetCertificate and GetClientCertificate callbacks above are for, and every service needs the current trust bundle, which changes when the CA rotates. The timing arithmetic, the four-phase root rotation and what happens to long-lived connections are the subject of mTLS certificate rotation; plan them before you go strict, because an expired certificate in strict mode is an outage.
Where TLS terminates: library, sidecar or node proxy
| Option | How it works | Strengths | Costs |
|---|---|---|---|
| In-process library | Each service configures TLS itself, as in the Go example | No extra hop; identity available directly in code; works anywhere | Every language and framework must get it right; rotation code in every service |
| Sidecar proxy | A proxy beside each instance terminates mTLS; the app speaks plaintext to localhost | Uniform policy and metrics; no app changes | Extra hop and memory per instance; app sees identity only if the proxy forwards it |
| Node or ambient proxy | A shared proxy per node handles mTLS for local workloads | Fewer proxies; no injection | Larger blast radius per proxy; identity tied to node-level attestation |
A mesh is usually the fastest way to get mTLS across many services in many languages, and the sidecar trade-offs are covered in service mesh sidecar architecture. Libraries remain the right choice for services outside the mesh, such as batch jobs, legacy clients and anything on plain VMs. Most organizations end up with both, which is why the identity scheme and trust bundle must be shared between them rather than owned by the mesh alone.
Rolling out: from plaintext to strict without an outage
Turning on strict mTLS in one step breaks every caller that was not ready, and in a large system you do not know all of them. The safe migration has a permissive phase in which servers accept both plaintext and mTLS, so you can measure who still sends plaintext. In Istio, PeerAuthentication has the modes UNSET, DISABLE, PERMISSIVE and STRICT; UNSET inherits from the parent and is otherwise treated as permissive, and a mesh-wide policy goes in the root namespace of the installation. Istio principals are SPIFFE IDs without the spiffe:// prefix, and the trust domain must be the one the mesh is configured with; Istio's default is cluster.local.
apiVersion: security.istio.io/v1
kind: PeerAuthentication
metadata:
name: default
namespace: payments
spec:
mtls:
mode: STRICT # was PERMISSIVE during the migration
---
apiVersion: security.istio.io/v1
kind: AuthorizationPolicy
metadata:
name: ledger-callers
namespace: payments
spec:
selector:
matchLabels:
app: ledger
action: ALLOW
rules:
- from:
- source:
principals: # trust domain must match the mesh config (Istio default: cluster.local)
- prod.example.internal/ns/payments/sa/api
- prod.example.internal/ns/payments/sa/reconciler- Inventory. List every caller of each service from traffic data, not from documentation.
- Issue identities everywhere. Every caller, including jobs and VMs outside the mesh, gets a certificate from the same trust domain.
- Permissive. Servers accept both. Measure plaintext connections per caller; in Istio, destination-reported request metrics carry a
connection_security_policylabel that readsmutual_tlsfor mTLS traffic. - Fix callers until plaintext traffic to a namespace is zero for a full business cycle, including monthly jobs.
- Strict per namespace, starting with the least critical, with a one-line rollback to permissive.
- Authorize. Only after strict is stable, add identity allowlists, first in a dry-run or audit mode if your policy engine supports one, then enforcing.
Outside a mesh the same shape works with two listeners: keep the plaintext port, add an mTLS port, move clients over with a flag, log every plaintext request with its source, and close the plaintext port when the log is empty.
Worked example: locking down the ledger
The payments namespace has three services: an API, a reconciler and the ledger that both call. The team wants the ledger reachable only by those two. In the permissive phase the metrics show a fourth caller: a nightly export job on a VM outside the cluster, using plaintext HTTP to the ledger's node port, which nobody had listed. They issue the job a SPIFFE ID under ns/payments/sa/export from the same trust domain, give it a client configuration like the Go example, and after a week of zero plaintext connections switch the namespace to STRICT.
Next comes authorization. Adding the AuthorizationPolicy with only the API and reconciler principals would have broken the export job, so they add its principal too and log denials before enforcing. In the first enforcing hour, a denial appears from the API's canary deployment, which ran under a new service account. That is the system working: identity changed, so access changed. The fix is to deploy the canary under the existing account, not to widen the policy.
Debugging handshakes
# What identity does this certificate carry, and when does it expire?
openssl x509 -in client.pem -noout -subject -enddate -ext subjectAltName,extendedKeyUsage
# Handshake as a client with a cert; prints the server chain and any alert.
openssl s_client -connect ledger.payments.svc:8443 -servername ledger.payments.svc \
-cert client.pem -key client.key -CAfile bundle.pem -verify_return_error </dev/null
# Same, without a client cert: a STRICT server should refuse you.
openssl s_client -connect ledger.payments.svc:8443 -CAfile bundle.pem </dev/null- unknown ca: the side sending the alert does not have the issuer of the other side's certificate in its bundle. Common during CA rotation or across trust domains.
- certificate required: the server demands a client certificate and none was sent; the client is not configured, or its certificate failed to load.
- bad certificate or a closed connection right after the handshake: chain verification passed but your authorization callback rejected the identity. Log the SPIFFE ID on rejection, so this is visible.
- certificate expired or not yet valid: a renewal that failed, or a clock off on one side; check NTP before checking the issuer.
One TLS 1.3 detail confuses many people: the client can consider the handshake finished before the server has validated the client certificate, so a rejected client often sees its first read fail rather than the connect call. Look for the alert on the first request, and on the server's logs.
Failure modes
- Authentication without authorization: any internal certificate can call any service.
- Strict before inventory: an unknown batch job or legacy client breaks at the switch.
- Expired leaves: renewal stopped silently and every service fails together at the same hour.
- Restart-only reload: a service reads its certificate once at start and stops working after the first rotation.
- Identity from headers: the app trusts an identity header that a caller can forge when it bypasses the proxy.
- Shared certificates: several services use one certificate, so authorization cannot tell them apart.
What to do next
- Write down your identity scheme: trust domain per environment, path per workload, one URI SAN per certificate.
- Pick one service, turn on permissive mTLS and list every plaintext caller over a full business cycle.
- Give every caller a short-lived certificate with hot reload, and test a rotation in staging.
- Switch that service to strict, keep the rollback ready, and only then add identity allowlists in audit mode.
- Make denied identities and certificate expiry dates visible on dashboards before repeating for the next service.