A service mesh moves the networking concerns every service needs, such as encryption, identity, retries, timeouts, load balancing and telemetry, out of application code and into a proxy that sits next to each workload. In the sidecar model that proxy runs as an extra container in every pod. The application makes plain HTTP or gRPC calls, and the proxies on each side turn them into authenticated, encrypted, observed and policy-checked exchanges.

That summary hides what matters when you run one. How does traffic reach a proxy the application does not know about? Where does the proxy get its routes and certificates? What happens when it starts after the application, or dies before it? This article takes the sidecar apart, traces one request through it, does the cost arithmetic and ends with failure modes and a checklist. For the higher-level case for a mesh and the product landscape, read service mesh architecture in the cloud first.

Advertisement

Two planes and one rule

The data plane is the set of proxies, one per pod. Envoy is the most common, used by Istio and others; Linkerd uses its own Rust proxy. The control plane watches the Kubernetes API, computes configuration for every proxy, issues workload certificates and pushes both down. The request path never depends on it: each proxy holds a local copy of what it needs, so a control plane outage freezes configuration but does not stop traffic.

One call from orders to payments through two sidecarsPod: ordersorders appplain HTTPsidecar proxyoutbound listeneriptables NAT rulesredirect all TCP to the proxyPod: paymentssidecar proxyinbound listenerpayments appplain HTTPiptables NAT rulesinbound redirectmTLS, SPIFFE identitiesControl planexDS config, certificates, service discoveryLDS RDS CDS EDS SDSKubernetes APIpods, servicesTelemetrymetrics, traces, logsThe applications never see TLS, retries or routing tables; the proxies do, and the control planeonly configures them. Nothing in the request path calls the control plane.
Each pod gets a proxy and a set of NAT rules. Configuration and certificates flow down from the control plane out of band.

Getting the proxy into the pod: injection

The mesh registers a mutating admission webhook with the Kubernetes API server. When a pod is created in a namespace labelled for injection, the API server sends the pod spec to the webhook, which returns a patch adding the proxy container, its configuration volume and an init step that sets up traffic redirection. Because injection happens at pod creation, changing the mesh version or its injection template affects only new pods; existing pods keep their old proxy until they are restarted, which is why mesh upgrades are rolled out as workload restarts.

If the webhook is down and fails closed, pods in injected namespaces cannot be created; if it fails open, pods start without a proxy and silently bypass policy. Choose deliberately, and alert on injected-namespace pods missing the proxy.

Advertisement

Getting traffic into the proxy: redirection

The application connects to payments:8080 as usual, and the kernel sends that connection to the proxy. An init container with the NET_ADMIN capability, or a CNI plugin that does the same work at pod network setup, installs NAT rules in the pod's network namespace. Every outbound TCP connection is redirected to the proxy's outbound listener, and every inbound connection to its inbound listener. The proxy recovers the original destination from the socket. Traffic from the proxy's own user ID is excluded, or its outgoing connections would loop back to itself.

# Simplified version of what an Istio-style init container installs (Istio defaults shown).
# The proxy runs as UID 1337 and listens on 15001 (outbound) and 15006 (inbound).
iptables -t nat -N MESH_REDIRECT
iptables -t nat -A MESH_REDIRECT -p tcp -j REDIRECT --to-ports 15001
iptables -t nat -N MESH_INBOUND
iptables -t nat -A PREROUTING -p tcp -j MESH_INBOUND
iptables -t nat -A MESH_INBOUND -p tcp --dport 15021 -j RETURN    # health port, not intercepted
iptables -t nat -A MESH_INBOUND -p tcp -j REDIRECT --to-ports 15006
iptables -t nat -N MESH_OUTPUT
iptables -t nat -A OUTPUT -p tcp -j MESH_OUTPUT
iptables -t nat -A MESH_OUTPUT -m owner --uid-owner 1337 -j RETURN  # proxy's own traffic: no loop
iptables -t nat -A MESH_OUTPUT -d 127.0.0.1/32 -j RETURN
iptables -t nat -A MESH_OUTPUT -j MESH_REDIRECT

The CNI variant exists because many platform teams refuse to grant NET_ADMIN to every pod. Either way the redirection is invisible to the application, which is also why it confuses: a protocol the proxy misparses on a port it treats as HTTP now fails inside the proxy, not in your code. Exclude ports and CIDR ranges that should not be meshed.

Configuration from the control plane: xDS

Envoy-based meshes configure proxies over the xDS family of gRPC streaming APIs. LDS delivers listeners (ports and filter chains), RDS routes (hosts and paths to clusters), CDS clusters (upstream services with load-balancing, pool and outlier settings), EDS endpoints (the pod IPs, which change most often) and SDS certificates and keys, all over one stream that updates as Kubernetes objects change.

Configuration is eventually consistent: just after a deployment some proxies still route to old endpoints, so retries on connection failure are not optional. By default every proxy also receives configuration for every service, so memory grows with mesh size; scope each workload to the services it calls (Istio's Sidecar resource) once you pass a few hundred services. Traffic policy is expressed in mesh resources and compiled into this configuration:

apiVersion: networking.istio.io/v1
kind: VirtualService
metadata: {name: payments}
spec:
  hosts: [payments]
  http:
  - timeout: 2s                 # whole call, including retries
    retries:
      attempts: 2               # retries after the first try
      perTryTimeout: 600ms
      retryOn: connect-failure,refused-stream,503
    route:
    - destination: {host: payments, subset: v1}
---
apiVersion: networking.istio.io/v1
kind: DestinationRule
metadata: {name: payments}
spec:
  host: payments
  subsets:
  - name: v1
    labels: {version: v1}
  trafficPolicy:
    connectionPool:
      http: {http2MaxRequests: 500}
    outlierDetection:           # eject a misbehaving endpoint for a while
      consecutive5xxErrors: 5
      interval: 10s
      baseEjectionTime: 30s
      maxEjectionPercent: 50

Be careful with retries at every layer. If three services in a chain each allow two retries, one user request can become 3 x 3 x 3 = 27 attempts at the deepest service during an outage, which turns a partial failure into a full one. Retry at one layer, normally the one closest to the caller that knows whether the operation is idempotent, and give every call an overall timeout that is shorter than its caller's.

Identity and mutual TLS

Pod IPs are reused within minutes, so mesh identity is not an address. Istio and Linkerd derive it from the Kubernetes service account, and Istio encodes it as a SPIFFE ID such as spiffe://cluster.local/ns/orders/sa/orders. At startup the proxy generates a private key that never leaves the pod, sends a certificate signing request to the mesh CA with its service account token as proof, and receives a short-lived certificate over SDS. It rotates the certificate well before expiry without a restart; Istio's default workload certificate lifetime is 24 hours.

Every connection between meshed pods is then mutual TLS: each side proves its identity and the channel is encrypted. Authorization policy can name callers by identity rather than by network location, which is what a zero-trust network means in practice. Zero trust in the cloud covers the wider model and mTLS certificate rotation covers what goes wrong when rotation does.

apiVersion: security.istio.io/v1
kind: PeerAuthentication
metadata: {name: default, namespace: payments}
spec:
  mtls: {mode: STRICT}          # reject plaintext once every client is meshed
---
apiVersion: security.istio.io/v1
kind: AuthorizationPolicy
metadata: {name: payments-callers, namespace: payments}
spec:
  selector: {matchLabels: {app: payments}}
  action: ALLOW
  rules:
  - from:
    - source:
        principals: ["cluster.local/ns/orders/sa/orders"]
    to:
    - operation: {methods: ["POST"], paths: ["/v1/charges"]}

Roll out in permissive mode, which accepts mTLS and plaintext, until telemetry shows every client using mTLS, then switch the namespace to strict. Strict first breaks every unmeshed caller.

One request, hop by hop

  1. orders calls http://payments:8080/v1/charges. DNS returns the Kubernetes service IP.
  2. The NAT rules redirect the connection to the orders proxy's outbound listener.
  3. The proxy matches the host to a route, picks a payments endpoint from its EDS list using its load-balancing policy, and applies the timeout and retry policy.
  4. It opens, or reuses, an mTLS connection to that pod, presenting the orders certificate and checking the payments certificate against the mesh CA.
  5. On the payments pod, the NAT rules redirect the incoming connection to the inbound listener, which terminates mTLS and learns the caller's SPIFFE identity.
  6. The inbound proxy evaluates the authorization policy for that identity, method and path, then forwards plain HTTP to the application on localhost.
  7. Both proxies emit metrics, access logs and trace spans. The application must still forward trace headers from inbound to outbound calls, or the spans will not join into one trace.

What it costs: a worked estimate

Measure your own costs, which depend on traffic, configuration size and proxy version, but do the arithmetic before adopting. Suppose a cluster runs 400 pods and each proxy uses 60 MiB of memory and 0.05 of a CPU core at typical load. The mesh then costs 400 x 60 MiB = 24,000 MiB, about 23.4 GiB, and 20 cores, before any control plane. If proxy memory is dominated by configuration for 1,000 services, scoping each workload to the 20 it calls can cut most of that.

Latency adds up per hop, and every service-to-service call crosses two proxies. A request that fans through a chain of five services crosses ten proxies. If each adds 0.3 ms at the median, the chain gains 3 ms, which is noise for a 200 ms page and a real cost for a 5 ms internal API. Tails suffer more, since each proxy can be throttled by a tight CPU limit. Give proxies CPU requests that match their real usage, and remember that every mesh upgrade restarts every pod.

Startup and shutdown ordering

A classic sidecar is just another container, and Kubernetes does not order container startup. If the application calls out before the proxy is ready, the redirected connection is refused; Istio added holdApplicationUntilProxyStarts as a workaround. On shutdown both receive SIGTERM and the proxy may exit first, failing in-flight requests. And in a batch Job the proxy keeps running after the main container finishes, so the Job never completes.

Kubernetes native sidecars fix all three. An init container with restartPolicy: Always starts before the application containers, can gate them with a startup probe, keeps running for the life of the pod, and is stopped only after the application containers have exited. The feature (KEP-753) was alpha in 1.28, beta in 1.29 and GA in 1.33. Meshes can inject the proxy this way; check whether your mesh version does so by default.

apiVersion: v1
kind: Pod
metadata: {name: report-job}
spec:
  restartPolicy: Never
  initContainers:
  - name: proxy
    image: example/mesh-proxy:1.0
    restartPolicy: Always        # this line makes it a native sidecar
    startupProbe:                # app containers wait until this passes
      httpGet: {path: /healthz/ready, port: 15021}
  containers:
  - name: report
    image: example/report:2.3   # when this exits, the sidecar is stopped and the Pod completes

Failure modes

SymptomLikely causeWhat to check
Connection refused at startupApp started before the proxyNative sidecar or hold-until-ready setting
503 with no upstreamEndpoints not yet pushed, or all ejected by outlier detectionProxy cluster and endpoint state from its admin interface
Works in permissive, fails in strictA caller without a proxyTelemetry for plaintext connections to the service
RBAC access deniedAuthorization policy names the wrong service accountCaller principal in the proxy access log
Error spike in a downstream outageRetries multiplied across layersRetry settings on every hop of the chain
Jobs never finishClassic sidecar keeps the pod aliveNative sidecars, or have the job stop the proxy

Always debug from what the proxy believes, not what the YAML says: its configuration dump, endpoint lists and the access log line for the failing request (istioctl proxy-config in Istio).

Sidecars versus sidecarless data planes

Sidecars give each pod a dedicated proxy with strong isolation, and the cost is one proxy per pod plus a restart for every upgrade. Sidecarless designs split the work. Istio's ambient mode, generally available since Istio 1.24, runs a shared per-node L4 proxy called ztunnel that handles mTLS and L4 policy, and optional waypoint proxies, deployed per namespace or service, for L7 features such as HTTP routing, retries and path-based authorization. Cilium can provide mesh features with eBPF plus a per-node proxy for L7. Upgrades do not restart application pods, and encryption-only workloads pay no L7 cost.

The trade-off is shared fate. A node proxy serves every pod on the node, so a bug or resource spike in it affects all of them, and L7 policy now runs in a separate hop rather than inside the pod. For many fleets, encryption-only L4 everywhere plus L7 waypoints for the handful of services that need routing is cheaper than sidecars on every pod.

When not to use a mesh

With a dozen services, a need for TLS only and a good client library for retries and timeouts, a mesh is a lot of machinery for little gain. An ingress or API gateway handles north-south traffic, as described in API gateway patterns, and the mesh earns its keep on east-west traffic between many services owned by many teams, where consistent identity, policy and telemetry are hard to get any other way. Load-balancing behaviour, in or out of a mesh, is covered in load balancing in system design.

What to do next

  1. Write down which problem the mesh solves for you: mTLS and identity, traffic policy, telemetry, or all three. If it is only mTLS, evaluate an L4 sidecarless mode first.
  2. In a test namespace, inspect one pod's injected spec, its NAT rules and its proxy configuration dump, so you know what normal looks like.
  3. Measure proxy memory, CPU and added p99 latency under your real traffic before rollout, and scale the arithmetic to your pod count.
  4. Roll out mTLS in permissive mode, confirm every caller is meshed from telemetry, then switch to strict namespace by namespace.
  5. Set retries at exactly one layer, give every call an overall timeout, and add outlier detection.
  6. Use native sidecars where your cluster supports them, and test pod startup, shutdown and Job completion explicitly.
  7. Scope proxy configuration per workload once the mesh passes a few hundred services, and alert on pods in injected namespaces without a proxy.
Key takeaway: A sidecar mesh is three mechanisms: kernel rules that send traffic to a local proxy, a control plane that streams configuration and certificates to it, and short-lived workload identities that make every hop mutual TLS. Understand those, measure the per-pod and per-hop cost, fix startup and shutdown ordering with native sidecars, keep retries to one layer, and choose a sidecarless data plane where L4 encryption is all you need.