A service mesh puts a proxy in the path of every call between services, and traffic management is the set of decisions that proxy makes for each request: which version receives it, how long to wait, whether to retry, when to stop sending to a struggling host, and where to fail over. Done well, it turns a canary release into a configuration change and contains failures that would otherwise cascade. Done carelessly, a retry setting multiplies one slow database into an outage.

This page covers the routing and resilience objects in depth, using Istio's API as the concrete example because it is the most widely deployed, and shows the Kubernetes Gateway API form that meshes are converging on. The mesh architecture itself, including the data plane, control plane and identity, is covered in service mesh architecture in depth.

Advertisement

Where the decisions are made

Traffic management happens in the caller's proxy, configured by the control planecontrol plane (istiod)VirtualService + DestinationRulecart podapp containercaller-side proxy (sidecar or waypoint)1. match route: host, path, headers2. pick subset by weight: v1 95 / v2 53. timeout, retries, fault injection4. connection pool limits5. load balance, skip ejected hostsxDS pushcheckout v1 podssubset v1checkout v2 podssubset v2ejected host5 consecutive 5xx95%5%skippedmirror (optional)copy to shadow service, response discardedEvery decision is made per request, on the client side, before a byte reaches the server.
The control plane compiles routing objects into proxy configuration and pushes it over xDS. The caller's proxy applies route match, weighted subset choice, timeouts and retries, connection limits and load balancing for each request.

The most important fact about mesh traffic management is that it runs on the client side. When the cart service calls checkout, the proxy next to cart, or the waypoint proxy that serves checkout in Istio's ambient mode, decides where the request goes. The checkout pods never see a weight or a retry policy. This explains several surprises: a retry policy on checkout's routes is enforced by every caller, a caller without a proxy bypasses all of it, and metrics for a canary are best read from the caller's side.

The control plane, istiod in Istio, watches the routing objects, compiles them into Envoy listeners, routes and clusters, and pushes the result to every proxy. Changes take effect within seconds without restarting anything, which is what makes progressive delivery practical.

Two objects, two questions

ObjectQuestion it answersKey fields
VirtualServiceWhere should this request go, and how should the call behave?hosts, http[].match, route[].destination.subset and weight, timeout, retries, fault, mirror
DestinationRuleWhat happens once traffic is headed to this service?subsets (label selectors), trafficPolicy: loadBalancer, connectionPool, outlierDetection, tls

A VirtualService is a routing table evaluated top to bottom; the first matching rule wins. A DestinationRule defines named subsets, typically by a version label, and the policies applied to connections to them. Routes refer to subsets by name, so a VirtualService that names a subset the DestinationRule does not define sends traffic nowhere. Apply the DestinationRule first, then the VirtualService, and remove them in the reverse order.

Advertisement

Worked example: a canary for checkout v2

The shop team is releasing checkout v2. They want internal testers, who send the header x-canary: true, to always reach v2, and five percent of everyone else to reach it, with limits that stop a bad version from dragging down its callers.

apiVersion: networking.istio.io/v1
kind: DestinationRule
metadata: {name: checkout, namespace: shop}
spec:
  host: checkout.shop.svc.cluster.local
  trafficPolicy:
    connectionPool:
      tcp: {maxConnections: 200}
      http: {http1MaxPendingRequests: 100, http2MaxRequests: 500}
    outlierDetection:
      consecutive5xxErrors: 5
      interval: 10s
      baseEjectionTime: 30s
      maxEjectionPercent: 50
  subsets:
  - name: v1
    labels: {version: v1}
  - name: v2
    labels: {version: v2}
---
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata: {name: checkout, namespace: shop}
spec:
  hosts: [checkout.shop.svc.cluster.local]
  http:
  - match:
    - headers:
        x-canary: {exact: "true"}
    route:
    - destination: {host: checkout.shop.svc.cluster.local, subset: v2}
  - route:
    - destination: {host: checkout.shop.svc.cluster.local, subset: v1}
      weight: 95
    - destination: {host: checkout.shop.svc.cluster.local, subset: v2}
      weight: 5
    timeout: 2s
    retries:
      attempts: 2
      perTryTimeout: 500ms
      retryOn: connect-failure,refused-stream,unavailable,cancelled,503

Weights split requests, not users: one shopper's requests can land on both versions. If v2 changes a response format or session state, route by a header or cookie that pins a user to one subset instead. Promotion is a sequence of edits, 5, 25, 50, 100, each held long enough to compare error rate and latency between subsets from the callers' metrics. Automate the comparison and the rollback; canary deployment covers the analysis side.

Timeouts and retries, and how they multiply

Two defaults matter. According to Istio's reference, a route has no timeout unless you set one, and the default retry policy, if none is specified, is two attempts on connect-failure,refused-stream,unavailable,cancelled. Those conditions cover failures where the request most likely never reached the application, which is why they are safe to retry by default. Adding 5xx codes is a decision about idempotency: retrying a non-idempotent POST that failed after doing its work can charge a customer twice.

Size the budget from the outside in. With a 2-second route timeout and a 500 ms per-try timeout, three tries take at most 1.5 seconds plus backoff, so the last retry has a chance to finish before the overall deadline. If the per-try timeout were 1 second, the third attempt would be cut off by the route timeout and add load without adding success.

Retries multiply across hops. Suppose the edge gateway, cart and checkout each make up to three attempts, and the payments service behind checkout starts timing out. One user request can then become 3 x 3 x 3 = 27 requests to payments, arriving exactly when it is least able to serve them. The rules that prevent this: retry at one layer only, usually the one closest to the failure; keep attempts low; make sure the caller's deadline shrinks as it is passed down; and treat retries as load in capacity planning. gRPC architecture explains deadline propagation, which the mesh cannot do for you.

Circuit breaking and outlier detection

Istio has two mechanisms that are often both called circuit breaking. Connection-pool limits cap concurrency from each client proxy to the service. In the example, each caller proxy holds at most 200 TCP connections and 500 concurrent HTTP/2 requests to checkout, and up to 100 HTTP/1 requests can wait for a connection. Beyond that, the proxy fails the request immediately with a 503 instead of queueing it, and Envoy's access log marks it with the response flag UO, upstream overflow. Failing fast protects the caller's threads and memory, which is the point of the circuit breaker pattern. Because the limits apply per client proxy, the load checkout sees is the limit times the number of caller pods; size it with that in mind.

Outlier detection is per host. Each proxy watches results from each checkout pod; a pod that returns five consecutive 5xx errors is ejected from that proxy's load-balancing pool for 30 seconds, and repeat offenders are ejected for longer. maxEjectionPercent: 50 guarantees that at least half the hosts stay in the pool, so a shared dependency failing behind every pod cannot eject them all and turn a partial outage into a total one. Outlier detection catches the single bad node, such as a pod with a broken disk or a stuck connection pool, that the overall error rate hides.

Locality-aware routing and failover

With locality load balancing, a proxy prefers endpoints in its own zone and region, which cuts latency and cross-zone transfer cost. Failover to another locality depends on outlier detection: Istio's locality failover documentation states that outlier detection is required for failover to work, because ejection is what tells the proxy that the local endpoints are unhealthy. A DestinationRule with localityLbSetting and no outlierDetection keeps sending traffic to a failing local zone.

trafficPolicy:
  loadBalancer:
    simple: LEAST_REQUEST
    localityLbSetting:
      enabled: true
      failover:
      - from: us-east1
        to: us-central1
  outlierDetection:
    consecutive5xxErrors: 5
    interval: 10s
    baseEjectionTime: 30s

Mirroring and fault injection

Mirroring sends a copy of live requests to another destination and discards the copy's response. Istio documents it as best effort: the proxy does not wait for the mirrored service before returning the primary response, and mirrorPercentage controls the share copied, defaulting to all traffic. It is ideal for testing a rewrite against real request shapes. The danger is side effects: a mirrored request to a service that writes to a database or calls a payment provider performs that action for real. Mirror only to services whose writes go to an isolated store.

Fault injection adds delays or aborts to a percentage of requests, which is how you test that timeouts, retries and fallbacks actually behave as designed. One trap is documented in the API: when faults are enabled on the client side, timeouts and retries are not enabled on that route. Run fault experiments in a dedicated route or environment rather than on the route whose retry behaviour you are trying to observe.

The Gateway API form

The Kubernetes Gateway API is becoming the portable way to express the same intent. For mesh traffic, an HTTPRoute attaches to a Service as its parent instead of to a Gateway, an approach developed by the Gateway API's GAMMA initiative and supported by Istio and Linkerd. Backends are Services, so each version needs its own Service rather than a subset.

apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata: {name: checkout, namespace: shop}
spec:
  parentRefs:
  - {group: "", kind: Service, name: checkout, port: 8080}
  rules:
  - backendRefs:
    - {name: checkout-v1, port: 8080, weight: 95}
    - {name: checkout-v2, port: 8080, weight: 5}
    timeouts:
      request: 2s

The portable core covers matching, weights and timeouts. Retries, outlier detection and connection limits are less uniform across implementations, so check which of them your mesh supports in the Gateway API form before migrating a policy that depends on them. Mixing both APIs for the same host is a reliable way to confuse everyone; pick one per service.

Debugging traffic rules

  • Run istioctl analyze before applying; it catches references to undefined subsets and conflicting rules.
  • Inspect what a proxy actually received with istioctl proxy-config routes <pod> and istioctl proxy-config clusters <pod>. If a rule is not there, the proxy is not enforcing it.
  • Read response flags in Envoy access logs: UO means a connection-pool limit rejected the request, UH no healthy upstream, UF an upstream connection failure, URX the retry limit was exceeded.
  • Compare subsets with the istio_requests_total metric split by destination version and response code, measured at the source.

Failure modes

  • Retry storms from retries at several layers. Retry at one layer and cap attempts.
  • Timeouts shorter than the work. A 2-second route timeout in front of a 3-second batch endpoint fails every call while the backend completes the work anyway. Set timeouts from measured latency percentiles per route.
  • Silent bypass. A workload outside the mesh, or a call addressed to a pod IP, skips every rule.
  • Over-ejection when maxEjectionPercent is high and a shared dependency fails, leaving no hosts.
  • Rule order bugs. A broad match placed above a specific one swallows its traffic.
  • mTLS mismatch. A DestinationRule with a TLS mode that disagrees with the server's policy yields connection resets that look like application errors; mTLS certificate rotation covers the identity side.

Trade-offs

Moving traffic policy into the mesh gives one consistent implementation across languages and lets operators change behaviour without redeploying. It costs a hop of latency per proxy, memory per sidecar or waypoint, and a second place where behaviour is defined: an engineer reading the application code cannot see that every call is retried twice. Keep business-aware resilience such as fallbacks, idempotency keys and deadlines in the application, and use the mesh for uniform transport policy. For a managed control plane and its trade-offs, see service mesh architecture in the cloud.

What to do next

  1. Inventory every route and record its timeout and retry policy; anything without an explicit timeout has none.
  2. Pick the single layer that retries on each call path and remove retries elsewhere.
  3. Add connection-pool limits and outlier detection with maxEjectionPercent of at most 50 to your most critical service.
  4. Run one canary using weights, with automated comparison of error rate and latency between subsets.
  5. If you use locality load balancing, confirm outlier detection is configured and test failover by failing a zone in staging.
  6. Inject a delay fault in staging to prove that timeouts and fallbacks fire as designed.
Key takeaway: Mesh traffic management is a set of per-request decisions made in the caller's proxy: route match, weighted subset choice, timeouts, retries, connection limits and outlier ejection. Use it for canaries and uniform transport resilience, set explicit timeouts, retry at one layer only to avoid multiplication across hops, cap ejection so failures stay partial, and remember that locality failover needs outlier detection.