A service mesh puts a proxy in the path of every call between services, and traffic management is the set of decisions that proxy makes for each request: which version receives it, how long to wait, whether to retry, when to stop sending to a struggling host, and where to fail over. Done well, it turns a canary release into a configuration change and contains failures that would otherwise cascade. Done carelessly, a retry setting multiplies one slow database into an outage.
This page covers the routing and resilience objects in depth, using Istio's API as the concrete example because it is the most widely deployed, and shows the Kubernetes Gateway API form that meshes are converging on. The mesh architecture itself, including the data plane, control plane and identity, is covered in service mesh architecture in depth.
Where the decisions are made
The most important fact about mesh traffic management is that it runs on the client side. When the cart service calls checkout, the proxy next to cart, or the waypoint proxy that serves checkout in Istio's ambient mode, decides where the request goes. The checkout pods never see a weight or a retry policy. This explains several surprises: a retry policy on checkout's routes is enforced by every caller, a caller without a proxy bypasses all of it, and metrics for a canary are best read from the caller's side.
The control plane, istiod in Istio, watches the routing objects, compiles them into Envoy listeners, routes and clusters, and pushes the result to every proxy. Changes take effect within seconds without restarting anything, which is what makes progressive delivery practical.
Two objects, two questions
| Object | Question it answers | Key fields |
|---|---|---|
| VirtualService | Where should this request go, and how should the call behave? | hosts, http[].match, route[].destination.subset and weight, timeout, retries, fault, mirror |
| DestinationRule | What happens once traffic is headed to this service? | subsets (label selectors), trafficPolicy: loadBalancer, connectionPool, outlierDetection, tls |
A VirtualService is a routing table evaluated top to bottom; the first matching rule wins. A DestinationRule defines named subsets, typically by a version label, and the policies applied to connections to them. Routes refer to subsets by name, so a VirtualService that names a subset the DestinationRule does not define sends traffic nowhere. Apply the DestinationRule first, then the VirtualService, and remove them in the reverse order.
Worked example: a canary for checkout v2
The shop team is releasing checkout v2. They want internal testers, who send the header x-canary: true, to always reach v2, and five percent of everyone else to reach it, with limits that stop a bad version from dragging down its callers.
apiVersion: networking.istio.io/v1
kind: DestinationRule
metadata: {name: checkout, namespace: shop}
spec:
host: checkout.shop.svc.cluster.local
trafficPolicy:
connectionPool:
tcp: {maxConnections: 200}
http: {http1MaxPendingRequests: 100, http2MaxRequests: 500}
outlierDetection:
consecutive5xxErrors: 5
interval: 10s
baseEjectionTime: 30s
maxEjectionPercent: 50
subsets:
- name: v1
labels: {version: v1}
- name: v2
labels: {version: v2}
---
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata: {name: checkout, namespace: shop}
spec:
hosts: [checkout.shop.svc.cluster.local]
http:
- match:
- headers:
x-canary: {exact: "true"}
route:
- destination: {host: checkout.shop.svc.cluster.local, subset: v2}
- route:
- destination: {host: checkout.shop.svc.cluster.local, subset: v1}
weight: 95
- destination: {host: checkout.shop.svc.cluster.local, subset: v2}
weight: 5
timeout: 2s
retries:
attempts: 2
perTryTimeout: 500ms
retryOn: connect-failure,refused-stream,unavailable,cancelled,503Weights split requests, not users: one shopper's requests can land on both versions. If v2 changes a response format or session state, route by a header or cookie that pins a user to one subset instead. Promotion is a sequence of edits, 5, 25, 50, 100, each held long enough to compare error rate and latency between subsets from the callers' metrics. Automate the comparison and the rollback; canary deployment covers the analysis side.
Timeouts and retries, and how they multiply
Two defaults matter. According to Istio's reference, a route has no timeout unless you set one, and the default retry policy, if none is specified, is two attempts on connect-failure,refused-stream,unavailable,cancelled. Those conditions cover failures where the request most likely never reached the application, which is why they are safe to retry by default. Adding 5xx codes is a decision about idempotency: retrying a non-idempotent POST that failed after doing its work can charge a customer twice.
Size the budget from the outside in. With a 2-second route timeout and a 500 ms per-try timeout, three tries take at most 1.5 seconds plus backoff, so the last retry has a chance to finish before the overall deadline. If the per-try timeout were 1 second, the third attempt would be cut off by the route timeout and add load without adding success.
Retries multiply across hops. Suppose the edge gateway, cart and checkout each make up to three attempts, and the payments service behind checkout starts timing out. One user request can then become 3 x 3 x 3 = 27 requests to payments, arriving exactly when it is least able to serve them. The rules that prevent this: retry at one layer only, usually the one closest to the failure; keep attempts low; make sure the caller's deadline shrinks as it is passed down; and treat retries as load in capacity planning. gRPC architecture explains deadline propagation, which the mesh cannot do for you.
Circuit breaking and outlier detection
Istio has two mechanisms that are often both called circuit breaking. Connection-pool limits cap concurrency from each client proxy to the service. In the example, each caller proxy holds at most 200 TCP connections and 500 concurrent HTTP/2 requests to checkout, and up to 100 HTTP/1 requests can wait for a connection. Beyond that, the proxy fails the request immediately with a 503 instead of queueing it, and Envoy's access log marks it with the response flag UO, upstream overflow. Failing fast protects the caller's threads and memory, which is the point of the circuit breaker pattern. Because the limits apply per client proxy, the load checkout sees is the limit times the number of caller pods; size it with that in mind.
Outlier detection is per host. Each proxy watches results from each checkout pod; a pod that returns five consecutive 5xx errors is ejected from that proxy's load-balancing pool for 30 seconds, and repeat offenders are ejected for longer. maxEjectionPercent: 50 guarantees that at least half the hosts stay in the pool, so a shared dependency failing behind every pod cannot eject them all and turn a partial outage into a total one. Outlier detection catches the single bad node, such as a pod with a broken disk or a stuck connection pool, that the overall error rate hides.
Locality-aware routing and failover
With locality load balancing, a proxy prefers endpoints in its own zone and region, which cuts latency and cross-zone transfer cost. Failover to another locality depends on outlier detection: Istio's locality failover documentation states that outlier detection is required for failover to work, because ejection is what tells the proxy that the local endpoints are unhealthy. A DestinationRule with localityLbSetting and no outlierDetection keeps sending traffic to a failing local zone.
trafficPolicy:
loadBalancer:
simple: LEAST_REQUEST
localityLbSetting:
enabled: true
failover:
- from: us-east1
to: us-central1
outlierDetection:
consecutive5xxErrors: 5
interval: 10s
baseEjectionTime: 30s
Mirroring and fault injection
Mirroring sends a copy of live requests to another destination and discards the copy's response. Istio documents it as best effort: the proxy does not wait for the mirrored service before returning the primary response, and mirrorPercentage controls the share copied, defaulting to all traffic. It is ideal for testing a rewrite against real request shapes. The danger is side effects: a mirrored request to a service that writes to a database or calls a payment provider performs that action for real. Mirror only to services whose writes go to an isolated store.
Fault injection adds delays or aborts to a percentage of requests, which is how you test that timeouts, retries and fallbacks actually behave as designed. One trap is documented in the API: when faults are enabled on the client side, timeouts and retries are not enabled on that route. Run fault experiments in a dedicated route or environment rather than on the route whose retry behaviour you are trying to observe.
The Gateway API form
The Kubernetes Gateway API is becoming the portable way to express the same intent. For mesh traffic, an HTTPRoute attaches to a Service as its parent instead of to a Gateway, an approach developed by the Gateway API's GAMMA initiative and supported by Istio and Linkerd. Backends are Services, so each version needs its own Service rather than a subset.
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata: {name: checkout, namespace: shop}
spec:
parentRefs:
- {group: "", kind: Service, name: checkout, port: 8080}
rules:
- backendRefs:
- {name: checkout-v1, port: 8080, weight: 95}
- {name: checkout-v2, port: 8080, weight: 5}
timeouts:
request: 2sThe portable core covers matching, weights and timeouts. Retries, outlier detection and connection limits are less uniform across implementations, so check which of them your mesh supports in the Gateway API form before migrating a policy that depends on them. Mixing both APIs for the same host is a reliable way to confuse everyone; pick one per service.
Debugging traffic rules
- Run
istioctl analyzebefore applying; it catches references to undefined subsets and conflicting rules. - Inspect what a proxy actually received with
istioctl proxy-config routes <pod>andistioctl proxy-config clusters <pod>. If a rule is not there, the proxy is not enforcing it. - Read response flags in Envoy access logs:
UOmeans a connection-pool limit rejected the request,UHno healthy upstream,UFan upstream connection failure,URXthe retry limit was exceeded. - Compare subsets with the
istio_requests_totalmetric split by destination version and response code, measured at the source.
Failure modes
- Retry storms from retries at several layers. Retry at one layer and cap attempts.
- Timeouts shorter than the work. A 2-second route timeout in front of a 3-second batch endpoint fails every call while the backend completes the work anyway. Set timeouts from measured latency percentiles per route.
- Silent bypass. A workload outside the mesh, or a call addressed to a pod IP, skips every rule.
- Over-ejection when maxEjectionPercent is high and a shared dependency fails, leaving no hosts.
- Rule order bugs. A broad match placed above a specific one swallows its traffic.
- mTLS mismatch. A DestinationRule with a TLS mode that disagrees with the server's policy yields connection resets that look like application errors; mTLS certificate rotation covers the identity side.
Trade-offs
Moving traffic policy into the mesh gives one consistent implementation across languages and lets operators change behaviour without redeploying. It costs a hop of latency per proxy, memory per sidecar or waypoint, and a second place where behaviour is defined: an engineer reading the application code cannot see that every call is retried twice. Keep business-aware resilience such as fallbacks, idempotency keys and deadlines in the application, and use the mesh for uniform transport policy. For a managed control plane and its trade-offs, see service mesh architecture in the cloud.
What to do next
- Inventory every route and record its timeout and retry policy; anything without an explicit timeout has none.
- Pick the single layer that retries on each call path and remove retries elsewhere.
- Add connection-pool limits and outlier detection with maxEjectionPercent of at most 50 to your most critical service.
- Run one canary using weights, with automated comparison of error rate and latency between subsets.
- If you use locality load balancing, confirm outlier detection is configured and test failover by failing a zone in staging.
- Inject a delay fault in staging to prove that timeouts and fallbacks fire as designed.