A service mesh puts a proxy, or a per-node tunnel, in front of every workload so that identity, encryption, authorization, retries, timeouts and telemetry become platform configuration rather than code in each service. Istio, Linkerd and Cilium's mesh all do this in different ways. For LLM systems the mesh is attractive for the usual reasons, and it also has defaults designed for short, cheap, idempotent RPCs that are wrong for generation calls that stream for a minute, cost GPU time and may trigger tool side effects.
This article covers what a mesh should own in an LLM stack, the policies to write for model servers, vector stores, agents and tools, how to tune traffic management for streaming inference, how egress control works, where model-aware routing such as the Kubernetes Gateway API Inference Extension fits, and the failure modes that show up in production. Examples use Istio because its APIs are widely deployed; the concepts carry to other meshes.
Why LLM traffic needs different mesh settings
Three properties make LLM traffic different from typical microservice calls. First, requests are long and streamed: a chat completion holds a connection open while tokens arrive over server-sent events, and reasoning models may think for a long time before the first byte. Second, requests are expensive and uneven: one call can occupy a GPU for tens of seconds, and two requests to the same endpoint can differ in cost by a factor of a hundred depending on prompt and output length. Third, requests can act: an agent's call to a tool service may create a ticket, send an email or move money, so replaying it is not harmless.
The security case is also sharper. Prompt injection turns the model into a confused deputy that will try to call whatever it can reach, so the network must refuse calls that the caller's identity should never make, regardless of what the text said. A mesh gives you that refusal at layer 7 with cryptographic identity, which plain Kubernetes NetworkPolicy, working on IPs and ports, cannot. What a mesh cannot do is understand tokens, prompts or tool semantics; content controls stay in the LLM gateway and the tools themselves.
Identity and authorization at every receiver
Start with strict mutual TLS everywhere. In Istio each workload receives an X.509 certificate encoding a SPIFFE identity derived from its namespace and service account, and a mesh-wide PeerAuthentication in STRICT mode rejects plaintext. Permissive mode is a migration aid; leaving it on means any pod without a proxy can still talk plaintext to your model servers. For the identity model behind this, see workload identity for LLM systems; for transport details, encryption in transit.
Then write ALLOW policies at every receiver. In Istio, once any ALLOW policy selects a workload, requests that match none of its rules are denied, so a policy is also a default deny. The model server should accept chat traffic only from the LLM gateway's identity, plus metrics scraping from the monitoring identity; tool services should accept only the agent runtime, and only the methods and paths each tool needs.
apiVersion: security.istio.io/v1
kind: PeerAuthentication
metadata: {name: default, namespace: istio-system}
spec:
mtls: {mode: STRICT}
---
apiVersion: security.istio.io/v1
kind: AuthorizationPolicy
metadata: {name: vllm-callers, namespace: inference}
spec:
selector: {matchLabels: {app: vllm-llama}}
action: ALLOW
rules:
- from: [{source: {principals: ["cluster.local/ns/llm-gateway/sa/gateway"]}}]
to: [{operation: {methods: ["POST"], paths: ["/v1/chat/completions", "/v1/completions"]}}]
- from: [{source: {principals: ["cluster.local/ns/monitoring/sa/prometheus"]}}]
to: [{operation: {methods: ["GET"], paths: ["/metrics"]}}]In Istio's ambient mode the split matters. The per-node ztunnel enforces layer 4 rules only; HTTP methods and paths need a waypoint proxy, and the policy must attach to it with targetRefs instead of a selector. Istio documents that an L7 policy left on a workload selector, and so enforced by ztunnel, fails safe by becoming a DENY policy. That is secure, but it looks like an outage, so test ambient policies in staging with real traffic before rollout.
Retries, timeouts and outlier detection for generation
Traffic management is where mesh defaults hurt LLM workloads most. Istio retries failed HTTP requests twice by default for certain failure conditions. For a generation call that failed after 40 seconds of GPU work, each retry burns the same work again, and under overload retries amplify load on the saturated pool. For a tool call, a retry after a timeout can execute the side effect twice. Set retries explicitly per route: zero for generation and side-effecting tools, a small number with idempotency keys for reads.
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata: {name: vllm-llama, namespace: inference}
spec:
hosts: ["vllm-llama.inference.svc.cluster.local"]
http:
- route: [{destination: {host: vllm-llama.inference.svc.cluster.local}}]
timeout: 300s # whole stream, sized to max_tokens at worst-case speed
retries: {attempts: 0} # the LLM gateway decides whether a retry is worth itTimeouts need the same care. A route timeout covers the whole response, including a long stream, so size it from your maximum output tokens and worst observed decode speed, then check every idle timeout between client and model: cloud load balancers, the ingress, and proxy stream idle settings can each cut a quiet stream while a reasoning model is thinking. Sending periodic SSE comment lines keeps intermediaries from treating it as idle.
Outlier detection ejects endpoints that return consecutive 5xx errors. If your model server sheds load with 503s when its queue is full, the mesh will eject the busiest pods, pushing their traffic onto the others until they also shed and are ejected. Either have the server signal overload with 429, which consecutive-5xx detection ignores, or cap ejection with a low maxEjectionPercent.
Model-aware routing and egress
Generic mesh load balancing counts requests or connections; it cannot see queue depth, KV-cache use or which pod already holds a prompt prefix in cache. For GPU serving those signals decide latency. The Kubernetes Gateway API Inference Extension addresses this: an InferencePool resource groups model-server pods and references an endpoint picker service, and the gateway's Envoy calls that picker through the external processing filter to choose a backend from live model-server metrics. Istio supports it on its gateway when installed with the inference extension feature enabled; check the current Istio and extension documentation for resource versions, which have changed across releases.
The clean division is: the inference gateway picks which replica serves a request; the mesh secures the hop to that replica and enforces who may call it. Do not let model-aware routing bypass mesh authorization by addressing pods directly on a port the policy does not cover.
Egress is the other half of routing. Agents and gateways that call hosted model APIs or the public web should leave through an egress gateway. Set the mesh's outbound traffic policy to REGISTRY_ONLY so only hosts declared with a ServiceEntry are reachable, then route those through the egress gateway, where you can log and authorize per source identity. A sidecar cannot stop a pod that has no sidecar, runs on the host network or has the privileges to rewrite iptables, so back egress control with NetworkPolicy and admission rules; egress control for LLM systems goes deeper.
apiVersion: networking.istio.io/v1
kind: ServiceEntry
metadata: {name: anthropic-api, namespace: llm-gateway}
spec:
hosts: ["api.anthropic.com"]
location: MESH_EXTERNAL
resolution: DNS
ports: [{number: 443, name: tls, protocol: TLS}]
Rolling it out without an outage
Rolling a mesh onto a running LLM platform is safest in a fixed order. First inject proxies, or enrol namespaces in ambient mode, with PERMISSIVE mTLS and no authorization policies, and confirm that streaming, large prompts and GPU metrics behave exactly as before. Second, switch each namespace to STRICT once telemetry shows no plaintext callers. Third, add authorization policies in dry-run form first: Istio's istio.io/dry-run: "true" annotation, an experimental feature, records what a policy would have decided without enforcing it, which lets you compare the caller identities you expect against the ones you actually see. Fourth, enforce ALLOW policies one receiver at a time, starting with tool services that have side effects, because those carry the most risk from an injected agent.
Treat policies as code. Keep them in the same repository as the services, review them like firewall changes, and generate a reachability matrix in CI that lists, for each identity, which services and paths it may call. A policy change that widens the matrix should need a security reviewer; a change that narrows it should need the owning team. During incidents, the access log line for a denied request, carrying source principal, destination, path and response flag, is usually the fastest way to tell a policy problem from an application bug.
Worked example: an injected refund agent
Consider a support assistant whose agent can read the knowledge base, search tickets and issue refunds through a billing tool. A red-team prompt hidden in a customer email tells the agent to call the billing service's admin endpoint and to fetch an attacker URL with the conversation attached. With the mesh in place, the agent runtime's identity is allowed to POST /refunds on billing and nothing else, so /admin/export returns 403 from the billing proxy and the denial appears in the access log with both identities. The attacker URL is not in the registry, so the outbound call fails at the sidecar and the egress gateway never sees it.
The same week, latency alarms fire during a traffic spike. Traces show each slow chat request reaching the model pool three times: the default retry policy was still active on a route created before the platform team's template. Setting retries to zero on generation routes cut GPU load by roughly a third during the next spike, and the refund route was given an idempotency key requirement so a future retry policy cannot double-pay.
Failure modes and trade-offs
Operational failure modes cluster around a few causes. Startup races: an application that calls the model server before its sidecar is ready fails its first requests; use native sidecars or Istio's holdApplicationUntilProxyStarts option. Forgotten scrapers and probes: a strict ALLOW policy silently blocks Prometheus, so GPU dashboards go blank on rollout day. Large bodies: long contexts and base64 images make requests large, and buffering filters can reject or slow them; test with your largest real prompt. Telemetry blind spots: the mesh reports status, bytes and total duration, which for streams says little; time to first token, tokens per second and cost must come from the gateway or model server.
The trade-offs: a sidecar adds memory and a little latency per hop, negligible against generation time but noticeable on chatty retrieval calls; ambient mode reduces per-pod cost at the price of a separate waypoint for L7 policy. A mesh adds an operational system to upgrade and debug. It earns its cost when you have many services and need identity-based policy; for a single model server behind one gateway, NetworkPolicy plus TLS may be enough. See zero trust for LLM systems for where the mesh fits in the wider design.