Traffic shifting is the act of moving a share of live traffic from one backend to another: a new release, a different model, another region, a different provider. It sounds like one number, a weight, changed from 0 to 100 in steps. For stateless web requests that is nearly true. For agents it is not, because an agent's traffic is made of multi-turn sessions, long streaming responses and model calls that draw on quotas and caches tied to one side.

This article is about the weight itself: where it can live, what it actually controls for an agent, why the share of traffic a backend serves lags the weight you set, what must be true of the target before you shift, and how to drive a shift with a controller that aborts safely. Deciding whether a candidate is better, with sample sizes and statistics, is covered in ADK Java Canary Deployment, and flag-based percentage rollouts in Progressive Rollouts and Kill Switches. Here we assume the decision is made and the job is to move traffic without hurting anyone.

Where the weight lives

A weight can be enforced at four layers, and they differ in what unit they split and how fast a change takes effect:

LayerUnit splitChange takes effectGood for
DNS weighted recordsResolver lookupsMinutes to hours (TTL, caching resolvers)Region evacuation, coarse moves
Load balancer or Cloud Run revisionsHTTP requestsSeconds, for new requestsRelease shifts of the whole service
Gateway API HTTPRoute weightsHTTP requestsSeconds, for new requestsKubernetes release shifts
In-app session routerSessions (or turns)Next session after the config reloadModel, provider and prompt shifts

The first three layers see HTTP requests, not conversations. On Cloud Run, a release shift is a deploy with no traffic and a tag, a test against the tagged URL, then explicit percentages. State both revisions' percentages rather than relying on how unspecified remainder is distributed:

gcloud run deploy support-agent --image "$IMAGE" --no-traffic --tag green
# test at https://green---support-agent-<hash>.a.run.app, then shift:
gcloud run services update-traffic support-agent \
  --to-revisions support-agent-00042-abc=90,support-agent-00043-def=10
# rollback is the same command with the old weights

On Kubernetes with the Gateway API, the weights are proportional integers on the route's backends:

apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata: {name: support-agent}
spec:
  parentRefs: [{name: public-gw}]
  rules:
  - backendRefs:
    - {name: support-agent-v42, port: 8080, weight: 90}
    - {name: support-agent-v43, port: 8080, weight: 10}

A shift that changes a model, a prompt or a provider behind the same release belongs in the application, because only the application knows which session a turn belongs to.

What a weight means for an agent

Three facts about agent traffic decide what a weight really does.

Turns of one session must not alternate. A request-level split sends turn 3 to the new release and turn 4 to the old one. If the two differ in tools, state keys or model, the conversation sees contradictory behaviour, and your metrics attribute the result to neither side. Split by session: hash the session id once, or record the assignment in session state on the first turn and honour it afterwards.

Long requests outlive the change. Cloud Run's documentation states that requests in progress continue to completion and may be served by either revision during the transition. A streaming turn that started on the old side finishes there. Setting a weight to 100 does not make the old side idle; it stops sending it new requests.

Affinity overrides weights. Cloud Run's session affinity uses a cookie with a 30-day TTL and, per its documentation, takes precedence over traffic splitting. With affinity on, returning clients keep reaching the instance they had, so the served split lags the configured one. Affinity is also best effort: an instance that is terminated or saturated breaks it. It is a cache-warmth optimisation, never a substitute for storing session state durably and pinning the version in the application.

The architecture: a session router

Putting those together gives the architecture in the figure. Edge weights, if used, split requests between releases. Inside the application, a session router assigns each new session to a side by weight and sends every later turn of that session to the side it started on. A controller steps the weight and reads guard metrics per side.

A shift moves weight; the served share follows only as sessions turn overClientnew or existing sessionEdge weightsLB / Cloud Run / HTTPRouteSession router (in app)existing session -> its side; new -> weightStable sideRunner: current modelquota A, warm cachesTarget sideRunner: target modelquota B, cold caches100 - wwShift controllerstep weight, hold, read guards on served shareGuards read per-side metrics; abort sets w to 0 for new sessions at once.
Weights act on new sessions. Existing sessions finish where they started, so the served share trails the weight.
public final class SessionRouter {
  private final Runner stable, target;
  private final AtomicInteger weight = new AtomicInteger(0);   // 0..100, set by controller
  private final SideStore sides;                               // sessionId -> side, durable
  private final TurnMetrics metrics;                           // served turns per side

  public Flowable<Event> run(String userId, String sessionId, Content msg) {
    Side side = sides.computeIfAbsent(sessionId, id ->
        bucket(id) < weight.get() ? Side.TARGET : Side.STABLE);
    Runner runner = side == Side.TARGET ? target : stable;
    return runner.runAsync(userId, sessionId, msg)
        .doOnComplete(() -> metrics.turn(side, true))
        .doOnError(e -> metrics.turn(side, false));
  }

  private static int bucket(String sessionId) {      // stable across pods and restarts
    return Math.floorMod(Hashing.murmur3_32_fixed().hashUnencodedChars(sessionId).asInt(), 100);
  }

  public void setWeight(int w) { weight.set(Math.max(0, Math.min(100, w))); }
}

The two runners share one durable session service and differ in the agent they run, for example an LlmAgent whose model id comes from configuration. Store the side durably: if it were derived only from the hash, raising the weight from 25 to 50 would move every session in buckets 25 to 49 to the target mid-conversation.

Served share lags the weight

Because weights act on new sessions, the share of turns the target serves converges to the weight only as old sessions end. If session lifetimes are roughly exponential with mean L, the fraction of active sessions started after a step at time 0 is 1 - exp(-t / L). With L = 12 minutes, the served share has covered 63 percent of the gap after 12 minutes, 86 percent after 24 and 95 percent after 36.

Two rules follow. Hold each step for at least two to three mean session lifetimes, or the guards are judging a share smaller than the one you set. And make the controller read the served share, turns completed per side, not the configured weight, when it computes whether enough evidence has accumulated to step again.

Can the target carry it?

Before the first step, check that the target can carry the final weight, not just the first one. Three resources are easy to forget:

  • Model quota. A different model, region or provider usually has its own tokens-per-minute and requests-per-minute limits. Multiply peak tokens per minute by the final weight and compare with the target's quota minus headroom. If the target's quota covers 60 percent of peak, the shift must stop at about 55 percent until it is raised.
  • Cold caches. Provider-side prompt or context caches, your own response caches and connection pools start empty on the target. The first steps show higher latency and cost than steady state. Compare the target with its own trend, not only with the warm stable side, and expect the gap to shrink as the share rises.
  • Instance capacity. Pre-scale the target's minimum instances to the next step's load. A shift onto instances that must cold-start under traffic measures the autoscaler, not the release.

A guarded shift controller

A shift controller turns a schedule into weight changes with guards between them:

int[] steps = {1, 5, 25, 50, 100};
Duration hold = Duration.ofMinutes(36);               // about 3 x mean session lifetime

for (int w : steps) {
  if (w > maxWeightFromTargetQuota()) { pauseAndPage("target quota caps shift at " + w); return; }
  router.setWeight(w);
  Instant stepStart = Instant.now(), until = stepStart.plus(hold);
  // hold for the full time AND until the target has served enough turns to judge
  while (Instant.now().isBefore(until) || metrics.targetTurnsSince(stepStart) < 200) {
    Guards g = metrics.window(Duration.ofMinutes(10));    // per side, served turns
    if (g.targetTurns() >= 200 && (
          g.targetErrorRate() > g.stableErrorRate() + 0.01
       || g.targetP95FirstToken().compareTo(g.stableP95FirstToken().multipliedBy(2)) > 0
       || g.targetRateLimited() > 0.005)) {
      router.setWeight(0);                              // new sessions go stable at once
      alert("shift aborted at " + w + "%: " + g);
      return;
    }
    sleep(Duration.ofSeconds(30));
  }
}

Abort sets the weight to zero for new sessions immediately. Sessions already on the target keep running there unless the target is actually broken; if it is, flip a second switch that sends every turn to stable, accepting that some conversations change behaviour mid-way, because a broken backend is worse. The 1 percent first step exists to surface configuration errors such as a wrong model id or missing permission on real traffic before many sessions are affected. Persist the controller's state, so a restart resumes or aborts rather than starting over at the top of the schedule.

Worked example: a model migration

An illustrative scenario; every figure in it is an assumption. A support agent moves from its current model to a target model in the same release. Peak load is 300 new sessions per minute, mean session lifetime is 12 minutes and an average session uses 40,000 tokens. That is 12 million tokens per minute at peak. The target model's quota in the project is 8 million tokens per minute, so the controller's cap is 8 / 12 = 67 percent, less headroom: 60 percent until the quota increase lands.

At 1 percent, three new sessions a minute reach the target and within the first hold a missing IAM permission on one tool surfaces as 100 percent tool errors on those sessions; nothing else changed, the fix is deployed, and the step restarts. At 5 and 25 percent, time to first token on the target is 40 percent higher for the first 20 minutes, then settles to 10 percent higher as caches warm. At 50 percent the served share reaches about 49 percent after 36 minutes, as the convergence formula predicts. The controller pauses at 60 percent for the quota, and finishes the next day.

Draining the old side

When the weight reaches 100, the stable side is not done. It still holds sessions started before the last step. Keep it running until its active session count reaches zero or until a cutoff, typically a few session lifetimes, after which remaining sessions are moved to the target at their next turn. Only then remove the old runner or revision. Draining in-flight streaming turns inside a pod is covered in Zero-Downtime Deployments for ADK Java, and evacuating a whole region with weights at the DNS or load balancer layer in Multi-Region Deployment for ADK Java.

Failure modes

  • Mid-session flips. Hash-only assignment moves sessions when the weight rises. Store the side on the first turn.
  • Judging on configured weight. Guards read a target share far below the weight and declare success on too little data. Count served turns per side.
  • Target quota exhaustion. Rate-limit errors appear only at high weights. Compute the cap before the shift and include rate-limited calls in the guards.
  • Affinity hiding the split. With Cloud Run session affinity on, returning clients stick to old instances. Expect a lag or pin in the application instead.
  • Controller restart. A stateless controller restarts at step one or, worse, at 100. Persist its step and hold deadline.
  • Deleting the old side early. Streaming turns and live sessions still on it fail. Delete only after its active count reaches zero or the cutoff passes.

What to do next

  1. Decide which layer owns each kind of shift: edge weights for releases and regions, the in-app router for models, prompts and providers.
  2. Persist a session's side on its first turn and route every later turn by it.
  3. Measure mean session lifetime and set hold times to two or three times it.
  4. Emit served turns, errors, time to first token and rate-limited calls per side.
  5. Before each shift, compute the weight the target's quota allows and pre-scale its instances.
  6. Run a shift with a 1 percent first step and a persisted controller in staging, including an abort.
  7. Drain and delete the old side only after its active session count reaches zero.
Key takeaway: For agents, a traffic weight should act on new sessions, never on individual turns. The served share then trails the weight by a few session lifetimes, so hold each step that long, judge on served turns, cap the shift at what the target's quota allows, abort by zeroing the weight, and drain the old side before deleting it.