Shipping a change to an agent is riskier than shipping a change to an ordinary service. A new prompt, a new model version or a new tool can pass every offline evaluation and still behave badly on real conversations, and the bad behaviour is often an action, such as a refund issued or an email sent, not just a wrong answer. Two controls make that risk manageable: progressive rollouts, which expose a change to a growing, deterministic fraction of users while metrics are watched, and kill switches, which turn a behaviour off in seconds without a deploy.

This article builds both for an ADK Java agent service: deterministic percentage buckets, a rule for which decisions to pin per turn and which to read live, a kill switch hierarchy enforced through a before-tool callback, safe behaviour when the flag service itself is down, and a ramp controller. Session-pinned canaries for whole releases, with sample size calculations, are covered in ADK Java canary deployment; this article is about the finer-grained, faster controls that sit inside one release.

Three controls, three speeds

There are three ways to change what an agent does, and they work at very different speeds. Design each for its own job and do not expect one to replace another.

ControlTypical speedGranularityUse it for
Deploy or canaryMinutes to hoursWhole releaseCode changes, dependency upgrades
Progressive rollout flagSeconds to minutesPercent of users or tenantsNew prompt, model, tool or behaviour inside a release
Kill switchSeconds, under failureGlobal, per tool, per tenantStopping harm now, while you investigate

The kill switch has the strictest requirements. It must work when the system is already unhealthy, it must not depend on the component that is failing, and it must fail safe. A rollout flag can tolerate a slower, richer evaluation path; a kill switch should be a boolean read from local memory.

Flags and kill switches around an ADK Java agent serviceFlag storerules, stages, killsBreak-glass filelocal overrideFlag client (per JVM)poll, snapshot, stalenesspollwins over storeRequest entrybucket, pin variantsnapshotRunner: stable agentcurrent prompt, model, toolsRunner: candidate agentnew behaviourbeforeToolCallbackreads kills livestablecandidateRamp controllerguardrail metricsper-variant metricsadvance or pauserollout assignment is pinned for the whole turn; kill switches are read at every tool call
One flag client per JVM keeps a snapshot. The request entry pins a variant for the turn; the tool callback checks kill switches live; the controller moves the ramp.

Deterministic percentage buckets

A percentage rollout must be deterministic, so the same user sees the same behaviour across turns and across JVMs, and monotonic, so raising 5 percent to 10 percent keeps the original 5 percent and adds new users rather than reshuffling. Hash the flag key together with a stable unit id into one of 10,000 buckets and enable the flag for buckets below the rollout threshold.

public final class Buckets {
  private Buckets() {}

  /** Stable bucket in [0, 10000) for this flag and unit (user id, tenant id). */
  public static int bucket(String flagKey, String unitId) {
    try {
      MessageDigest md = MessageDigest.getInstance("SHA-256");
      byte[] h = md.digest((flagKey + ":" + unitId).getBytes(StandardCharsets.UTF_8));
      long v = ((h[0] & 0xffL) << 24) | ((h[1] & 0xffL) << 16)
             | ((h[2] & 0xffL) << 8) | (h[3] & 0xffL);
      return (int) (v % 10_000);
    } catch (NoSuchAlgorithmException e) {
      throw new IllegalStateException(e);
    }
  }

  /** basisPoints: 500 means 5 percent. Monotonic: raising it only adds units. */
  public static boolean enabled(String flagKey, String unitId, int basisPoints) {
    return bucket(flagKey, unitId) < basisPoints;
  }
}

Salting with the flag key matters: without it, the same 5 percent of users would receive every experiment at once, and their metrics would mix all changes. Choose the unit deliberately. User id is right for conversational behaviour; tenant id is right when a business customer must see consistent behaviour across all its users, or when a contract requires opting in. Never bucket on session id for a behaviour change users can notice: a user who opens a second session would see the agent change personality.

Pin rollout decisions, never kills

Rollout decisions and kill switches need opposite consistency rules. A rollout decision should be pinned for the whole turn: if a flag flips while the agent is between model calls, half the turn would run with the old prompt and tools and half with the new, which produces transcripts that neither variant would produce and corrupts the metrics. A kill switch must not be pinned: a turn that started before the kill must still be stopped at its next tool call.

The simplest way to pin is to build one agent per variant at startup, each with its own runner, and choose the runner before the turn begins. The agent's configuration then cannot change mid-turn, because it is a different object.

public final class VariantRouter {
  private final Runner stable;
  private final Runner candidate;
  private final FlagClient flags;
  private final TurnMetrics metrics;

  public VariantRouter(Runner stable, Runner candidate, FlagClient flags, TurnMetrics metrics) {
    this.stable = stable; this.candidate = candidate; this.flags = flags; this.metrics = metrics;
  }

  public Flowable<Event> runTurn(String userId, String sessionId, Content msg) {
    FlagSnapshot snap = flags.current();                 // one read per turn
    int bp = snap.rolloutBasisPoints("refund_tool_v2");
    boolean useCandidate = !snap.killed("refund_tool_v2")
        && Buckets.enabled("refund_tool_v2", userId, bp);
    Runner runner = useCandidate ? candidate : stable;
    metrics.tagTurn(sessionId, useCandidate ? "candidate" : "stable", snap.version());
    return runner.runAsync(userId, sessionId, msg);
  }
}

Here FlagClient, FlagSnapshot and TurnMetrics are your own small classes, not ADK APIs. Both runners must share the same session service and the same app name, because sessions are looked up by app, user and session id; a plain InMemoryRunner creates its own session service, so build both runners over one shared instance. Log the snapshot version with every turn: when a metric moves, you need to know which flag state produced it. Hot-swapping configuration objects safely is covered in dynamic configuration reload.

Kill switches as a hierarchy

Organise kill switches as a hierarchy, checked from the broadest to the narrowest, so an operator can choose the smallest blast radius that stops the harm:

  • Global agent kill: the request entry returns a fixed maintenance reply without calling the model at all.
  • Variant kill: forces everyone onto the stable runner, which is what snap.killed(...) does in the router above.
  • Tool kill: the tool is skipped and the model receives an unavailable result it can explain to the user.
  • Tenant kill: the same, scoped to one customer.

Tool kills belong in a before-tool callback. ADK Java calls it with four arguments (invocation context, tool, call arguments, tool context) just before the tool body would execute. It answers with a Maybe<Map<String, Object>>: returning nothing lets the call proceed, while returning a map short-circuits the tool and hands that map to the model as if the tool had produced it. Check the exact callback types for your ADK version, as they have evolved between releases.

public final class KillSwitchGuard {
  private final FlagClient flags;

  public KillSwitchGuard(FlagClient flags) { this.flags = flags; }

  public Maybe<Map<String, Object>> beforeTool(
      InvocationContext ctx, BaseTool tool, Map<String, Object> args, ToolContext toolContext) {
    FlagSnapshot snap = flags.current();          // live read, deliberately not pinned
    String tenant = String.valueOf(ctx.session().state().getOrDefault("tenant_id", "unknown"));
    String reason = snap.killReason(tool.name(), tenant);
    if (reason == null && flags.isStale() && SIDE_EFFECTING.contains(tool.name())) {
      reason = "flag state unavailable";          // fail closed for actions
    }
    if (reason == null) {
      return Maybe.empty();                       // run the tool normally
    }
    audit.toolBlocked(ctx.invocationId(), tool.name(), tenant, reason, snap.version());
    return Maybe.just(Map.of(
        "status", "unavailable",
        "message", "This action is temporarily disabled. Do not retry; offer a human handoff."));
  }
}

As in the router, SIDE_EFFECTING (a set of tool names) and audit are your own code. Register the guard with .beforeToolCallback(guard::beforeTool) on every agent builder, stable and candidate. The returned message is written for the model: it tells it not to retry and gives it something useful to say. A tool kill that returns an error string the model does not understand often causes the model to call the tool again, or to invent a success. How tool results flow back into the loop is covered in advanced function calling.

When the flag service is down

The flag service is a dependency, and it will fail at the worst moment: incidents often involve the same network or control plane. Decide its failure behaviour in advance:

  • Keep the last known snapshot. The client polls with a short interval and replaces its snapshot atomically only after a complete, validated response. A failed poll leaves the previous snapshot in place.
  • Bound staleness. Track the age of the snapshot. Past a limit, such as five minutes, rollouts freeze at their current stage and side-effecting tools fail closed, as in the guard above, while read-only tools keep working.
  • Start safe. A JVM that boots without ever reaching the flag store uses compiled-in defaults: every rollout at 0 percent, every candidate off.
  • Break glass locally. A file or environment override on each instance takes precedence over the store, so operators can kill a tool even when the store is unreachable.

Never let a missing kill switch evaluate to enabled for something that can cause harm. The asymmetry is intentional: a false kill degrades the product for minutes; a failed kill lets damage continue.

A ramp controller with guardrails

Advancing a rollout by hand works for a few changes a month. Past that, encode the ramp as stages with guardrails, and let a controller move only forwards automatically and backwards immediately.

record Stage(int basisPoints, Duration minDuration, int minTurns) {}

List<Stage> ramp = List.of(
    new Stage(100,    Duration.ofHours(2),  300),   // 1 percent
    new Stage(500,    Duration.ofHours(6),  1500),  // 5 percent
    new Stage(2500,   Duration.ofHours(12), 5000),  // 25 percent
    new Stage(5000,   Duration.ofHours(24), 10000), // 50 percent
    new Stage(10000,  Duration.ZERO,        0));    // 100 percent

Decision decide(Stage s, VariantStats cand, VariantStats stable) {
  if (cand.toolErrorRate() > stable.toolErrorRate() * 2 + 0.005) return Decision.ROLL_BACK;
  if (cand.humanEscalationRate() > stable.humanEscalationRate() * 1.5) return Decision.PAUSE;
  if (cand.costPerTurn() > stable.costPerTurn() * 1.3) return Decision.PAUSE;
  if (cand.turns() < s.minTurns() || cand.age().compareTo(s.minDuration()) < 0) return Decision.HOLD;
  return Decision.ADVANCE;
}

Compare against the stable variant over the same time window, not against yesterday's baseline, so traffic patterns cancel out. Pick guardrails an agent change actually moves: tool error rate, human escalation rate, cost and latency per turn, and an online quality score if you have one. Roll back on harm; pause on ambiguity and page a human. The thresholds above are starting points to tune, and the minimum turn counts should come from a sample size calculation on your own baseline rates. Per-variant metric tagging is described in ADK Java observability.

Worked example: a refund tool rollout

An illustrative scenario: a support agent gains a new refund tool, gated by the refund_tool_v2 flag and bucketed by user id. The service handles about 40,000 turns a day. At 1 percent, roughly 400 turns a day reach the candidate, so the 2-hour stage lasts until the 300-turn minimum is met. At 5 percent the candidate's tool error rate climbs to 3.1 percent against 0.8 percent on stable, because one payment region rejects a currency format.

The controller returns ROLL_BACK, which does two things: it sets the rollout stage to 0 and it sets a tool kill on issueRefund, the candidate's refund tool. The two are separate on purpose: the flag key gates which variant a turn uses, while the tool name gates each call. Turns already in progress on the candidate are not abandoned: their next issueRefund call hits the kill check in the callback, receives the unavailable result, and the model offers a human handoff. Within one poll interval every JVM has the new snapshot, and new turns go to the stable runner. The fix ships in the next release, and the ramp restarts at 1 percent with the same salt, so the same users see it first.

Failure modes

  • Unsalted buckets. The same users get every experiment, and their metrics are a mixture.
  • Pinned kills. A kill read once per turn cannot stop a long turn already in progress.
  • Unpinned rollouts. A mid-turn flip mixes two prompts in one transcript and two variants in one metric.
  • Fail-open on outage. A flag client that returns defaults of true when the store is down turns every candidate on during an incident.
  • Opaque tool refusals. A blocked tool that returns a bare error string makes the model retry or invent a result.
  • Flag debt. Flags at 100 percent for months become hidden configuration; remove the flag and the losing code path after the ramp completes.

What to do next

  1. List every behaviour you might need to stop in seconds, and give each a tool or variant kill switch.
  2. Implement salted, monotonic bucketing and choose user or tenant as the unit for each flag.
  3. Build one runner per variant at startup and pick the runner once per turn.
  4. Add the kill-switch guard as a before-tool callback on every agent, with model-readable refusal messages.
  5. Define staleness limits, safe startup defaults and a local break-glass override, and test them by blocking the flag store in staging.
  6. Encode ramp stages and guardrails, roll back automatically on harm, and remove flags once they reach 100 percent.
  7. Rehearse a kill in a game day: measure the time from decision to the last JVM honouring it, and keep the deploy path ready for the permanent fix.
Key takeaway: Use deploys for code, rollout flags for behaviour, and kill switches for stopping harm. Bucket deterministically with a salted hash, pin the variant for a whole turn by choosing a runner up front, but read kill switches live in a before-tool callback that returns a refusal the model can explain. Keep the last known flags, fail closed for side-effecting tools when they go stale, and let a controller advance slowly and roll back fast.