Most rollout guides treat a new release as a percentage of traffic: 1%, 5%, 25%, 100%. The progressive rollouts article covers that mechanism for ADK Java in detail, with deterministic buckets, pinning and kill switches, and the canary article covers how many sessions each stage needs. This article adds the second axis that agents need and ordinary services do not: autonomy. An agent that can unlock accounts, change group memberships or restart servers is dangerous in a different way from one that answers questions, and the safe way to launch it is to widen who can use it and how much it may do on its own as separate, gated steps.

You will see a rollout grid of audience rings and autonomy levels, how to assign both per session, a plugin that enforces the level on every tool call, how ADK's tool confirmation fits the middle level, the statistics for deciding a stage is done, and a worked example for an IT helpdesk agent.

Why a percentage ramp is not enough

A percentage ramp answers the question whether the new version works for a random sample of users. For an agent with write tools it leaves two problems open. First, the most expensive failures are rare actions with side effects, such as disabling the wrong account, and at 5% of traffic you may see too few of them to learn anything before going wider. Second, a random sample mixes your most forgiving users with your least, so an early incident lands on a customer instead of on your own staff.

Rings fix the second problem by ordering audiences by tolerance: internal staff first, then a friendly site, then a fraction of sites, then everyone. Autonomy levels fix the first by letting the agent make its decisions in production while a human still controls whether they take effect. Decisions are what you need evidence about, and you can collect them long before you let the agent act alone.

Two axes: rings and autonomy levels

Rollout grid: audience rings across, autonomy levels upR0 IT staffR1 one siteR2 25% sitesR3 allL3 actL2 confirmL1 suggestL0 observestage 1stage 2stage 3stage 4stage 5stage 6stage 7stage 8stage 9stage 10widermoreautonomyWithin a level the audience widens; raising autonomy restarts at ring 0. A severe incident drops a level.
The ten-stage plan for account writes in the worked example. Within a level the audience widens; raising autonomy restarts at ring 0.
LevelRead toolsWrite toolsEvidence it produces
L0 observeRunNot executed; the proposed call is loggedWould the agent have done the right thing?
L1 suggestRunNot executed; shown to a human operator as a proposalOperator agreement rate
L2 confirmRunExecuted only after the user or operator approves that callApproval rate, rejected-call reasons
L3 actRunExecuted within limits; auditedIncident rate, reversals

Levels apply per tool class, not per agent: the same session can be at L3 for password resets and L1 for group changes. Keep the level decision outside the model. The instruction can tell the model that writes may need approval, so it explains itself well, but the model must never be the component that decides whether a write happens.

Assigning ring and level per session

Resolve a session's ring and levels once, when the session is created, and store them in session state with no prefix, so they are scoped to that session and survive a move to another replica. Resolving per turn would let a conversation change autonomy halfway through when the plan is edited, which makes incidents hard to reason about. The plan itself is versioned configuration, not code.

record Stage(int ring, Map<String, Integer> levelByToolClass) {}

Session open(String userId, String siteId) {
  Stage st = rolloutPlan.stageFor(siteId);                 // ring from site, levels from plan
  Map<String, Object> state = new HashMap<>();
  state.put("rollout_plan", rolloutPlan.version());
  state.put("rollout_ring", st.ring());
  st.levelByToolClass().forEach((cls, lvl) -> state.put("autonomy_" + cls, lvl));
  return sessions.createSession("helpdesk", userId, state, null)   // Map overload, not the
      .blockingGet();                                     // deprecated ConcurrentMap one
}

Rings follow the audience: IT staff accounts map to ring 0, sites to rings by a list the release owner edits. Within a ring, the percentage mechanics from the earlier article still apply if a ring is large.

Enforcing the level in a plugin

Enforcement lives in a runner plugin, so it covers every agent in the graph and cannot be forgotten by whoever adds a new tool. The plugin classifies the tool, reads the session's level from state and either lets the call through, returns a substitute result, or asks for confirmation. A non-empty result from beforeToolCallback replaces the tool's output and the tool does not run.

final class AutonomyGatePlugin extends BasePlugin {
  private final ToolClasses classes;          // tool name -> "read" | "account_write" | "group_write"
  private final ProposalLog proposals;

  AutonomyGatePlugin(ToolClasses classes, ProposalLog proposals) {
    super("autonomy_gate");
    this.classes = classes; this.proposals = proposals;
  }

  @Override
  public Maybe<Map<String, Object>> beforeToolCallback(
      BaseTool tool, Map<String, Object> args, ToolContext ctx) {
    String cls = classes.of(tool.name());
    if (cls.equals("read")) return Maybe.empty();
    Object raw = ctx.state().get("autonomy_" + cls);
    int level = raw instanceof Number n ? n.intValue() : 0;     // missing means L0: fail safe
    switch (level) {
      case 0, 1 -> {
        proposals.record(ctx, tool.name(), args, level);       // L1 also notifies an operator
        return Maybe.just(Map.of("status", "not_executed",
            "detail", level == 0 ? "recorded for review" : "sent to an operator for action"));
      }
      case 2 -> {
        if (ctx.toolConfirmation().isEmpty()) {
          ctx.requestConfirmation("Approve " + tool.name() + " " + args + "?");
          return Maybe.just(Map.of("status", "awaiting_confirmation"));
        }
        return ctx.toolConfirmation().get().confirmed()
            ? Maybe.empty()
            : Maybe.just(Map.of("status", "rejected_by_user"));
      }
      default -> { return Maybe.empty(); }                      // L3: run, audited elsewhere
    }
  }
}

An unknown level reads as zero, so a missing state key or a bad plan fails safe. The plan-version comparison described under downgrades below is left out of the snippet for brevity; it belongs before the switch. Tool classification is an explicit registry: a new tool without a class should fail the build, not default to read.

Confirmation at L2

L2 mirrors what FunctionTool does when created with requireConfirmation set to true: with no confirmation on the context it calls requestConfirmation and returns an error map; with a rejected confirmation it returns a rejection. ADK then emits a function call named adk_request_confirmation to the client, which shows the hint and answers with a function response carrying the decision. On the next turn the tool call is retried with toolConfirmation() filled in. The built-in flag is fixed when the tool is constructed, which is why the plugin does it dynamically per session instead.

Two cautions. Check the exact response payload your ADK version expects for the confirmation answer before you build the client side, because this article does not pin it. And add an integration test that the confirmation event is emitted when the request comes from a plugin rather than the tool, since the plugin path is the less common one.

Exit rules with numbers

Each stage needs an exit rule written before it starts. Two statistics cover most cases.

Zero-incident bounds. If a stage sees n write actions and no severe incident, the 95% upper bound on the true incident rate is 1 - 0.05^(1/n), which is close to 3/n, the rule of three. To claim the rate is below one in a thousand you need about 3,000 incident-free actions; 600 actions only bound it below roughly 0.5%. This is why low-volume rings cannot prove much: plan rings so each can produce the count the next stage requires.

Agreement and approval rates. At L1 and L2, use a lower confidence bound, not the raw rate. With 1,180 approvals out of 1,220 confirmations the raw rate is 96.7% and the Wilson 95% interval is about 95.6% to 97.6%, so a rule requiring a lower bound of at least 95% passes. With 412 of 430 the raw rate is 95.8% but the lower bound is 93.5%, so the same rule says wait for more data.

Add guardrail signals that end a stage early regardless of counts: any severe incident, a rise in reversals of agent actions, and complaints tagged to the agent. Define them with the same care as the agent SLOs they protect.

Worked example: an IT helpdesk agent

An IT helpdesk agent has three tool classes: reads such as account lookup, account writes (unlock, password reset link), and group writes (add or remove group membership). Account writes follow the ten-stage plan in the grid; group writes follow a slower plan of their own.

Stages 1 to 3 run L0 for both write classes across rings 0 to 2. Reviewers label a sample of logged proposals; account-write proposals match the human decision in 97% of reviewed cases, but group-write proposals often add users to a parent group instead of the narrower one asked for. That is fixed in the tool description and the group search tool, and stage 3 repeats.

Stages 4 and 5 move account writes to L1 for rings 0 and 1, so operators see each proposal and act on it themselves. Early in stage 5 operators agree with 412 of 430 proposals: the lower bound of 93.5% misses the 95% rule, so the stage continues until agreement reaches 1,180 of 1,220, a lower bound of 95.6%. Stages 6 to 8 move account writes to L2, restarting at ring 0. By the end of stage 8, ring 2, users have approved 3,450 of 3,560 confirmations, a Wilson interval of about 96.3% to 97.4%, and the 3,450 approved unlocks ran with no severe incident, which bounds the incident rate below 0.09%. Stages 9 and 10 give account writes L3 for rings 0 and 1, with everyone later.

Group writes reach only L2, because the reviewed error rate never met the bar for autonomy. At L2 a human still approves each change, and the audit trail from the audit log article records who approved what.

Downgrades and incidents

Rolling back an autonomy level is different from rolling back a release. The code does not change; the permission does, and sessions that are already open hold the old level in their state. Give the plan a version and store it in the session, as the assignment code does, and have the gate plugin compare the session's plan version with the current one on every write. If the current plan gives a lower level for that tool class, use the lower level and record the downgrade; if it gives a higher one, keep the session's level until a new session starts. Downgrades apply at once, upgrades only to new conversations.

When a severe incident happens, the response has a fixed order. Drop the affected tool class to L1 for the whole ring through the plan, which takes effect on the next write in every session. Reverse the action where the tool supports it. Then find every action of the same shape in the audit trail since the stage began, because a single visible incident usually has quieter siblings. Only after that review does the stage restart, and its incident-free count starts again from zero, since the earlier count measured a system you have since changed.

Failure modes

  • Level decided by the model. An instruction saying "ask before writing" is not a control. Enforce in the plugin.
  • Unclassified new tools. A tool added later defaults to run. Fail the build when a tool has no class.
  • Mid-session level changes. Re-resolving per turn changes behaviour inside a conversation. Resolve once and store it.
  • Confirmation fatigue. Users approve everything after a while, so approval rate stops measuring quality. Sample approved actions for review.
  • Rings too small to exit. A ring that produces 200 actions a month cannot bound a rare incident rate. Size rings from the counts the exit rule needs.
  • Rollback that strands sessions. Lowering a level must apply to open sessions too; add a plan version check that downgrades, never upgrades, mid-session.

Trade-offs

The grid is slower than a percentage ramp: the helpdesk example took ten stages. In return, the risky part, autonomy, is decided on evidence gathered from real requests at no risk, and incidents land on staff first. Observe and suggest levels cost reviewer time and do not test latency under full load, so pair them with a percentage ramp on the serving side. For read-only agents the autonomy axis collapses and an ordinary canary is enough.

What to do next

  1. List every tool and assign it a class; make a missing class a build failure.
  2. Write the rollout plan as versioned configuration: rings, levels per class, exit rules.
  3. Store ring and levels in session state at creation, and add the gate plugin to the runner.
  4. Log L0 and L1 proposals with enough context for reviewers to label them.
  5. Decide each stage's exit counts with the rule of three and Wilson bounds before it starts.
  6. Add integration tests for each level, including a rejected confirmation.
  7. Rehearse a downgrade on an open session in staging.
Key takeaway: An agent with write tools should be rolled out along two axes: who can use it, in rings ordered by tolerance, and how much it may do alone, from observing to acting. Assign both once per session, enforce them in a runner plugin so the model never decides whether a write happens, and use confirmation as the bridge to autonomy. Exit each stage on numbers fixed in advance, and size rings so they can produce those numbers.