Most rollout guides treat a new release as a percentage of traffic: 1%, 5%, 25%, 100%. The progressive rollouts article covers that mechanism for ADK Java in detail, with deterministic buckets, pinning and kill switches, and the canary article covers how many sessions each stage needs. This article adds the second axis that agents need and ordinary services do not: autonomy. An agent that can unlock accounts, change group memberships or restart servers is dangerous in a different way from one that answers questions, and the safe way to launch it is to widen who can use it and how much it may do on its own as separate, gated steps.
You will see a rollout grid of audience rings and autonomy levels, how to assign both per session, a plugin that enforces the level on every tool call, how ADK's tool confirmation fits the middle level, the statistics for deciding a stage is done, and a worked example for an IT helpdesk agent.
Why a percentage ramp is not enough
A percentage ramp answers the question whether the new version works for a random sample of users. For an agent with write tools it leaves two problems open. First, the most expensive failures are rare actions with side effects, such as disabling the wrong account, and at 5% of traffic you may see too few of them to learn anything before going wider. Second, a random sample mixes your most forgiving users with your least, so an early incident lands on a customer instead of on your own staff.
Rings fix the second problem by ordering audiences by tolerance: internal staff first, then a friendly site, then a fraction of sites, then everyone. Autonomy levels fix the first by letting the agent make its decisions in production while a human still controls whether they take effect. Decisions are what you need evidence about, and you can collect them long before you let the agent act alone.
Two axes: rings and autonomy levels
| Level | Read tools | Write tools | Evidence it produces |
|---|---|---|---|
| L0 observe | Run | Not executed; the proposed call is logged | Would the agent have done the right thing? |
| L1 suggest | Run | Not executed; shown to a human operator as a proposal | Operator agreement rate |
| L2 confirm | Run | Executed only after the user or operator approves that call | Approval rate, rejected-call reasons |
| L3 act | Run | Executed within limits; audited | Incident rate, reversals |
Levels apply per tool class, not per agent: the same session can be at L3 for password resets and L1 for group changes. Keep the level decision outside the model. The instruction can tell the model that writes may need approval, so it explains itself well, but the model must never be the component that decides whether a write happens.
Assigning ring and level per session
Resolve a session's ring and levels once, when the session is created, and store them in session state with no prefix, so they are scoped to that session and survive a move to another replica. Resolving per turn would let a conversation change autonomy halfway through when the plan is edited, which makes incidents hard to reason about. The plan itself is versioned configuration, not code.
record Stage(int ring, Map<String, Integer> levelByToolClass) {}
Session open(String userId, String siteId) {
Stage st = rolloutPlan.stageFor(siteId); // ring from site, levels from plan
Map<String, Object> state = new HashMap<>();
state.put("rollout_plan", rolloutPlan.version());
state.put("rollout_ring", st.ring());
st.levelByToolClass().forEach((cls, lvl) -> state.put("autonomy_" + cls, lvl));
return sessions.createSession("helpdesk", userId, state, null) // Map overload, not the
.blockingGet(); // deprecated ConcurrentMap one
}Rings follow the audience: IT staff accounts map to ring 0, sites to rings by a list the release owner edits. Within a ring, the percentage mechanics from the earlier article still apply if a ring is large.
Enforcing the level in a plugin
Enforcement lives in a runner plugin, so it covers every agent in the graph and cannot be forgotten by whoever adds a new tool. The plugin classifies the tool, reads the session's level from state and either lets the call through, returns a substitute result, or asks for confirmation. A non-empty result from beforeToolCallback replaces the tool's output and the tool does not run.
final class AutonomyGatePlugin extends BasePlugin {
private final ToolClasses classes; // tool name -> "read" | "account_write" | "group_write"
private final ProposalLog proposals;
AutonomyGatePlugin(ToolClasses classes, ProposalLog proposals) {
super("autonomy_gate");
this.classes = classes; this.proposals = proposals;
}
@Override
public Maybe<Map<String, Object>> beforeToolCallback(
BaseTool tool, Map<String, Object> args, ToolContext ctx) {
String cls = classes.of(tool.name());
if (cls.equals("read")) return Maybe.empty();
Object raw = ctx.state().get("autonomy_" + cls);
int level = raw instanceof Number n ? n.intValue() : 0; // missing means L0: fail safe
switch (level) {
case 0, 1 -> {
proposals.record(ctx, tool.name(), args, level); // L1 also notifies an operator
return Maybe.just(Map.of("status", "not_executed",
"detail", level == 0 ? "recorded for review" : "sent to an operator for action"));
}
case 2 -> {
if (ctx.toolConfirmation().isEmpty()) {
ctx.requestConfirmation("Approve " + tool.name() + " " + args + "?");
return Maybe.just(Map.of("status", "awaiting_confirmation"));
}
return ctx.toolConfirmation().get().confirmed()
? Maybe.empty()
: Maybe.just(Map.of("status", "rejected_by_user"));
}
default -> { return Maybe.empty(); } // L3: run, audited elsewhere
}
}
}An unknown level reads as zero, so a missing state key or a bad plan fails safe. The plan-version comparison described under downgrades below is left out of the snippet for brevity; it belongs before the switch. Tool classification is an explicit registry: a new tool without a class should fail the build, not default to read.
Confirmation at L2
L2 mirrors what FunctionTool does when created with requireConfirmation set to true: with no confirmation on the context it calls requestConfirmation and returns an error map; with a rejected confirmation it returns a rejection. ADK then emits a function call named adk_request_confirmation to the client, which shows the hint and answers with a function response carrying the decision. On the next turn the tool call is retried with toolConfirmation() filled in. The built-in flag is fixed when the tool is constructed, which is why the plugin does it dynamically per session instead.
Two cautions. Check the exact response payload your ADK version expects for the confirmation answer before you build the client side, because this article does not pin it. And add an integration test that the confirmation event is emitted when the request comes from a plugin rather than the tool, since the plugin path is the less common one.
Exit rules with numbers
Each stage needs an exit rule written before it starts. Two statistics cover most cases.
Zero-incident bounds. If a stage sees n write actions and no severe incident, the 95% upper bound on the true incident rate is 1 - 0.05^(1/n), which is close to 3/n, the rule of three. To claim the rate is below one in a thousand you need about 3,000 incident-free actions; 600 actions only bound it below roughly 0.5%. This is why low-volume rings cannot prove much: plan rings so each can produce the count the next stage requires.
Agreement and approval rates. At L1 and L2, use a lower confidence bound, not the raw rate. With 1,180 approvals out of 1,220 confirmations the raw rate is 96.7% and the Wilson 95% interval is about 95.6% to 97.6%, so a rule requiring a lower bound of at least 95% passes. With 412 of 430 the raw rate is 95.8% but the lower bound is 93.5%, so the same rule says wait for more data.
Add guardrail signals that end a stage early regardless of counts: any severe incident, a rise in reversals of agent actions, and complaints tagged to the agent. Define them with the same care as the agent SLOs they protect.
Worked example: an IT helpdesk agent
An IT helpdesk agent has three tool classes: reads such as account lookup, account writes (unlock, password reset link), and group writes (add or remove group membership). Account writes follow the ten-stage plan in the grid; group writes follow a slower plan of their own.
Stages 1 to 3 run L0 for both write classes across rings 0 to 2. Reviewers label a sample of logged proposals; account-write proposals match the human decision in 97% of reviewed cases, but group-write proposals often add users to a parent group instead of the narrower one asked for. That is fixed in the tool description and the group search tool, and stage 3 repeats.
Stages 4 and 5 move account writes to L1 for rings 0 and 1, so operators see each proposal and act on it themselves. Early in stage 5 operators agree with 412 of 430 proposals: the lower bound of 93.5% misses the 95% rule, so the stage continues until agreement reaches 1,180 of 1,220, a lower bound of 95.6%. Stages 6 to 8 move account writes to L2, restarting at ring 0. By the end of stage 8, ring 2, users have approved 3,450 of 3,560 confirmations, a Wilson interval of about 96.3% to 97.4%, and the 3,450 approved unlocks ran with no severe incident, which bounds the incident rate below 0.09%. Stages 9 and 10 give account writes L3 for rings 0 and 1, with everyone later.
Group writes reach only L2, because the reviewed error rate never met the bar for autonomy. At L2 a human still approves each change, and the audit trail from the audit log article records who approved what.
Downgrades and incidents
Rolling back an autonomy level is different from rolling back a release. The code does not change; the permission does, and sessions that are already open hold the old level in their state. Give the plan a version and store it in the session, as the assignment code does, and have the gate plugin compare the session's plan version with the current one on every write. If the current plan gives a lower level for that tool class, use the lower level and record the downgrade; if it gives a higher one, keep the session's level until a new session starts. Downgrades apply at once, upgrades only to new conversations.
When a severe incident happens, the response has a fixed order. Drop the affected tool class to L1 for the whole ring through the plan, which takes effect on the next write in every session. Reverse the action where the tool supports it. Then find every action of the same shape in the audit trail since the stage began, because a single visible incident usually has quieter siblings. Only after that review does the stage restart, and its incident-free count starts again from zero, since the earlier count measured a system you have since changed.
Failure modes
- Level decided by the model. An instruction saying "ask before writing" is not a control. Enforce in the plugin.
- Unclassified new tools. A tool added later defaults to run. Fail the build when a tool has no class.
- Mid-session level changes. Re-resolving per turn changes behaviour inside a conversation. Resolve once and store it.
- Confirmation fatigue. Users approve everything after a while, so approval rate stops measuring quality. Sample approved actions for review.
- Rings too small to exit. A ring that produces 200 actions a month cannot bound a rare incident rate. Size rings from the counts the exit rule needs.
- Rollback that strands sessions. Lowering a level must apply to open sessions too; add a plan version check that downgrades, never upgrades, mid-session.
Trade-offs
The grid is slower than a percentage ramp: the helpdesk example took ten stages. In return, the risky part, autonomy, is decided on evidence gathered from real requests at no risk, and incidents land on staff first. Observe and suggest levels cost reviewer time and do not test latency under full load, so pair them with a percentage ramp on the serving side. For read-only agents the autonomy axis collapses and an ordinary canary is enough.
What to do next
- List every tool and assign it a class; make a missing class a build failure.
- Write the rollout plan as versioned configuration: rings, levels per class, exit rules.
- Store ring and levels in session state at creation, and add the gate plugin to the runner.
- Log L0 and L1 proposals with enough context for reviewers to label them.
- Decide each stage's exit counts with the rule of three and Wilson bounds before it starts.
- Add integration tests for each level, including a rejected confirmation.
- Rehearse a downgrade on an open session in staging.