A feature flag in an ordinary service switches between two code paths you can read. In an agent, the most valuable things to flag are not code paths. They are a prompt, the set of tools the model may call, the model or its sampling settings, and the limits on a run. Changing any of those changes behaviour probabilistically, across turns, in ways a unit test will not catch, which is exactly why you want to ship them dark and turn them on for a few users first.
This article is about where flags plug into ADK for Java and how to keep a flagged agent debuggable. It maps each behaviour surface to the ADK hook that can change it, argues for evaluating every flag once per invocation and persisting the result, and wires a complete example. Percentage buckets, ramp schedules and kill-switch hierarchies are covered in progressive rollouts and kill switches, and atomic configuration swaps in dynamic configuration reload; this page builds on both rather than repeating them. API names were checked against the google-adk 1.11.0 jar.
What a flag can change in an agent
Each surface has a natural hook, and the hook decides how often the flag is consulted. The table lists them from cheapest to most invasive.
| Behaviour | ADK hook | When it runs | Notes |
|---|---|---|---|
| Instruction text | Instruction.Provider on the agent | Each model call | A function from ReadonlyContext to Single<String> |
| Which tools are offered | A BaseToolset whose getTools(ctx) filters | When a request is built | Hiding is a hint; also enforce in beforeToolCallback |
| Sampling settings | beforeModelCallback editing LlmRequest.Builder | Each model call | Temperature, max output tokens, safety settings |
| Model name, same provider | beforeModelCallback setting model(...) | Each model call | The Gemini class uses the request's model when present; other adapters may not |
| Run limits | RunConfig passed to runAsync | Once per invocation | For example setMaxLlmCalls |
| Agent tree or provider | Choose between prebuilt Runners | Once per invocation | The only safe way to switch providers or sub-agents |
Two rows deserve a warning. The model's BaseLlm instance is resolved from the agent, not from the request, so rewriting the model string in a callback renames the model within the same adapter and cannot move a call from one provider to another; build a second agent for that. And removing a tool from the declaration does not remove the model's memory of earlier turns in which it existed, so a flag that withdraws a capability must also refuse calls to it.
Evaluate once per invocation
The instruction provider and the model callback run on every model call, and one invocation can make many. If each of them asked the flag service directly, a flag flipped halfway through a turn would give step one the old prompt and step four the new one, and a percentage rollout keyed on anything unstable would give different steps different variants. Neither failure is visible in metrics; both produce baffling transcripts.
The fix is to evaluate every flag the agent uses exactly once, at the edge, before calling the Runner, and to hand the result to the run as data. Runner.runAsync has an overload that takes a Map<String, Object> state delta, which is merged into the session state for that run. Put the snapshot there under a session-scoped key. Every hook then reads ctx.state() instead of the flag service, all steps agree, and the snapshot is in the event log next to the turn it governed. Do not use the temp: prefix, which is never persisted, or user:, which would carry one session's variants into the user's other sessions. Verify once with an event-log dump that the key appears where you expect, because that persistence is what makes the turn replayable.
The flag client
Use OpenFeature, the CNCF vendor-neutral flag API, as the client so the provider behind it, whether a commercial service, an open-source server or a file in tests, can change without touching agent code. The evaluation context needs a stable targeting key; use the user id so one person sees one variant across sessions, and add the tenant and app as attributes for targeting rules.
public final class AgentFlags {
public static final String KEY = "flags"; // session-scoped state key
private final Client client = OpenFeatureAPI.getInstance().getClient("agent");
/** Evaluate every flag the agent reads, once, before the run. */
public Map<String, Object> snapshot(String tenantId, String userId) {
EvaluationContext ctx = new ImmutableContext(userId, Map.of(
"tenant", new Value(tenantId)));
Map<String, Object> f = new LinkedHashMap<>();
f.put("refund_prompt", client.getStringValue("refund_prompt", "v1", ctx));
f.put("refund_tool", client.getBooleanValue("refund_tool", false, ctx));
f.put("fast_model", client.getBooleanValue("fast_model", false, ctx));
f.put("strict_limits", client.getBooleanValue("strict_limits", false, ctx));
return Map.copyOf(f);
}
@SuppressWarnings("unchecked")
public static Map<String, Object> of(Map<String, Object> state) {
Object v = state.get(KEY);
return v instanceof Map<?, ?> m ? (Map<String, Object>) m : Map.of();
}
}Every default in that snapshot is the current production behaviour. If the provider is unreachable, OpenFeature returns the defaults and the agent behaves exactly as before the flag existed, which is the property that lets you ship flag code ahead of the feature.
Wiring the surfaces
With the snapshot in state, each surface becomes a few lines. The instruction provider picks a prompt version; the toolset offers the refund tool only when its flag is on; a plugin adjusts the model and enforces the tool gate; the handler chooses the run limits.
LlmAgent support = LlmAgent.builder()
.name("support")
.model("gemini-2.5-pro")
.instruction(new Instruction.Provider(ctx -> Single.just(
prompts.get("support", (String) AgentFlags.of(ctx.state()).getOrDefault("refund_prompt", "v1")))))
.tools(new FlaggedToolset(List.of(orderLookup, refundTool), Map.of(refundTool, "refund_tool")))
.build();
final class FlaggedToolset implements BaseToolset {
private final List<BaseTool> tools;
private final Map<BaseTool, String> gate; // tool -> flag that must be true
FlaggedToolset(List<BaseTool> tools, Map<BaseTool, String> gate) { this.tools = tools; this.gate = gate; }
@Override public Flowable<BaseTool> getTools(ReadonlyContext ctx) {
Map<String, Object> f = AgentFlags.of(ctx.state());
return Flowable.fromIterable(tools)
.filter(t -> !gate.containsKey(t) || Boolean.TRUE.equals(f.get(gate.get(t))));
}
@Override public void close() {}
}
final class FlagPlugin extends BasePlugin {
FlagPlugin() { super("flags"); }
@Override public Maybe<LlmResponse> beforeModelCallback(CallbackContext cb, LlmRequest.Builder req) {
if (Boolean.TRUE.equals(AgentFlags.of(cb.state()).get("fast_model"))) {
req.model("gemini-2.5-flash"); // same adapter, different model name
}
return Maybe.empty();
}
@Override public Maybe<Map<String, Object>> beforeToolCallback(
BaseTool tool, Map<String, Object> args, ToolContext tc) {
if (tool.name().equals("issue_refund")
&& !Boolean.TRUE.equals(AgentFlags.of(tc.state()).get("refund_tool"))) {
return Maybe.just(Map.of("status", "unavailable", "reason", "feature not enabled"));
}
return Maybe.empty();
}
}
// API handler: evaluate once, choose limits, pass the snapshot as the run's state delta.
Map<String, Object> flags = agentFlags.snapshot(tenantId, userId);
RunConfig rc = RunConfig.builder()
.setMaxLlmCalls(Boolean.TRUE.equals(flags.get("strict_limits")) ? 15 : 40).build();
return runner.runAsync(userId, sessionId, message, rc, Map.of(AgentFlags.KEY, flags));Returning a map from beforeToolCallback short-circuits the call and uses the map as the tool's result, so the model receives a clear refusal instead of an exception. The model names are examples; use the ones your project has access to. Where a flag changes the provider or the sub-agent graph, build both trees at startup and pick the Runner in the handler, exactly as the canary article does for deployment variants.
Worked example: the refund pair
A team wants a rewritten refund prompt and a new issue_refund tool for 10 percent of users. Both are flagged separately, because the prompt can ship without the tool but not the reverse: the targeting rule for refund_tool requires refund_prompt to be v2, so no user gets the tool with a prompt that never mentions it. Two days in, a user complains that the agent promised a refund that never arrived.
Support pastes the session id into the event dump. The turn's state delta shows refund_prompt=v2 and refund_tool=false: the user was in the prompt cohort but not the tool cohort, which the targeting rule should have prevented. The flag service's audit log shows someone widened the prompt flag to 25 percent without updating the dependent rule. The new prompt tells the model it can process refunds; without the tool, it promised one anyway. The fix was in two places: the prompt now conditions its refund language on the tool being present, and the dependency between the flags moved from a convention into a validation check in the deployment pipeline. Because the snapshot was persisted, the team replayed the exact turn with both combinations in a test before re-enabling the ramp.
Testing a flagged agent
Flags multiply the agent's configurations. Four boolean flags are sixteen agents, and you will not run evaluation suites on all of them. Test each flag on and off with the others at their defaults, test known interacting pairs explicitly, and declare dependencies, like the refund pair above, so impossible combinations are rejected rather than tested. Because every hook reads the snapshot from state, a test can inject a snapshot directly through runAsync's state delta with no flag provider at all. Keep a small set of golden conversations per variant and compare tool-call sequences, not exact text.
Operating flags in production
A flag you cannot measure is a guess. Because the snapshot sits in session state, a plugin can attach the variant names to every metric it emits: tool-call counts, model calls per invocation, escalation to a human, user feedback and cost per turn. Compare cohorts on those, not on a single quality score, and keep the comparison window long enough to cover weekly traffic shapes. Watch the flag service itself too: evaluation latency adds directly to every turn, so set a short client timeout and let the defaults apply when it is exceeded.
Treat flag changes as production changes. Require a reason and a reviewer for changes to flags that alter prompts or tools, keep the provider's audit log, and annotate dashboards when a flag moves so a latency step can be matched to it in seconds. When a flag reaches full rollout, delete the code path, the prompt variant and the flag in that order.
Failure modes
- Mid-invocation flips. A hook that queries the flag service directly sees two values in one turn. Read only the snapshot.
- Unstable targeting keys. Bucketing on session id gives one person different behaviour every conversation, which reads as a flaky product. Use the user id.
- Prompt-cache misses. Each instruction variant is a different prefix, so context caching hit rates drop during a split and costs rise. Budget for it.
- Hidden but callable tools. A withdrawn tool that appears in history can still be requested; the
beforeToolCallbackgate is the real control. - Snapshot in the wrong scope.
temp:loses it,user:leaks it across sessions. - Flag debt. A flag at 100 percent for months is dead code plus a configuration risk. Give every flag an owner and a removal date when it is created.
Trade-offs
Flags, configuration and deployments overlap. A flag is right when you need per-user targeting, a fast off switch and an experiment readout. Versioned configuration is better for per-tenant settings that are not experiments. A canary deployment is better when the change is code, a new dependency or a different agent graph. Evaluating at the edge costs one flag round trip per turn and a little state per event, and in exchange makes every turn explainable and replayable, which is the trade worth making for anything that changes what the model is told or allowed to do. Tool exposure decisions also belong in the tool registry so a flag never offers a tool the agent's profile forbids.
What to do next
- List the behaviours you change most often: prompts, tools, model, limits.
- Map each to its hook using the surfaces table, and build a second Runner for anything provider- or graph-level.
- Add an OpenFeature client and a snapshot function whose defaults equal today's behaviour.
- Pass the snapshot through
runAsync's state delta under a session-scoped key and confirm it in an event dump. - Make every hook read flags from
ctx.state()only. - Gate withdrawn tools in
beforeToolCallbackas well as in the toolset. - Encode flag dependencies as validation, not as a wiki convention.
- Give each flag an owner and a removal date, and delete it after full rollout.