Dynamic configuration reload means changing how a running agent service behaves (its instruction, model, sampling settings, tool allow-list or call budget) without restarting the JVM. For ADK Java services this is tempting for good reasons: prompt fixes ship faster than code, a misbehaving tool can be switched off in seconds, and a restart drops in-flight streaming responses and anything held in memory.
It is also easy to do badly. An agent run is a multi-step conversation of model calls and tool calls; if configuration changes underneath it, step three can run with a different instruction or tool set than step one. A half-written file can take the service down, and rebuilding the wrong object can silently throw away every user's session. This article builds a reload design from first principles on ADK Java's real APIs: immutable snapshots, validation before publication, one atomic swap, per-invocation pinning, and the operational guard rails that make reload safe to use at 3 a.m. The configuration layers themselves are covered in runtime configuration for ADK Java and environment-driven config management.
What can safely change at runtime
Start by sorting every setting into one of three classes, because they need different machinery.
| Class | Examples | How it changes |
|---|---|---|
| Per-call tunable | temperature, max output tokens, maxLlmCalls | Read on each call or each run; no rebuild needed |
| Agent definition | instruction, model name, tool list, sub-agents | Rebuild the LlmAgent and the Runner that executes it |
| Process wiring | session store, credentials provider, thread pools, ports | Restart or redeploy; do not hot-reload |
The third class is where outages come from. Swapping the session service in place orphans every live session; swapping thread pools leaks threads; rotating credentials belongs to the secrets system, not to a file watcher. Make the reloadable set an explicit allow-list in the config schema and reject a change to anything else with a clear message telling the operator to redeploy.
Then make the configuration a single immutable value with a version:
public record AgentConfig(
long version,
String model,
String instruction,
float temperature,
int maxLlmCalls,
Set<String> enabledTools) {
public AgentConfig {
enabledTools = Set.copyOf(enabledTools); // defensive, immutable copy
}
public List<String> validate(Set<String> knownTools, Set<String> allowedModels) {
List<String> errors = new ArrayList<>();
if (!allowedModels.contains(model)) errors.add("unknown model " + model);
if (instruction == null || instruction.isBlank()) errors.add("instruction is empty");
if (temperature < 0f || temperature > 2f) errors.add("temperature out of range");
if (maxLlmCalls < 1 || maxLlmCalls > 200) errors.add("maxLlmCalls out of range");
for (String t : enabledTools)
if (!knownTools.contains(t)) errors.add("unknown tool " + t);
return errors;
}
}Validation is semantic, not just syntactic: a well-formed file naming a tool that does not exist, or a model your project cannot call, must be rejected before it is published, not discovered on the next user request.
Snapshots, atomic swaps and pinned invocations
ADK Java agents are built once with LlmAgent.builder() and are not meant to be mutated, which is exactly the property reload needs. Treat the agent and the runner that executes it as part of the snapshot, and publish new snapshots through an AtomicReference:
public record Snapshot(AgentConfig config, Runner runner, Instant loadedAt) {}
public final class AgentHolder {
private final AtomicReference<Snapshot> current = new AtomicReference<>();
private final BaseSessionService sessions; // created once at startup
private final BaseArtifactService artifacts;
private final Map<String, BaseTool> toolCatalog;
public Snapshot get() { return current.get(); }
Snapshot build(AgentConfig cfg) {
LlmAgent agent = LlmAgent.builder()
.name("support_agent")
.model(cfg.model())
.instruction(cfg.instruction())
.generateContentConfig(GenerateContentConfig.builder()
.temperature(cfg.temperature())
.build())
.tools(cfg.enabledTools().stream().map(toolCatalog::get).toList())
.build();
Runner runner = Runner.builder()
.agent(agent)
.appName("support") // never changes across reloads
.sessionService(sessions) // SAME instances every time
.artifactService(artifacts)
.build();
return new Snapshot(cfg, runner, Instant.now());
}
boolean publish(Snapshot expected, Snapshot next) {
return current.compareAndSet(expected, next); // loses a race instead of clobbering
}
}The comment on the session service is the most important line in this article. InMemoryRunner creates its own in-memory session service, so rebuilding an InMemoryRunner on reload discards every session in the process: users see their conversation history vanish mid-chat. Build the full Runner instead, passing the same session and artifact service instances (and memory service, if you use one) to every snapshot, and keep the app name constant so session lookups still match. Runner.builder() is the current API; older ADK Java releases expose a constructor taking the agent, app name, artifact service and session service, now deprecated, so check which your version has.
Per-invocation pinning falls out of this design. The request handler reads the reference once and uses that runner for the whole run:
Flowable<Event> handle(String userId, String sessionId, Content msg) {
Snapshot s = holder.get(); // read exactly once
RunConfig rc = RunConfig.builder().maxLlmCalls(s.config().maxLlmCalls()).build();
return s.runner().runAsync(userId, sessionId, msg, rc)
.doOnSubscribe(x -> metrics.tagVersion(s.config().version()));
}An invocation that started on version 41 finishes on version 41, even if version 42 is published while its tools are running; the next request picks up 42. The old snapshot is garbage-collected when the last invocation holding it completes. Nothing needs a lock on the request path, which costs one volatile read.
Read-through settings without a rebuild
Rebuilding is not the only option. ADK Java's Instruction is a sealed interface with two forms, Instruction.Static and Instruction.Provider, and the provider wraps a Function<ReadonlyContext, Single<String>> that is evaluated when the agent assembles a request. Per-call settings can likewise be patched in a beforeModelCallbackSync by starting from the request's config with toBuilder(). Both read the holder at call time instead of at build time:
LlmAgent agent = LlmAgent.builder()
.name("support_agent")
.model("gemini-2.5-flash")
.instruction(new Instruction.Provider(ctx ->
Single.just(holder.get().config().instruction())))
.build();This avoids rebuilding but gives up pinning: a long run makes several model calls, and each one evaluates the provider again, so a reload between calls changes the instruction mid-conversation. You can restore pinning by recording the version in session state at the start of a run and looking the instruction up by that version, but at that point you are maintaining a versioned config store. A reasonable rule: use read-through only for settings where a mid-run change is harmless, such as a logging level or a temperature nudge, and rebuild-and-swap for anything that defines what the agent is. The callback mechanics are covered in ADK Java callbacks.
Detecting changes without reading half a file
How the process notices a change matters more than it looks. Java's WatchService is the obvious choice and the most fragile one: editors write files in several steps, so you can receive an event for a truncated file; many editors replace the file by rename, so a watch on the file itself goes stale; and in Kubernetes a mounted ConfigMap is updated by atomically repointing a ..data symlink, which produces events on directories you were not watching. Kubernetes also delays ConfigMap propagation to the pod by up to the kubelet sync period plus cache time, so a reload is not instantaneous after kubectl apply.
Polling with a content hash is boring and correct. Read the whole file, hash it, and act only when the hash changes and the content parses:
void pollOnce() {
byte[] bytes;
try { bytes = Files.readAllBytes(path); }
catch (IOException e) { metrics.reloadError("read"); return; }
String hash = HexFormat.of().formatHex(sha256(bytes));
if (hash.equals(lastHash)) return;
try {
AgentConfig cfg = mapper.readValue(bytes, AgentConfig.class);
List<String> errors = cfg.validate(knownTools, allowedModels);
Snapshot prev = holder.get();
if (!errors.isEmpty()) { reject(cfg, errors); lastHash = hash; return; }
if (cfg.version() <= prev.config().version()) { reject(cfg, List.of("version not newer")); lastHash = hash; return; }
Snapshot next = holder.build(cfg); // all expensive work before the swap
if (holder.publish(prev, next)) {
lastHash = hash;
log.info("config v{} -> v{}", prev.config().version(), cfg.version());
metrics.reloadSuccess(cfg.version());
}
} catch (Exception e) {
metrics.reloadError("parse"); // keep last-known-good, retry next tick
}
}Run it on a single scheduled thread every five to thirty seconds. Recording the hash of a rejected file stops the service from logging the same rejection on every tick, while a parse failure is retried because it may be a partially written file. Requiring a strictly increasing version guards against an old file being restored by accident. The same loop works against a remote store (a database row, a parameter store, a config service) by replacing the read with a fetch that returns the payload and its version.
Operating reload in production
Reload is a deployment mechanism without the safety net of a deployment pipeline, so give it one.
- Expose the version. Publish a
config_versiongauge, success and rejection counters, and seconds since the last successful reload. Tag every trace and log line with the snapshot version so an incident can be tied to a change. - Keep last-known-good on disk. On every successful publish, write the config to a local file, and boot from it if the source is unreachable at startup. A config outage should not become a service outage.
- Stage changes. In a fleet, every replica reloading at once is a fleet-wide deploy. Roll the change through one replica or a canary cohort first, as described in canary releases for ADK Java agents, and compare outcomes per version.
- Offer a kill switch. A reloadable tool allow-list is the fastest way to stop a tool that is misbehaving. Test that removing a tool takes effect on new invocations and does not break runs whose pinned snapshot still has it.
- Audit. Log who changed what, with the diff between versions. Prompt changes alter behaviour as much as code changes and deserve the same review.
Test the machinery itself: a unit test that publishes v2 while an invocation on v1 is blocked in a tool and asserts the invocation completes on v1; a test that a malformed file leaves the current snapshot untouched; and a test that a session created before a reload is still readable after it.
Finally, decide what a rollback is. With versioned snapshots it is just another publish: restore the previous file with a new, higher version number. Do not lower the version to roll back, or the monotonic check will rightly reject it; the version counts changes, not content.
Failure modes and trade-offs
The failure modes, in rough order of how often they hurt:
- Sessions lost on reload. Rebuilding
InMemoryRunneror creating a new session service per snapshot. Share the services; assert in a test that sessions survive. - Partial files published. Acting on a watch event before the write completes. Hash whole reads, parse fully, validate, then swap.
- Torn configuration. Updating fields one at a time through several volatile variables, so a request sees the new model with the old instruction. Publish one immutable record.
- Mid-run drift. Read-through providers changing instructions between model calls of one run. Pin per invocation.
- Thundering reload. Every replica rebuilding at the same moment while also warming caches. Add jitter to the poll interval and stage rollouts.
- Silent rejection. A bad config rejected with only a log line, while operators believe it is live. Alert when the source's version and the running version differ for more than a few polls.
The trade-off is speed against control. Reload makes behaviour changes take seconds, which is exactly why it must be validated, versioned, staged and observable; without those, a restart-based deploy is the safer tool.
What to do next
- Classify every setting your agent service reads as per-call, agent-definition or process wiring, and make the reloadable set an explicit allow-list.
- Model the reloadable part as one immutable versioned record with a semantic
validate(). - Replace any
InMemoryRunnerthat might be rebuilt with aRunnerthat receives shared session and artifact services. - Publish snapshots through an
AtomicReferenceand read it exactly once per request. - Detect changes by polling with a content hash and a strictly increasing version; keep last-known-good on disk.
- Add the version gauge, reload counters and a divergence alert, then write the three machinery tests above.