An agent service built on the Agent Development Kit for Java can start, pass its health check and take traffic while being unable to answer a single request. The reason is that most of what can be wrong with its configuration is checked late. In current source, LlmAgent's builder rejects an empty agent name and little else; the model name is stored as a string and only resolved to a client through LlmRegistry.getLlm when the agent first needs it. A typo in an environment variable therefore surfaces as IllegalArgumentException: Unsupported model: ... on the first customer request, not at deploy time.
Configuration validation at startup is the discipline of finding every such problem before the process says it is ready. This article builds a validation phase that runs between reading configuration and building the Runner: typed parsing, cross-field rules, checks on the agent tree itself, and optional time-boxed live probes. It reports all problems at once with secrets redacted, gates readiness, and can run on its own as a validate-only command in CI. The sequence of boot steps is covered in ADK Java Runtime in Depth, and where configuration comes from is covered in Environment and Configuration Management; here we focus on the gate itself.
Why validate at startup
A configuration error costs least when it is caught before the old version stops serving. Caught in a pull request, it costs a red build. Caught at process start, a rolling deployment halts with the previous replicas still running. Caught on the first request, users see errors and the rollout may already have replaced every healthy pod, because each new pod reported ready. Startup validation moves errors as far left as their nature allows.
Three properties make it worth more than scattered Objects.requireNonNull calls. It is exhaustive: each tier reports all of its problems at once. It is ordered: checks that need no network run first and can run anywhere, including CI. And it is authoritative: the runner is not built and readiness does not flip until validation passes.
Typed parsing that collects every error
Start by turning raw strings into a typed, immutable configuration object, and collect errors instead of throwing on the first one. A small Problems accumulator does the work. Each check records the configuration key and a human message, and sensitive values are never put into the message.
public final class Problems {
public record Problem(String key, String message) {}
private final List<Problem> items = new ArrayList<>();
public void add(String key, String message) { items.add(new Problem(key, message)); }
public boolean isEmpty() { return items.isEmpty(); }
public List<Problem> items() { return List.copyOf(items); }
}
public record AgentConfig(
String model, // e.g. gemini-2.5-flash
boolean useVertex, // GOOGLE_GENAI_USE_VERTEXAI
String project, // GOOGLE_CLOUD_PROJECT (Vertex only)
String location, // GOOGLE_CLOUD_LOCATION (Vertex only)
Secret apiKey, // GOOGLE_API_KEY (API-key mode only)
Duration toolTimeout,
int maxTurns,
URI ordersBaseUrl) {}
public final class ConfigLoader {
public static AgentConfig load(Map<String, String> env, Problems p) {
String model = required(env, "AGENT_MODEL", p);
boolean vertex = bool(env, "GOOGLE_GENAI_USE_VERTEXAI", false, p);
Duration timeout = duration(env, "TOOL_TIMEOUT", Duration.ofSeconds(10), p);
int maxTurns = intInRange(env, "AGENT_MAX_TURNS", 8, 1, 50, p);
URI orders = uri(env, "ORDERS_BASE_URL", p);
Secret key = Secret.ofNullable(env.get("GOOGLE_API_KEY"));
return new AgentConfig(model, vertex, env.get("GOOGLE_CLOUD_PROJECT"),
env.get("GOOGLE_CLOUD_LOCATION"), key, timeout, maxTurns, orders);
}
static boolean bool(Map<String, String> env, String k, boolean dflt, Problems p) {
String raw = env.get(k);
if (raw == null || raw.isBlank()) return dflt;
return switch (raw.trim().toLowerCase(Locale.ROOT)) {
case "true", "1" -> true;
case "false", "0" -> false;
default -> { p.add(k, "must be true/false/1/0, got '" + raw + "'"); yield dflt; }
};
}
// required(), duration(), uri() and intInRange() follow the same pattern.
}Two details matter. Parsing functions return a default after recording a problem, so later checks still run and the report is complete. And Secret is a tiny wrapper of your own whose toString() returns ***, so even an accidental log statement or exception message cannot leak the key. Notice also the strict boolean parser: Boolean.parseBoolean("ture") returns false without complaint, which is exactly the silent misconfiguration this phase exists to catch.
Cross-field rules
Many real mistakes are not wrong values but wrong combinations. The google-genai client underneath ADK's Gemini support reads GOOGLE_GENAI_USE_VERTEXAI to choose between the Gemini API with an API key and Vertex AI with a project and location. Each mode makes other fields required, and a field that belongs to the other mode is a sign of a half-migrated deployment.
static void crossField(AgentConfig c, Problems p) {
if (c.useVertex()) {
if (isBlank(c.project())) p.add("GOOGLE_CLOUD_PROJECT", "required when GOOGLE_GENAI_USE_VERTEXAI=true");
if (isBlank(c.location())) p.add("GOOGLE_CLOUD_LOCATION", "required when GOOGLE_GENAI_USE_VERTEXAI=true");
if (c.apiKey().isPresent()) p.add("GOOGLE_API_KEY", "set but ignored in Vertex mode; remove it");
} else if (c.apiKey().isEmpty()) {
p.add("GOOGLE_API_KEY", "required when GOOGLE_GENAI_USE_VERTEXAI is false or unset");
}
if (c.ordersBaseUrl() != null && !"https".equals(c.ordersBaseUrl().getScheme())
&& !"localhost".equals(c.ordersBaseUrl().getHost())) {
p.add("ORDERS_BASE_URL", "must use https outside localhost");
}
}Whether a stray key in Vertex mode is an error or a warning is a policy choice; an error stops the confusing case where someone rotates an unused key and sees no effect. Every rule should name the key to change.
Validating the agent tree
The third tier validates what you are about to build rather than the raw values. It needs no network, so it runs in unit tests and CI as well as at boot.
- Model name resolves. Calling
LlmRegistry.getLlm(model)runs the same pattern match the first request would. In current source the registry ships patterns forgemini-.*,gemma-.*andapigee/.*; register any custom factory withregisterLlmbefore this check. CatchIllegalArgumentExceptionand record it as a problem onAGENT_MODEL. Add your own allowlist on top, because a name can matchgemini-.*and still not exist. - Agent names are unique. Walk the tree through
subAgents()and reject duplicates; transfers between agents are by name, so a collision misroutes silently. - Tools are present.
FunctionTool.create(Class, String)finds a method by reflection; build the tool list in a factory method and assert the expected tool names, so a renamed method fails here. - Instruction placeholders are seeded. Placeholders such as
{user_tier}are filled from session state per request. Extract them with a regular expression and check each against the set of keys your session-creation code promises to seed.
static void validateTree(BaseAgent root, Set<String> seededKeys,
Map<String, String> instructions, Problems p) {
Set<String> seen = new HashSet<>();
Deque<BaseAgent> stack = new ArrayDeque<>(List.of(root));
while (!stack.isEmpty()) {
BaseAgent a = stack.pop();
if (!seen.add(a.name())) p.add("agent:" + a.name(), "duplicate agent name");
stack.addAll(a.subAgents());
}
Pattern ph = Pattern.compile("\\{([A-Za-z_][A-Za-z0-9_]*)\\}");
instructions.forEach((agent, text) -> {
Matcher m = ph.matcher(text);
while (m.find()) {
if (!seededKeys.contains(m.group(1))) {
p.add("agent:" + agent, "instruction uses {" + m.group(1) + "} but no session seeds it");
}
}
});
}
static void validateModel(String model, Set<String> allowed, Problems p) {
if (!allowed.contains(model)) p.add("AGENT_MODEL", "not in allowlist " + allowed);
try { LlmRegistry.getLlm(model); }
catch (IllegalArgumentException e) { p.add("AGENT_MODEL", e.getMessage()); }
}
Bounded live probes
Some things can only be checked against the real world: the credential is accepted, the tool backend answers, the session store is reachable. These probes are the fourth tier and need care, because a probe that blocks forever or fails on a transient blip turns validation into an outage of its own.
Give every probe a hard deadline, run them concurrently, and classify each as required or advisory. A required probe failing keeps readiness false; an advisory one only logs a warning. The session store is usually required, since every turn writes to it. A model test call is usually advisory: it costs money, and a provider blip should not block an incident fix. Do not assume getLlm proves anything about credentials; it constructs a client, and whether that touches the network is an implementation detail you should not depend on.
record Probe(String name, boolean required, Callable<Void> check) {}
static void runProbes(List<Probe> probes, Duration deadline, Problems p, Logger log) {
ExecutorService pool = Executors.newVirtualThreadPerTaskExecutor();
try {
Map<Probe, Future<Void>> futures = new LinkedHashMap<>();
for (Probe pr : probes) futures.put(pr, pool.submit(pr.check()));
long end = System.nanoTime() + deadline.toNanos();
for (var e : futures.entrySet()) {
try {
e.getValue().get(Math.max(0, end - System.nanoTime()), TimeUnit.NANOSECONDS);
} catch (Exception ex) {
String msg = "probe failed: " + ex.getClass().getSimpleName();
if (e.getKey().required()) p.add("probe:" + e.getKey().name(), msg);
else log.warn("advisory probe {} failed: {}", e.getKey().name(), msg);
e.getValue().cancel(true);
}
}
} finally {
pool.shutdownNow();
}
}The message records the exception class, not its text, because provider and driver exceptions sometimes embed URLs with tokens or connection strings with passwords. Virtual threads need Java 21; on older runtimes use a small fixed pool.
Reporting failures safely
When validation fails, print one block that an on-call engineer can act on without reading code: every key, what is wrong and what is expected, sorted by key, with sensitive values never shown. Then exit with a distinct status code, for example 2 for invalid configuration as opposed to 1 for a crash, so an orchestrator or CI step can tell the difference. On Kubernetes, also write the report to /dev/termination-log (the default termination message path) so it appears in kubectl describe pod.
CONFIG INVALID (2 problems, exit 2)
AGENT_MODEL not in allowlist [gemini-2.5-flash, gemini-2.5-pro]
agent:support_agent instruction uses {user_tier} but no session seeds itNote that a typo such as gemini-2.5-flsh still matches gemini-.*, so the registry alone would accept it; the first line comes from the allowlist rule, which is why you need both. If validation passes, freeze the result. Build the runner from the validated AgentConfig and nothing else, so no code path can reach around the gate and read a raw environment variable later. Log a fingerprint, such as a SHA-256 of the canonicalised non-secret values, so you can tell from logs which configuration a replica is running.
Putting it in main
Wiring it together gives a main method with two modes. The normal mode validates, builds and serves. The --validate-only mode runs tiers one to three, prints the report and exits, which is what CI and a pre-deploy job call against the rendered configuration for each environment.
public static void main(String[] args) {
boolean validateOnly = Arrays.asList(args).contains("--validate-only");
Problems p = new Problems();
AgentConfig cfg = ConfigLoader.load(System.getenv(), p);
if (p.isEmpty()) ConfigRules.crossField(cfg, p);
LlmAgent root = null;
if (p.isEmpty()) {
ConfigRules.validateModel(cfg.model(), ALLOWED_MODELS, p);
root = Agents.root(cfg); // pure construction, no I/O
ConfigRules.validateTree(root, SEEDED_KEYS, Agents.instructions(), p);
}
if (p.isEmpty() && !validateOnly) Probes.runProbes(Probes.forConfig(cfg), Duration.ofSeconds(8), p, LOG);
if (!p.isEmpty()) { Report.print(p); System.exit(2); }
if (validateOnly) { System.out.println("CONFIG OK " + Fingerprint.of(cfg)); return; }
Server.start(root, cfg); // readiness flips true only after this
}Tiers stop early on purpose: if parsing failed, cross-field rules would report noise about defaults. On Kubernetes, keep the readiness probe false until Server.start finishes, and give the startup probe enough budget to cover the probe deadline; probe configuration is covered in Running ADK Java Agents on Kubernetes.
Worked example: a half-finished Vertex migration
A team moves its support agent from the Gemini API to Vertex AI. The pull request sets GOOGLE_GENAI_USE_VERTEXAI=true and GOOGLE_CLOUD_PROJECT in the staging overlay, but forgets GOOGLE_CLOUD_LOCATION and leaves the old API key in place. In the same change someone edits the instruction to say “Customers on the {user_tier} plan get priority”, but session creation only seeds user_name.
Without the gate, staging pods start, report ready, and every request fails at the model call. With the gate, the CI job runs --validate-only against the rendered staging configuration. Tier two fails with two lines, the missing location and the ignored API key, and stops. The author fixes both, and the next run reaches tier three and reports the unseeded {user_tier}. Two red builds, minutes apart, and no staging outage.
Failure modes
| Failure mode | Symptom | Mitigation |
|---|---|---|
| Fail-first validation | Several redeploys to find several errors | Accumulate problems; report once |
| Lenient parsing | Misspelled boolean silently false | Strict parsers with explicit accepted values |
| Secret in error text | API key in logs or termination message | Secret wrapper; log exception class only |
| Probe hangs | Pod never ready, crash loop on startup probe | Shared deadline, cancel on timeout |
| Flaky required probe | Deploys blocked by provider blips | Mark model calls advisory; retry once |
| Bypass of the gate | Code reads env vars after boot | Build only from the frozen config object |
Trade-offs
Strictness has a price. Every new rule can block a deploy, including an emergency one, so keep an explicit and audited override, such as a single CONFIG_VALIDATION_WAIVE variable that lists waived rule identifiers and is itself logged loudly. Live probes lengthen startup and couple your deploy to other systems' health; keep them few and bounded. A model allowlist needs an owner and a review path for new models. Prompt quality belongs in evals, not here.
Frameworks can help. If you run under Spring Boot, @ConfigurationProperties classes annotated with @Validated and Jakarta Bean Validation constraints fail context startup with a list of binding errors, which covers tier one. The cross-field, agent-tree and probe tiers still need code like the above.
What to do next
- Inventory every configuration key the agent reads and route all reads through one loader that returns a typed, immutable record.
- Replace throw-on-first-error parsing with a problem accumulator and strict parsers for booleans, durations and URIs.
- Write cross-field rules for the Gemini API versus Vertex AI choice and any other mode switches.
- Add tree checks:
LlmRegistry.getLlmplus an allowlist, unique agent names, expected tools and seeded placeholder keys. - Add bounded concurrent probes, classify each as required or advisory, and keep readiness false until they pass.
- Ship a
--validate-onlymode, run it in CI for every environment, and pair it with per-invocation limits from ADK Java Runtime Configuration.