Every ADK Java service answers the same questions on every request. Which model handles this call, with which credentials, at what temperature, and how many model calls is it allowed? The answer is rarely written in one place. A default sits in code, a YAML file overrides it, an environment variable overrides that, an agent builder hard-codes something else, and a callback rewrites the request just before it leaves the process. When those layers disagree the service does not crash. It quietly runs with a value nobody chose, and you find out from the bill or from an incident review.

This article is about precedence: the rules that decide which layer wins. It covers the precedence already built into the libraries ADK Java sits on, including how the google-genai client chooses between an API key and a Google Cloud project; the inheritance rules of the agent tree, which are different for the model and for generation settings; and how to build your own layered resolver that records where every value came from. The library behaviour described here was read from the current adk-java and java-genai sources. It changes between releases, so confirm it against the versions you pin.

The layers and their order

Think of configuration as a stack of layers, lowest first. Each layer may set a key or leave it alone, and a resolver walks the stack to find the effective value. The usual order for a service is shown below, but the order is a design decision, not a law. Write yours down.

LayerTypical ownerChanges whenExample keys
1. Code defaultsDeveloperReleasedefault model id, retry counts, RunConfig's own builder defaults
2. Packaged fileDeveloperReleaseagent names, tool timeouts
3. Profile or mounted filePlatform teamDeployproject, region, endpoints per environment
4. Environment variablesPlatform teamDeploy or restartsecrets, GOOGLE_CLOUD_LOCATION
5. CLI args or system propertiesOperatorProcess startone-off overrides during an incident
6. Dynamic overridesProduct or SREAny timekill switches, per-tenant limits
7. Per-invocation RunConfigRequest pathEvery callLLM call budget, streaming mode
8. Before-model callbackRequest pathEvery model calltoken caps, temperature for eval traffic
Where an ADK Java value can come from (higher wins, unless the rule says otherwise)8 Before-model callbackrewrites LlmRequest.Builder per model call7 Per-invocation RunConfigbuilt per request by your factory6 Dynamic overridesflags service, tenant policy5 CLI args / system propertiesper process launch4 Environment variablesyours, and GOOGLE_* read by libraries3 Profile fileapplication-prod.yaml, ConfigMap2 Packaged fileapplication.yaml in the jar1 Code defaultsconstants, builder defaultsYour resolvermerges 1-6, records the sourcegoogle-genai Clientreads GOOGLE_* env vars itselfAgent treemodel inherited, config notprivate readbuilder argsGemini instanceThe seam: values your resolver never sawstill reach the model through the client.
Figure 1. The eight layers and the seam. Your resolver merges layers 1 to 6, but the google-genai client reads its own environment variables, so a value can reach the model without ever passing through your code.

There are two kinds of layer in this stack. Some you own: you decide the merge rules and can log the result. Others are resolved privately by a library, and the most important one is the google-genai client, which reads GOOGLE_* variables straight from the process environment. Most precedence bugs live at that seam. Your resolver believes the model runs in one region while the client has read a different region from the environment, and nothing in your logs disagrees with you.

Precedence inside the google-genai client

ADK's Gemini model class wraps a google-genai Client. The client resolves its backend and credentials with rules that are worth knowing exactly, because they mix three behaviours: explicit values beat environment values, some conflicts throw, and other conflicts only log a warning.

DecisionRule in the java-genai source
BackendBuilder enterprise(...) wins, then builder vertexAI(...). If both are set and differ, the build throws IllegalArgumentException.
Backend from environmentOnly if neither builder flag is set: GOOGLE_GENAI_USE_ENTERPRISE, then GOOGLE_GENAI_USE_VERTEXAI. If both are set and differ, it logs a warning and the ENTERPRISE value wins. Neither set means the Gemini Developer API.
Builder project or location on the Developer APIThrows: the Developer API does not accept them. A stray environment project does not throw.
API key from environmentGOOGLE_API_KEY beats GEMINI_API_KEY, with a warning.
Explicit credentials plus explicit API key (Vertex backend)Throws: choose one.
Explicit API key, project only in environmentThe explicit key wins and the environment project and location are dropped, with a warning.
Explicit project or location, key only in environmentThe explicit project wins and the environment key is dropped, with a warning.
Both only in environmentProject and location win over the key, with a warning.
Nothing sets a locationWith no API key and no custom base URL, the location defaults to global.

Two lessons follow. First, the rules are asymmetric: a conflict between builder arguments fails at startup, but the same conflict between environment variables is a log line. If your deployment sets GOOGLE_GENAI_USE_VERTEXAI=true and a newer base image adds GOOGLE_GENAI_USE_ENTERPRISE=false, the ENTERPRISE value wins, the client quietly uses the Developer API, and the warning is the only trace. Second, the enterprise flag is recent. Gemini Enterprise Agent Platform is the April 2026 rebrand of Vertex AI, and current java-genai releases read both variable names, while older releases read only the VERTEXAI one. ADK Java's main branch pins java-genai 1.75.0, which has both builder methods. Check the version your build actually resolves.

On top of the client sits the Gemini builder, which picks an explicit apiClient first, then an apiKey, then vertexCredentials, and only then a default client built from the environment. Passing a model as a string, such as .model("gemini-2.5-flash"), goes through LlmRegistry, which builds that default client. So a string model name means your credentials come entirely from the environment. The Gemini class deep dive walks through the rest of that class. The safe pattern is to resolve everything yourself and hand the client explicit values:

Client client = Client.builder()
    .enterprise(true)                          // older java-genai: .vertexAI(true)
    .project(cfg.require("gcp.project"))       // explicit beats GOOGLE_CLOUD_PROJECT
    .location(cfg.require("gcp.location"))     // explicit beats GOOGLE_CLOUD_LOCATION
    .build();

BaseLlm model = Gemini.builder()
    .modelName(cfg.require("agent.model"))
    .apiClient(client)                         // wins over apiKey and vertexCredentials
    .build();

LlmAgent root = LlmAgent.builder()
    .name("triage")
    .model(model)                              // an instance, not a string: no registry lookup
    .instruction(cfg.require("agent.triage.instruction"))
    .build();

Inheritance in the agent tree

The agent tree has its own precedence rules, and they are inconsistent in a way that surprises people.

The model is inherited. When an LlmAgent has no model, it walks up its parents to the nearest LlmAgent ancestor and uses that agent's model. Workflow agents such as a sequential agent in between are skipped. If no ancestor has a model, resolution throws IllegalStateException naming the agent.

Generation settings are not inherited. When the request is built, the agent's own generateContentConfig is copied in, or an empty config if the agent has none. Nothing is merged from the parent. Temperature, output token limits and safety settings set on the root agent do not reach a sub-agent that sets nothing; the sub-agent runs with the model's defaults. The globalInstruction is different again: it is read from the root agent and applied across the tree.

Consider a real tree. A support system has a root triage agent with the model, a temperature of 0.1, a 1,024 token output cap and strict safety settings. Its refunds sub-agent was written quickly and sets only a name and an instruction. Refunds inherits the model, so it works in testing and nobody notices that it runs at default temperature, with no output cap and default safety thresholds. The fix is to make the base explicit and merge it field by field when you build each agent:

static GenerateContentConfig withBase(GenerateContentConfig base, GenerateContentConfig own) {
    GenerateContentConfig.Builder b = base.toBuilder();
    own.temperature().ifPresent(b::temperature);           // scalars: own value replaces base
    own.maxOutputTokens().ifPresent(b::maxOutputTokens);
    own.safetySettings().ifPresent(b::safetySettings);     // lists: replace, never concatenate
    return b.build();
}

LlmAgent refunds = LlmAgent.builder()
    .name("refunds")
    .instruction(cfg.require("agent.refunds.instruction"))
    .generateContentConfig(withBase(baseConfig, refundsOverrides))
    .build();

Then add a startup check that walks the tree and fails if any LlmAgent has an empty generateContentConfig or no model of its own. It turns a silent inheritance gap into a failed deploy.

A layered resolver with provenance

For the layers you own, a resolver needs three things: a fixed order, explicit merge rules, and provenance, meaning a record of which layer supplied each final value. If you use a framework such as Spring Boot, adopt its documented property-source order rather than fighting it. Otherwise the core is small:

public final class LayeredConfig {
    public record Layer(String name, Map<String, Object> values) {}
    public record Resolved(Object value, String source) {}

    private final Map<String, Resolved> effective = new TreeMap<>();

    public LayeredConfig(List<Layer> lowestFirst) {
        for (Layer layer : lowestFirst) {
            flatten("", layer.values(), layer.name());
        }
    }

    @SuppressWarnings("unchecked")
    private void flatten(String prefix, Map<String, Object> m, String source) {
        for (var e : m.entrySet()) {
            String key = prefix.isEmpty() ? e.getKey() : prefix + "." + e.getKey();
            Object v = e.getValue();
            if (v instanceof Map<?, ?> nested) {          // maps: deep merge, key by key
                flatten(key, (Map<String, Object>) nested, source);
            } else if (v == null) {                        // explicit null: unset at this layer
                effective.keySet().removeIf(k -> k.equals(key) || k.startsWith(key + "."));
            } else {                                       // scalars and lists: replace
                effective.put(key, new Resolved(v, source));
            }
        }
    }

    public String require(String key) {
        Resolved r = effective.get(key);
        if (r == null) throw new IllegalStateException("missing config key " + key);
        return String.valueOf(r.value());
    }

    public Map<String, Resolved> dump() { return Collections.unmodifiableMap(effective); }
}

Three rules in that code deserve a decision of their own. Maps merge deeply, so a profile file that sets only agent.refunds.timeout does not erase the rest of agent.refunds. Lists replace, because concatenation makes it impossible for a higher layer to remove an entry, such as a tool or an allowed region. An explicit null removes a key, which gives operators a way to unset something without inventing a sentinel value. Environment variables need a name mapping into the same key space, for example APP_AGENT_REFUNDS_TIMEOUT to agent.refunds.timeout; give your variables a prefix so you never accidentally map the library's own GOOGLE_* variables into your tree.

Per-invocation and per-call layers

Layers 7 and 8 are evaluated on the request path. The RunConfig deep dive covers the fields; the precedence point is that "later wins" is the wrong rule for limits. A tenant policy should be able to lower the LLM call budget, but not raise it above the operator's ceiling, so a limit resolves to the minimum across layers, not the last value written:

int ceiling = Integer.parseInt(cfg.require("limits.maxLlmCalls"));      // operator, layers 1-6
int tenant  = tenantPolicy.maxLlmCalls(tenantId).orElse(ceiling);        // layer 6, dynamic
RunConfig run = RunConfig.builder()
    .maxLlmCalls(Math.min(tenant, ceiling))                              // most restrictive wins
    .build();

The before-model callback is the top of the stack. It receives the LlmRequest.Builder for one model call and can change anything, so it overrides every layer beneath it. Use it for values that genuinely depend on the call, and always clamp rather than set. See the callback architecture for how callbacks chain.

LlmAgent agent = LlmAgent.builder()
    .name("refunds")
    .model(model)
    .instruction(instruction)
    .beforeModelCallbackSync((ctx, req) -> {
        GenerateContentConfig cur = req.config()
            .orElseGet(() -> GenerateContentConfig.builder().build());
        int cap = tenantPolicy.maxOutputTokens(ctx.userId());
        int asked = cur.maxOutputTokens().orElse(cap);
        if (asked > cap) {
            log.info("clamped maxOutputTokens {} -> {} for agent {}", asked, cap, ctx.agentName());
            req.config(cur.toBuilder().maxOutputTokens(cap).build());
        }
        return Optional.empty();                 // empty: continue to the model
    })
    .build();

Making the winner visible

Precedence is only debuggable if the winner is visible. At startup, log the effective configuration once: every key, its value, and the layer that supplied it, with secrets replaced by a short hash so you can tell two keys apart without leaking either. Compute a hash of the effective configuration and attach it to traces and to the service's health endpoint, so an incident review can ask "which config was this request running under" and get an answer.

Treat library warnings about precedence as errors in production. A log appender or a startup test that fails when it sees "will take precedence" or "conflicting values" from the client turns the asymmetric behaviour above into a consistent one. The boot sequence article shows where in startup such checks fit.

Failure modes

  • Stale environment from a base image. A GOOGLE_* variable set in a base image or CI runner silently changes the backend or project. Pass explicit values and unset the variables you do not intend to use.
  • Assuming config inheritance. Sub-agents inherit the model but not temperature, output caps or safety settings.
  • Two resolvers for one key. Your YAML sets gcp.location and the client reads GOOGLE_CLOUD_LOCATION; they drift apart and only one is logged.

Testing precedence

Test precedence as a table: one row per key and combination of layers, with the expected winner and source. For the client, test the startup failure paths too, because they are part of the contract.

@ParameterizedTest
@CsvSource({
    // file,   env,     cli,     expected, source
    "us-east1, ,        ,        us-east1, profile",
    "us-east1, europe-west4, ,   europe-west4, env",
    "us-east1, europe-west4, asia-south1, asia-south1, cli",
})
void locationPrecedence(String file, String env, String cli, String expected, String source) {
    LayeredConfig cfg = new LayeredConfig(List.of(
        layer("profile", "gcp.location", file),
        layer("env", "gcp.location", env),
        layer("cli", "gcp.location", cli)));
    assertEquals(expected, cfg.require("gcp.location"));
    assertEquals(source, cfg.dump().get("gcp.location").source());
}

@Test
void conflictingBackendFlagsFailFast() {
    assertThrows(IllegalArgumentException.class,
        () -> Client.builder().enterprise(true).vertexAI(false).build());
}

Here layer(...) is a test helper that returns an empty layer when the value is blank. Add one more test that builds your real agent tree and asserts that every LlmAgent has an explicit model and configuration. For the broader topic of profiles and secrets, see environment and configuration management.

What to do next

  1. Write down your layer order, including the two request-path layers, and put it in the repository next to the code.
  2. Replace string model names with explicit Gemini instances built from an explicit Client.
  3. Search your images and CI runners for GOOGLE_GENAI_USE_*, GOOGLE_API_KEY and GOOGLE_CLOUD_*, and remove any you do not set on purpose.
  4. Add a base GenerateContentConfig and merge it into every agent; add the tree-walking startup check.
  5. Resolve limits with minimum-wins, and make every callback clamp and log rather than set.
  6. Log the effective configuration with sources and a hash at startup, and fail on client precedence warnings.
  7. Add a parameterised precedence test for each key that has caused an incident, then for each new key.
Key takeaway: Write down your layer order and make every value's source visible. Hand the google-genai client explicit backend, project and location values, because its environment conflicts only warn. Remember that sub-agents inherit the model but not temperature, token caps or safety settings, so merge a base config into each agent. Merge maps deeply, replace lists, resolve limits to the minimum, and let callbacks clamp rather than set.