An ADK Java agent names the model it wants in one of two ways: it holds a BaseLlm object, or it holds a string such as gemini-2.5-flash. Strings are what make agents portable. They can live in YAML agent configs, differ per environment and be changed without recompiling. But a string is only useful if something turns it into a working client with the right credentials, endpoint and retry policy. In ADK Java that something is LlmRegistry, a small static class that maps regular-expression patterns to factories and caches the instances it builds.
The registry is deliberately minimal, and most production problems with model selection come from treating it as more than it is. This article reads its behaviour from the source and shows how resolution works for an agent tree. It covers the semantics that bite: full-match regexes, an undefined pattern order, the cache and its interaction with exceptions. It then builds the discovery layer ADK does not provide: a model catalog with aliases, boot-time registration and validation, and routing without deadlocks. A worked example spans three environments, and the article closes with testing, failure modes, trade-offs and a checklist. Two earlier pages cover the BaseLlm contract and writing a custom adapter; this one assumes you have an adapter and need to wire many of them up reliably.
How an agent resolves its model
Resolution happens lazily, the first time a flow needs the model. LlmAgent.resolvedModel() checks three things in order. If the agent was built with a BaseLlm instance, that instance is used and the registry is never consulted. If it was built with a string, the string goes to LlmRegistry.getLlm(name). If it has neither, the method walks up the parent chain to the nearest LlmAgent and uses that agent's resolved model; if no ancestor has one, it throws IllegalStateException.
The result is memoised per agent with double-checked locking, so resolution runs once per agent object for the life of the process. Two consequences follow. First, sub-agents of a coordinator inherit the coordinator's model unless you set one, which is convenient but means a change at the root silently changes every child. Second, nothing you register after an agent has resolved will affect that agent. YAML configs follow the same path: the model: field of an agent config is a plain string handed to builder.model(String), so every config-defined agent depends on the registry.
The name that reaches the provider is not always the name you typed. The basic request processor sets the request's model to the resolved instance's own model() value. A factory that maps an alias such as fast to Gemini.builder().modelName("gemini-2.5-flash") therefore sends the real model name upstream, and capability checks keyed on the model name see the real name too.
Inside LlmRegistry
The whole registry is two static ConcurrentHashMaps. One maps a pattern string to an LlmFactory, a functional interface with a single create(String modelName) method. The other maps a concrete model name to the BaseLlm already built for it. A static initialiser registers three defaults: gemini-.* and gemma-.*, both served by the Gemini class, and apigee/.*. Lookup is a computeIfAbsent on the instance map whose mapping function iterates the factories and calls the first one whose pattern matches.
| Behaviour | What it means for you |
|---|---|
Patterns use String.matches, a full match | gpt-4 does not match gpt-4o. An unescaped . in model-4.1 matches any character. |
| Factories live in a hash map | When two patterns match one name, which wins is unspecified. Keep patterns disjoint, ideally with a provider prefix. |
| The factory key is the pattern string | Registering exactly gemini-.* again replaces the default factory deterministically. A narrower overlapping pattern does not. |
| One instance per name, shared | Every agent and thread that asks for a name gets the same object, so adapters must be thread-safe. |
| Exceptions are not cached | A factory that throws leaves no entry, so every later call retries it; a missing credential fails on every request. |
No match throws IllegalArgumentException | Raised at first resolution, which is the first request unless you validate at boot. |
| No unregister, no cache clear | The test-only reset is package-private, so production code cannot evict an instance. |
The default Gemini factory builds a client without explicit credentials. The google-genai Java client then reads its environment: GOOGLE_API_KEY or GEMINI_API_KEY for the Gemini API, or GOOGLE_GENAI_USE_VERTEXAI with GOOGLE_CLOUD_PROJECT and GOOGLE_CLOUD_LOCATION for Vertex AI. That is convenient on a laptop and risky in production, where you want credentials to come from your secret store and to fail loudly. Overriding the exact gemini-.* pattern with your own factory is the clean way to take control.
Writing patterns that cannot collide
Treat patterns as a namespace, not a convenience. The safest convention is a provider prefix ending in a slash, because a slash never appears in Gemini model names and two prefixes cannot overlap:
// Disjoint by construction: each pattern owns one prefix.
LlmRegistry.registerLlm("openai/.*", name -> OpenAiLlm.create(name, secrets));
LlmRegistry.registerLlm("local/.*", name -> VllmLlm.create(name, localUrl));
LlmRegistry.registerLlm("bedrock/.*", name -> BedrockLlm.create(name, awsClient));
// Literal names: quote them so dots are not wildcards.
LlmRegistry.registerLlm(Pattern.quote("acme-chat-4.1"), name -> AcmeLlm.create(name));Each factory should strip its prefix before talking to the provider, and should build the instance's model() from the bare provider name if you want capability checks and logs to show the real model. Avoid catch-all patterns such as .*: they overlap every other pattern, so whether they win is down to hash order and can change when you add an unrelated registration.
Discovery: a model catalog on top of the registry
ADK has no discovery mechanism: no service-provider interface, no config file and no listing API. Discovery is something your application does before agents resolve. The pattern that scales is a small catalog, loaded from configuration, that lists every model the deployment may use, who owns it and how to build it. The registry then becomes a thin index over the catalog. Everything in this section is application code layered on ADK's public API.
record ModelEntry(String name, String provider, String target, Map<String, String> opts) {}
final class ModelCatalog {
private final Map<String, ModelEntry> entries; // keyed by exact name or alias
private final Map<String, ProviderFactory> providers; // "vertex", "openai", "local" ...
ModelCatalog(List<ModelEntry> list, Map<String, ProviderFactory> providers) {
this.entries = list.stream().collect(Collectors.toUnmodifiableMap(ModelEntry::name, e -> e));
this.providers = Map.copyOf(providers);
}
/** Register one literal pattern per entry, so no two patterns can overlap. */
void registerAll() {
for (ModelEntry e : entries.values()) {
ProviderFactory pf = providers.get(e.provider());
if (pf == null) throw new IllegalStateException("Unknown provider " + e.provider() + " for " + e.name());
LlmRegistry.registerLlm(Pattern.quote(e.name()), n -> pf.build(e.target(), e.opts()));
}
}
/** Resolve every name now, so a bad entry fails the deploy, not the first user. */
void validate() {
for (String name : entries.keySet()) {
BaseLlm llm = LlmRegistry.getLlm(name);
log.info("model {} -> {} ({})", name, llm.model(), llm.getClass().getSimpleName());
}
}
Set<String> names() { return entries.keySet(); }
}Registering literal, quoted names instead of wildcards removes the overlap problem entirely and makes the catalog the single list of permitted models: an agent config that names anything else fails validation. Aliases such as fast and reasoning are just entries whose target is a provider model name, so moving every fast agent to a new model is one catalog edit. If adapters ship in separate JARs, java.util.ServiceLoader can discover ProviderFactory implementations on the classpath and fill the providers map; keep the catalog itself in configuration so that what is allowed stays reviewable.
Validation calls getLlm for every name, which builds every client at boot. That is the point: it surfaces missing secrets, wrong endpoints and typos while the old version is still serving. Keep factories cheap; network health checks belong in readiness probes, not in constructors.
Routing and fallback without recursive lookups
Fallback and routing look like natural registry features, but a factory must not call LlmRegistry.getLlm for another name. Factories run inside computeIfAbsent, and ConcurrentHashMap forbids updating the map from inside its own mapping function; a recursive resolution may throw IllegalStateException or deadlock, depending on where the keys hash. Build delegates first and close over them:
final class FallbackLlm extends BaseLlm {
private final BaseLlm primary, secondary;
FallbackLlm(String name, BaseLlm primary, BaseLlm secondary) {
super(name);
this.primary = primary;
this.secondary = secondary;
}
@Override
public Flowable<LlmResponse> generateContent(LlmRequest req, boolean stream) {
// Only retry before anything was emitted; never splice two partial streams.
AtomicBoolean emitted = new AtomicBoolean(false);
return primary.generateContent(as(primary, req), stream)
.doOnNext(r -> emitted.set(true))
.onErrorResumeNext(err -> emitted.get() || !isRetryable(err)
? Flowable.error(err)
: secondary.generateContent(as(secondary, req), stream));
}
@Override
public BaseLlmConnection connect(LlmRequest req) {
return primary.connect(as(primary, req)); // live sessions are not failed over
}
// The request arrives named "support-default"; each delegate needs its own name.
private static LlmRequest as(BaseLlm target, LlmRequest req) {
return req.toBuilder().model(target.model()).build();
}
}
// At boot: resolve delegates OUTSIDE any factory, then register the composite.
BaseLlm p = LlmRegistry.getLlm("vertex-flash");
BaseLlm s = LlmRegistry.getLlm("openai/gpt-4.1-mini");
LlmRegistry.registerLlm(Pattern.quote("support-default"), n -> new FallbackLlm(n, p, s));BaseLlm has two abstract methods, generateContent and connect, so a composite must implement both. The as helper matters: the request processor names the request after the composite, and the Gemini adapter sends req.model() upstream when it is present, so without the copy the provider would receive support-default as a model name. For per-request hedging rather than failover, see request hedging.
Worked example: one agent tree, three environments
A support product runs three agents: a router on a fast model, a resolver on a reasoning model and a summariser that inherits from the router. Developers run open models locally; staging and production use Vertex AI with a second provider as fallback. The catalog files differ per environment, and the agent YAML never changes:
# agents/router.yaml (identical in every environment)
name: router
model: fast
instruction: Route the ticket to billing, outage or account.
# catalog-dev.yaml # catalog-prod.yaml
- name: fast - name: fast
provider: local provider: vertex
target: qwen2.5-7b-instruct target: gemini-2.5-flash
- name: reasoning - name: reasoning
provider: local provider: vertex
target: qwen2.5-32b-instruct target: gemini-2.5-proBoot runs in a fixed order: load the catalog for the active environment, register every entry, load agent configs, then call validate() and resolve each root agent's model once. The order matters because resolution is memoised; the runtime boot sequence shows where this slots in. In production, validate() builds two Vertex clients from secrets fetched at startup, and the deploy fails if either secret is missing. When the team later moves the router to a newer flash model, the change is one line in catalog-prod.yaml, rolled out like any configuration change, with the old process serving until the new one passes validation. Logs from model call orchestration show the provider model name, not the alias, so dashboards keep working.
Testing with a static registry
Because the registry is static and process-wide, tests that register fakes leak into each other. Three rules keep suites deterministic. Give every test its own unique name, such as test-<uuid>, so cached instances never collide. Prefer passing a BaseLlm instance to the agent builder in unit tests, which bypasses the registry. And reserve registry-based tests for the boot path itself: load a test catalog, run validate() and assert that an unknown name fails with IllegalArgumentException.
Failure modes
- Unsupported model on first request. A typo in YAML passes agent construction and fails at first resolution. Fix: validate every configured name at boot.
- The wrong factory wins. Two patterns match one name and hash order picks one; it can change when an unrelated pattern is added. Fix: prefixes or quoted literals only.
- Registration after resolution has no effect. Agents memoise their model and the registry caches instances. Fix: register everything before loading agents; changing models means a restart or a new agent object.
- Credential retries on every call. A factory that throws is not cached, so a missing key costs a failed construction per request. Fix: fail the deploy instead.
- Recursive resolution. A router factory calling
getLlmcan throw or hang. Fix: resolve delegates before registering the composite. - Shared mutable state. One instance per name is shared by all agents; per-request fields in an adapter race. Fix: keep adapters stateless.
- Silent environment credentials. The default Gemini factory picks up whatever key is in the environment. Fix: override
gemini-.*with a factory that reads your secret store.
Trade-offs
Strings plus a registry buy you configuration-driven agents, environment portability and one place to audit which models run. They cost you compile-time safety and introduce a global, mutable, process-wide singleton. Instances passed directly to builders are the opposite: explicit, testable and immune to ordering bugs, but every model change is a code change. A sensible split is instances in libraries and tests, strings resolved through a validated catalog in deployable applications. If you need per-tenant credentials, encode the tenant in the model name, so each tenant gets its own cached instance, and avoid building per-request clients inside the registry, which it was never designed to hold.
What to do next
- List every model string your agents and YAML configs use today, and check each against the registered patterns for overlaps and unescaped dots.
- Create a catalog file per environment with exact names and aliases, and register each entry as a quoted literal.
- Override the default gemini-.* factory so Gemini credentials come from your secret store rather than ambient environment variables.
- Call getLlm for every catalog name at boot and fail the deploy on any exception.
- Move any fallback or routing logic into composites whose delegates are resolved before registration.
- Switch unit tests to builder-supplied instances and give registry-based tests unique names.
- Log the resolved provider model name on every call so aliases never hide what actually ran.