Most ADK Java examples connect an agent to Gemini with one line: .model("gemini-2.5-flash"). That line hides a chain of decisions that matter the first time you deploy: which backend the request goes to, which credentials it uses, which HTTP client carries it, how long that client waits before giving up, and when a typo in the model name is detected. Getting any of them wrong produces failures that look unrelated to configuration, such as a request that hangs forever or an exception on the first user message instead of at startup.
This article follows that chain from the agent builder to the network, using the ADK Java and Google Gen AI Java SDK source on their main branches at the time of writing. It covers the three ways to hand an agent a Gemini model, choosing between the Gemini Developer API and Vertex AI, building a client with explicit timeouts and retries, a fail-fast boot sequence, and the failure modes of each choice. Tool calling and model features have their own articles: Gemini function calling in ADK Java and Gemini 2.5 features in ADK Java. Model ids below are examples; check Google's current model list before you pin one.
From agent builder to network
When an LlmAgent needs to call the model, it asks for its resolved model. If you passed a BaseLlm instance to model(...), that instance is used. If you passed a string, the agent resolves it lazily, on first use, through LlmRegistry.getLlm(name). The registry holds regular-expression patterns mapped to factories; on main it registers gemini-.*, gemma-.* and apigee/.*, and the Gemini factory is simply Gemini.builder().modelName(modelName).build(). The resulting instance is cached per model name, so every agent in the JVM that uses the same string shares one Gemini object and one HTTP client.
The Gemini class extends BaseLlm and owns a com.google.genai.Client from the Google Gen AI Java SDK. For each turn it prepares the request and calls either generateContent or generateContentStream on the client's async models API. In streaming mode the chunks are aggregated so the runtime sees partial events followed by one final aggregated response; the runtime side of that is covered in the streaming architecture article, and the contract a model class must meet is in the BaseLlm interface article.
Three ways to give an agent a Gemini model
There are three ways to give an agent a Gemini model, in increasing order of control. Gemini.Builder exposes modelName, apiClient, apiKey, vertexCredentials and httpExecutorService. If several are set, the builder uses an explicit client first, then an API key, then Vertex credentials, and otherwise builds a default client from the environment.
import com.google.adk.agents.LlmAgent;
import com.google.adk.models.Gemini;
import com.google.adk.models.VertexCredentials;
import com.google.auth.oauth2.GoogleCredentials;
// 1. A model id string: resolved lazily through LlmRegistry, client built from env vars.
LlmAgent quick = LlmAgent.builder()
.name("support")
.model("gemini-2.5-flash")
.instruction("Answer billing questions briefly.")
.build();
// 2. An explicit Gemini instance with a key from your secret store, not from env.
Gemini viaKey = Gemini.builder()
.modelName("gemini-2.5-flash")
.apiKey(secrets.read("gemini-api-key"))
.build();
// 3. Vertex AI project, region and credentials (see the backend flag caveat below).
Gemini viaVertex = Gemini.builder()
.modelName("gemini-2.5-flash")
.vertexCredentials(VertexCredentials.builder()
.project("acme-prod")
.location("europe-west4")
.credentials(GoogleCredentials.getApplicationDefault()) // throws IOException
.build())
.build();
LlmAgent support = LlmAgent.builder()
.name("support")
.model(viaVertex) // a BaseLlm: no registry lookup, no shared cache entry
.instruction("Answer billing questions briefly.")
.build();Option one is right for prototypes. Options two and three make the credential source explicit, which is what you want in production, and they give each agent its own instance instead of the registry's shared one. The fourth option, passing your own Client through apiClient(...), is the only one that lets you control timeouts and retries, and it is covered below.
Choosing a backend and credentials
The SDK talks to two backends with the same request types. The Gemini Developer API authenticates with an API key and is the quickest start. Vertex AI on Google Cloud authenticates with Google credentials (Application Default Credentials, a service account or workload identity), is addressed by project and location, and brings Cloud IAM, quotas, audit logging and regional data handling. The Gen AI Java SDK README now calls the second backend the Gemini Enterprise Agent Platform; the code accepts both names.
The SDK chooses the backend in this order: an explicit enterprise(true) or vertexAI(true) on the client builder, then the environment variables GOOGLE_GENAI_USE_ENTERPRISE and GOOGLE_GENAI_USE_VERTEXAI. Setting the two builder flags to different values throws IllegalArgumentException; if the two environment variables conflict, the SDK logs a warning and uses the enterprise one. For the Developer API it reads GOOGLE_API_KEY and, at lower precedence, GEMINI_API_KEY; for Vertex AI it reads GOOGLE_CLOUD_PROJECT and GOOGLE_CLOUD_LOCATION.
One interaction is easy to miss. On main, ADK's Vertex path sets project, location and credentials on the client but does not set the backend flag, and the SDK constructor rejects a project or location when the backend resolves to the Developer API, with the message Gemini API does not support project/location. So option three above works only when one of the two environment variables is set to true. If you would rather not depend on the environment, build the client yourself and set the flag explicitly. Older SDK versions may predate the enterprise naming, so check the google-genai version your ADK release pulls in.
Timeouts and retries: own the HTTP client
On main, when ADK builds the client itself, it creates an OkHttp client with connect, read and write timeouts set to Duration.ZERO. In OkHttp zero means no timeout. A stalled TCP connection or a load balancer that stops forwarding bytes will therefore hold the calling thread or subscription indefinitely, and ADK has no built-in model retry either; the error recovery article covers what the runtime does when a model call does fail. The fix is to own the client.
import com.google.genai.Client;
import com.google.genai.types.ClientOptions;
import com.google.genai.types.HttpOptions;
import com.google.genai.types.HttpRetryOptions;
import okhttp3.OkHttpClient;
import java.time.Duration;
static Client vertexClient(AppConfig cfg) throws IOException {
OkHttpClient http = new OkHttpClient.Builder()
.connectTimeout(Duration.ofSeconds(5))
.readTimeout(Duration.ofSeconds(60)) // max silence between bytes, not total time
.writeTimeout(Duration.ofSeconds(30))
.build();
HttpOptions httpOptions = HttpOptions.builder()
.retryOptions(HttpRetryOptions.builder()
.attempts(3)
.httpStatusCodes(429, 503))
.build();
return Client.builder()
.vertexAI(true) // explicit: no dependence on env flags
.project(cfg.project())
.location(cfg.location())
.credentials(GoogleCredentials.getApplicationDefault())
.httpOptions(httpOptions)
.clientOptions(ClientOptions.builder().customHttpClient(http).build())
.build();
}
Client client = vertexClient(cfg); // build once, share across agents
Gemini model = Gemini.builder().modelName(cfg.modelId()).apiClient(client).build();The SDK clones a custom OkHttp client and adds its own retry interceptor, so your timeouts survive. Three details matter. The read timeout bounds silence between bytes, not total duration, which is what a streaming call needs; if you add OkHttp's callTimeout for a hard total limit, keep it on a client used only for non-streaming calls or it will cut long streams. Retries multiply cost and latency, so retry only throttling and unavailability codes and budget them alongside the agent's maxLlmCalls limit. And a retry after a stream has started delivering text cannot be transparent: the user has already seen the first half.
Worked example: a boot sequence that fails fast
Consider a billing-support agent deployed to a Kubernetes cluster on Google Cloud. The goal is that every configuration error surfaces at startup, and that a sick endpoint costs one request a bounded wait, not a stuck pod. The boot code builds the client explicitly, forces model resolution, and runs one smoke turn with its own deadline before the readiness probe goes green.
LlmAgent agent = LlmAgent.builder()
.name("billing")
.model(Gemini.builder().modelName(cfg.modelId()).apiClient(vertexClient(cfg)).build())
.instruction("Answer billing questions. Never promise refunds.")
.build();
agent.resolvedModel(); // a string model would fail here, not on turn one
InMemoryRunner runner = new InMemoryRunner(agent, "billing");
Session s = runner.sessionService().createSession("billing", "boot-check").blockingGet();
Event last = runner.runAsync("boot-check", s.id(),
Content.fromParts(Part.fromText("Reply with the single word OK.")))
.timeout(30, TimeUnit.SECONDS) // RxJava deadline over the whole turn
.blockingLast();
log.info("gemini smoke ok model={} text={}", cfg.modelId(), last.stringifyContent());
readiness.markReady();| Misconfiguration | Without this boot sequence | With it |
|---|---|---|
Model id typed as Gemini-2.5-flash | first user turn fails with Unsupported model (the registry pattern is case-sensitive) | with a string id, resolvedModel() throws at boot; with an instance, the smoke turn fails |
| Service account lacks Vertex AI permission | 403 on the first real request | 403 in the smoke turn; pod never ready |
| Model not offered in the chosen location | 404 on first use in that region only | 404 at boot in that region |
| Egress to the endpoint black-holed | request threads hang forever | connect timeout after 5 s, smoke turn fails |
| Quota exhausted | user-facing 429s | 429 retried; the smoke turn fails if all three attempts are throttled |
Keep the smoke turn cheap and run it once per pod start, not per readiness probe, or the probe becomes a cost and quota consumer.
Testing without calling Gemini
Unit tests should never reach Gemini. Because agents accept any BaseLlm, the cleanest seam is to inject a scripted model that returns canned LlmResponse objects, so tests assert on behaviour without network, quota or nondeterminism. If agents are built from string ids in configuration, register a test pattern before any agent resolves its model: LlmRegistry.registerLlm("test-.*", name -> new ScriptedLlm(name)), then set the model id to test-billing in the test profile. Remember the cache: an instance created for a name stays for the life of the JVM, so register early. How to write the scripted model is in implementing a custom LLM in ADK Java.
Keep a small, separate integration suite that does call the real endpoint with the production client configuration: one plain turn, one tool-calling turn and one streaming turn. Run it on a schedule as well as on deploy, because model versions and quotas change underneath you.
Operating the integration
Build one Client per backend and region and share it; each client owns a connection pool, and creating one per request wastes connections and handshakes. Read model ids from configuration so a model change is a config rollout with an evaluation gate, not a code change. Store API keys in a secret manager and prefer workload identity over keys on Google Cloud, because a key in an environment variable leaks through crash dumps and debug endpoints.
Log, per model call, the model id, backend, location, latency, token usage from the response metadata, and the status of any retried attempt. Alert on the share of turns that hit a timeout rather than on raw error counts. If you need regional failover, build one Gemini per region and switch at the agent factory level with a circuit breaker; the SDK's retry interceptor retries the same endpoint and cannot fail over by itself.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| String model id | shortest code, config-driven | env-dependent client, no timeouts, errors on first use, shared instance |
Gemini with apiKey | explicit credential, own instance | key management; Developer API only |
Gemini with vertexCredentials | IAM, regions, audit | still needs the env backend flag on main; no timeout control |
Own Client via apiClient | timeouts, retries, explicit backend | more code; you track SDK option changes |
What to do next
- Find every
.model("...")string in your codebase and decide, per agent, whether it should become an explicitGeminiinstance with its own client. - Build one shared
Clientper backend and region with explicit backend flag, connect, read and write timeouts, and a retry policy limited to 429 and 503. - Check which
google-genaiversion your ADK release depends on and confirm the backend flag and environment variable names against that version. - Add the boot sequence: call
resolvedModel(), run one smoke turn with a deadline, and only then mark the pod ready. - Test a black-holed endpoint in staging, for example by blocking egress to it, and confirm the call fails within your timeout instead of hanging.
- Inject a scripted
BaseLlmin unit tests and keep a small scheduled integration suite against the real endpoint. - Log model id, backend, location, latency and token usage per call, and alert on timeout rate.