In ADK for Java, the "Gemini LLM class" is com.google.adk.models.Gemini; there is no class named GeminiLLM. It is the class that actually talks to Gemini. Every LlmAgent that names a gemini- model ends up calling it, either because you constructed one or because the model registry built one for you from a string. It is a small class, but it makes decisions that show up in production: which backend and credentials you get, what is rewritten in your request before it is sent, how streamed chunks become ADK events, what happens when the model pauses a long generation, and how long a call may hang.

This article reads the class as it ships in adk-java 1.11.0, with google-genai 1.75.0, the client version that release pins. It covers construction, backend selection and its main trap, request preparation, the non-streaming and streaming paths, live connections and the registry cache, then works through a traced request and lists failure modes. For wiring choices, credentials strategy and retry policy see Gemini integration in ADK Java; for the abstract contract the class implements, see the BaseLlm overview.

What the class is

Inside com.google.adk.models.Gemini (adk-java 1.11.0)LlmAgentmodel(String) or model(BaseLlm)LlmRegistrygemini-.*, gemma-.* -> GeminiGemini.builder()apiClient > apiKey > vertex > envStringGemini (extends BaseLlm)holds one google-genai ClientGeminiUtil.prepare...strip labels, displayName (API key only)drop client function-call idsappend user turn if last is not usergenerateContentstream=false: one responsestream=true: partials + finalresume on CONTINUATION (max 256)connectlive session overclient.async.live.connectreturns GeminiLlmConnectiongoogle-genai Client (async)OkHttp with connect/read/write timeouts set to zero by ADKGemini APIAPI keyVertex AIproject, location, ADC
One Gemini instance wraps one google-genai Client. Every request is rewritten by GeminiUtil before it leaves, and the backend the Client chose decides which rewrites apply.

public class Gemini extends BaseLlm is not final and holds exactly one field of interest: a com.google.genai.Client. It overrides two methods. generateContent(LlmRequest, boolean stream) returns a Flowable<LlmResponse> and serves ordinary request-response turns, streamed or not. connect(LlmRequest) returns a BaseLlmConnection for bidirectional live sessions. Everything else, including tools, history and callbacks, is assembled by the agent runtime into the LlmRequest before this class sees it.

The model name is stored by the BaseLlm constructor, but a request can override it: the class sends to llmRequest.model() if present and falls back to its own name. That is why a single instance can serve several model names, and why logs should record the effective name rather than the configured one.

Constructors and builder precedence

There are three public constructors, Gemini(String modelName, Client apiClient), Gemini(String modelName, String apiKey) and Gemini(String modelName, VertexCredentials vertexCredentials), plus a builder. The builder requires modelName and throws a NullPointerException without it. It then picks a client in strict precedence order:

  1. An explicit apiClient(Client) wins, and the API key, Vertex credentials and executor settings are all ignored.
  2. Otherwise apiKey(String) builds a Gemini API client, even if Vertex credentials were also set.
  3. Otherwise vertexCredentials(VertexCredentials) builds a client from the optional project, location and GoogleCredentials it holds.
  4. Otherwise a default client is built entirely from environment variables.

When ADK builds the client itself (cases 2 to 4), it does two more things. It adds tracking headers, x-goog-api-client and user-agent, with the value google-adk/<version> gl-java/<java version>. And it installs an OkHttp client whose connect, read and write timeouts are all Duration.ZERO, meaning no timeout. That suits long streamed generations, but it also means a stalled connection waits forever unless something above it imposes a deadline. The optional httpExecutorService sets the executor for the HTTP dispatcher threads, for example so a command-line JVM can exit or an application server can manage the threads. It is ignored when you pass your own client.

Backend selection and the Vertex trap

The google-genai Client decides the backend, and the rule is easy to get wrong. If the client was built with an explicit enterprise(boolean) or vertexAI(boolean) flag, that flag decides. Otherwise it reads GOOGLE_GENAI_USE_ENTERPRISE, then GOOGLE_GENAI_USE_VERTEXAI, and treats the value "true" as Vertex AI. If neither variable is set, it uses the Gemini API. API keys come from GOOGLE_API_KEY or GEMINI_API_KEY, and the former wins if both are set. Vertex project and location come from GOOGLE_CLOUD_PROJECT and GOOGLE_CLOUD_LOCATION.

Now the trap. ADK's Vertex path copies project, location and credentials onto the client builder but never sets the Vertex flag. And the Client constructor in 1.75.0 throws IllegalArgumentException: Gemini API does not support project/location. when a project or location is present but the backend resolved to the Gemini API. So Gemini.builder().vertexCredentials(...) works on a machine where GOOGLE_GENAI_USE_VERTEXAI=true is exported and fails at startup on one where it is not. The same code behaves differently on a laptop, in CI and in a container. The robust fix is to build the client yourself and say what you mean:

import com.google.adk.agents.LlmAgent;
import com.google.adk.models.Gemini;
import com.google.auth.oauth2.GoogleCredentials;
import com.google.genai.Client;
import com.google.genai.types.HttpOptions;

GoogleCredentials creds = GoogleCredentials.getApplicationDefault();

Client client = Client.builder()
    .vertexAI(true)                       // explicit: no dependence on env flags
    .project("my-project")
    .location("us-central1")
    .credentials(creds)
    .httpOptions(HttpOptions.builder()
        .timeout(60_000)                  // milliseconds; ADK's own clients set none
        .build())
    .build();

Gemini model = new Gemini("gemini-2.5-flash", client);

LlmAgent agent = LlmAgent.builder()
    .name("triage")
    .model(model)                         // the BaseLlm overload, not the String one
    .instruction("Classify the ticket and draft a reply.")
    .build();

Passing your own client has two side effects: ADK's tracking headers are not added, and the zero-timeout OkHttp client is not installed, so google-genai's own HTTP defaults apply. Both are usually what you want, but verify the timeout actually fires with a test against an endpoint that never answers. The model name above is an example; use whichever model your project has access to.

What happens to your request

Before any call, generateContent passes the request through GeminiUtil.prepareGenenerateContentRequest (the misspelling is in the source). Four rewrites can happen, and two depend on the backend:

RewriteWhenWhy it matters
Clear config.labelsGemini API backend onlyLabels are a Vertex AI feature; on the Gemini API they would fail the request. Cost attribution by label silently disappears if you switch backends.
Strip displayName from inline and file dataGemini API backend onlySame reason: the field is not accepted there.
Remove client-side function-call idsAlways, for ids ADK generatedIds ADK generated locally are not sent back to the model, because some APIs reject them; ids the model assigned are kept.
Append a user turnWhen the last content is not from the userThe model only answers a user turn, so ADK appends the text of CONTINUE_OUTPUT_MESSAGE.

The appended message reads "Continue output. DO NOT look at this line. ONLY look at the content before this line and system instruction." If you see that sentence in a trace, it was not your prompt or a prompt injection; the history ended with a model or tool turn. Thought parts are not stripped on this path (stripThoughts is false), so thought content and signatures from earlier turns are sent back as the model returned them.

The non-streaming path and paused generations

With stream=false, the class calls client.async.models.generateContent, which returns a CompletableFuture, and wraps it in Flowable.defer. The HTTP request starts when generateContent is called, not on subscription, so a Flowable you build and then drop has still sent, and been billed for, its request. The same detail breaks naive retries: resubscribing replays the same future, so .retry(3) on the returned Flowable never re-sends anything. Wrap the call instead: Flowable.defer(() -> model.generateContent(req, stream)).retry(3). The result is a single LlmResponse.

In 1.11.0 both paths also handle a paused generation. If the first candidate's finish reason is CONTINUATION and it carries a non-empty continuationToken, the class sends a follow-up request containing the original contents, the model output collected so far as a model turn, and the token in the config. It repeats until the model finishes, then returns one response with all parts joined and token usage summed across requests. Three guards stop the loop and return the partial output with a warning: a missing token, the same token twice (no progress), or MAX_RESUMES (256) resumes. Source code does not say which models or backends emit this finish reason, so do not build on it; just be aware that one logical call can be many HTTP requests, and that usage on the final response is the total.

The streaming path

With stream=true, the class calls generateContentStream and feeds each chunk to a StreamingResponseAggregator. Knowing its behavior saves debugging time:

  • Every chunk carrying content is re-emitted at once as an LlmResponse with partial=true, so a UI can render tokens as they arrive.
  • Text is buffered into runs. Consecutive text of the same kind is joined; a switch between thought text and answer text starts a new part, so reasoning never merges into the answer.
  • Function calls whose arguments are streamed across chunks (partialArgs or willContinue) are accumulated and appear in the final response as one complete call.
  • A function call that arrives without an id gets a client-generated id, reused in both the partial and final responses so consumers can correlate them. On the next request that id is removed again, as the table above shows.
  • Chunks with no content (metadata only) are not emitted as partials, and the empty-text part that ends a Gemini 3 stream is dropped.
  • When the stream ends, a final non-partial response carries the aggregated content, provided something accumulated or the last chunk had a finish reason; otherwise no final response is emitted. Only the final response should be persisted or used to run tools.

How the runner turns these into session events is covered in streaming events in ADK Java.

Live connections and the registry cache

connect applies only the Gemini API sanitizing step, then opens a live session with client.async.live.connect(model, liveConnectConfig), using the LiveConnectConfig carried on the request, and wraps it in a GeminiLlmConnection. The transport factory, connectLiveTransport, is a protected method, so a test subclass can supply an in-process double and exercise the real connection logic without a network.

Model strings go through LlmRegistry. Its built-in patterns map gemini-.* and gemma-.* to Gemini.builder().modelName(name).build() and apigee/.* to the Apigee model. getLlm caches with computeIfAbsent, so there is one instance per model name for the whole JVM, built from the environment-default client. Every agent naming "gemini-2.5-flash" shares one client, one set of credentials and the zero-timeout HTTP client. Per-tenant credentials, separate quotas or different timeouts therefore require constructing instances yourself, or registering your own factory with LlmRegistry.registerLlm. Patterns live in a hash map, so if two registered patterns match the same name, which factory wins is unspecified; keep patterns disjoint. The model registration guide goes deeper.

Worked example: tracing a tool-calling turn

To see the rewrites, subclass the class and log what goes in. The override sees the request before preparation, so it shows what the runtime built; the comments note what is actually sent.

import com.google.adk.models.Gemini;
import com.google.adk.models.LlmRequest;
import com.google.adk.models.LlmResponse;
import com.google.genai.Client;
import io.reactivex.rxjava3.core.Flowable;

public final class TracingGemini extends Gemini {
  public TracingGemini(String modelName, Client client) {
    super(modelName, client);
  }

  @Override
  public Flowable<LlmResponse> generateContent(LlmRequest request, boolean stream) {
    String model = request.model().orElse(model());
    long started = System.nanoTime();
    System.out.printf("-> %s turns=%d stream=%s%n", model, request.contents().size(), stream);
    return super.generateContent(request, stream)
        .doOnNext(r -> {
          if (!r.partial().orElse(false)) {
            System.out.printf("<- final in %d ms, usage=%s%n",
                (System.nanoTime() - started) / 1_000_000, r.usageMetadata().orElse(null));
          }
        });
  }
}

Trace one tool-calling turn on the Gemini API backend, with stream=true:

  1. The user asks for an order's status. The request holds one user turn; preparation leaves it alone apart from clearing any labels.
  2. The stream returns a function call get_order(id="A17") with no id. The aggregator assigns a client id, emits a partial, and the final response carries the same id. The runner executes the tool and records a function response with that id.
  3. Second call: the history ends with the tool's function response, which is not a user turn, so the continue message is appended. The client-generated id is removed from the call and its response before sending.
  4. The model streams its answer. Partials go to the UI; the final aggregated response is stored. Usage is logged once, from the final response.

If the logs show a startup IllegalArgumentException about project and location, you hit the backend trap. If a call logs its request and then nothing at all, neither partials nor an error, suspect the missing timeout; an empty stream can also end without a final response, so confirm with the HTTP client's own metrics.

Failure modes

FailureCause in the classFix
Startup error about project/locationVertex credentials without a Vertex flag in the environmentBuild the Client with vertexAI(true) and pass it in
Silent backend switchBuilder fell through to the env-default clientAlways set one of apiClient, apiKey or vertex credentials; log client.vertexAI() at boot
Calls hang indefinitelyADK-built clients have zero HTTP timeoutsOwn the Client and set HttpOptions.timeout, or add a Flowable timeout
Retries never re-sendThe request starts eagerly; resubscribing replays one futureRetry around Flowable.defer of the call
Labels vanish from billingLabels cleared on the Gemini API backendUse Vertex AI if you rely on labels
Strange "Continue output" text in tracesHistory ended with a non-user turnExpected; do not filter it as an attack
Tools run twice or on half a callActing on partial responsesRun tools only from the non-partial final response
Wrong credentials for a tenantRegistry caches one instance per name per JVMConstruct per-tenant instances; pass model(BaseLlm)

Trade-offs

Use model strings and the registry for prototypes and single-tenant services: it is one line and the environment decides everything. Construct Gemini explicitly in production, so the backend, credentials and timeout are in code and reviewed. Subclass it, as above, for tracing; for retries, rate limiting or fallbacks, prefer a decorator around BaseLlm, which composes across models. Write a fully custom BaseLlm only for a provider the class cannot reach; implementing a custom LLM explains how.

What to do next

  1. Grep your code for model strings and Gemini.builder() calls; list which backend each resolves to in each environment.
  2. Replace vertexCredentials(...) builders with an explicitly built Vertex Client.
  3. Set a request timeout and prove it fires against a black-hole endpoint.
  4. Log the effective model name, client.vertexAI() and final-response usage for every call.
  5. Make sure tools, persistence and billing read only non-partial responses.
  6. If you serve several tenants, stop sharing registry instances across them.
  7. Re-read Gemini.java and GeminiUtil.java whenever you upgrade adk-java; these internals change between releases.
Key takeaway: The ADK Java Gemini class wraps one google-genai Client, rewrites every request before sending it, aggregates streamed chunks into partials plus one final response, and resumes paused generations. In production, build the Client yourself with an explicit backend flag and a timeout, pass the instance to LlmAgent, act only on final responses, and remember that model strings share one cached instance per JVM.