Every model call an ADK Java agent makes goes through one small type: com.google.adk.models.BaseLlm. Gemini, Anthropic's Claude, models behind Apigee and anything bridged through LangChain4j all plug into the agent runtime by extending it. If you understand what the runtime asks of that type, and what it assumes about the answer, you can choose models confidently, wrap them with retries and metrics, and debug the odd behaviour that shows up only under streaming or concurrency.

This page is the reference from the runtime's side: who calls BaseLlm, when, with what, and what callers may rely on. Writing a new adapter step by step is covered in implementing a custom LLM; here the adapter is the thing being called. Names and behaviour were checked against the google/adk-java repository on 2026-10-01. The framework is young and still changing, so confirm against the version you build with.

Advertisement

An abstract class, not an interface

Despite the common name, BaseLlm is an abstract class. It holds one field, the model name, and declares two abstract methods. This is the whole surface:

package com.google.adk.models;

public abstract class BaseLlm {
  private final String model;

  public BaseLlm(String model) { this.model = model; }

  /** The model name, e.g. "gemini-2.5-flash". */
  public String model() { return model; }

  /** One request, one logical response. Non-streaming: one LlmResponse.
   *  Streaming: several LlmResponses that together form one content. */
  public abstract Flowable<LlmResponse> generateContent(LlmRequest llmRequest, boolean stream);

  /** A bidirectional live session, used for audio and video. */
  public abstract BaseLlmConnection connect(LlmRequest llmRequest);
}

Being a class has practical consequences. A subclass cannot extend anything else, so adapters wrap a vendor client rather than inherit from it. The constructor requires a model name, which the runtime uses for logging and which LlmRequest may also carry. And because there is no default implementation of either method, every subclass must decide what connect does, even if the answer is to throw UnsupportedOperationException for models without a live API.

The repository's models package ships Gemini, Claude and ApigeeLlm implementations; a LangChain4j bridge class lives in a separate contrib module and opens up the providers LangChain4j supports, including local models.

Where BaseLlm sits in a turn

A user message enters through the Runner, which drives the agent's flow. The flow runs request processors that assemble the LlmRequest: the agent's instructions, the conversation contents taken from session events, tool declarations and generation config. It then runs before-model callbacks, calls the model, runs after-model callbacks and turns the result into events. If a response contains function calls, the flow executes the tools, appends their results and loops, which is covered in model call orchestration.

RunnerrunAsync / runLiveLlmAgentmodel: instance or nameBaseLlmFlowrequest processors build LlmRequestBefore-model callbacksmay return a response: skip modelBaseLlm.generateContent(request, stream = mode is SSE)LlmRegistrypattern to factory, cachedresolve by nameAfter-model callbacksmay replace the responseEventspartial, final, tool callsconnect()BIDI: BaseLlmConnectionrunLiveGemini, Claude, ApigeeLlm and contrib adapters all sit behind the same two abstract methods.
The runtime's path to BaseLlm. Text and SSE streaming go through generateContent; live audio and video go through connect.

Two details from the flow's source are worth knowing. First, the stream argument is not a per-agent choice: the flow passes true exactly when the run's streaming mode is SSE. Second, the live path, runLive with bidirectional streaming, never calls generateContent; it calls connect and talks to the returned connection.

Advertisement

How a model is resolved

An LlmAgent can be given a model as a BaseLlm instance or as a string. When the flow needs the model, it uses the agent's instance if one is present; otherwise it looks up the name with LlmRegistry.getLlm(name). The registry maps regular-expression patterns to factories. Out of the box it registers gemini-.*, gemma-.* (both served by the Gemini class) and apigee/.*. You add your own:

// once, at startup, before any agent runs
LlmRegistry.registerLlm("acme-.*", modelName -> new AcmeLlm(modelName, acmeClient));

LlmAgent byName = LlmAgent.builder()
    .name("triage")
    .model("acme-large-2")                 // resolved through the registry
    .instruction("Classify the ticket.")
    .build();

LlmAgent byInstance = LlmAgent.builder()
    .name("support")
    .model(Gemini.builder()
        .modelName("gemini-2.5-flash")
        .apiKey(System.getenv("GOOGLE_API_KEY"))
        .build())                          // used directly, registry skipped
    .instruction("Answer support questions.")
    .build();

The registry's implementation has three consequences that are easy to miss. Lookups use computeIfAbsent on a concurrent map, so one instance per model name is created and then shared by every agent and thread that asks for it; your BaseLlm must be thread-safe and must not keep per-request state in fields. Registering a pattern again does not evict an instance already cached for a matching name, so register everything before the first call. And factories are held in a hash map, so when two patterns match the same name, which one wins is not guaranteed; keep patterns disjoint. A name that matches nothing throws IllegalArgumentException at first use, not at agent build time, so a typo in a model string surfaces on the first request.

What arrives in LlmRequest

By the time generateContent is called, the request is complete. Its accessors are model() (an Optional name), contents() (the conversation as a list of Content), config() (an Optional GenerateContentConfig holding system instruction, temperature, response schema and tool declarations), tools() (a map from tool name to BaseTool) and liveConnectConfig() for live sessions. Helpers getSystemInstructions() and getFirstSystemInstruction() extract instruction text.

Treat the request as read-only input. Before-model callbacks receive a builder and are the sanctioned place to modify a request; an adapter that mutates what it receives makes behaviour depend on call order. The types come from the google-genai Java library, so an adapter for another vendor is fundamentally a translation layer from these types to that vendor's wire format.

The generateContent contract

The return type is an RxJava 3 Flowable<LlmResponse>. The javadoc states the rule the runtime depends on: a non-streaming call yields one LlmResponse; a streaming call may yield several, and all of them together should be treated as one content by merging their parts. The bundled Gemini implementation does this by marking chunk responses partial and then emitting a final aggregated response that contains the accumulated text and the complete function calls.

In practice callers rely on the following, and adapters should honour it:

  • Cold and lazy. Nothing is sent until the Flowable is subscribed, and each subscription is a new call. Build the HTTP request inside Flowable.defer or a create callback, not in the method body.
  • One logical answer. Partial responses are for display. Tool calls should be acted on from complete, final content, never from a half-streamed argument string.
  • Cancellation works. When the subscriber disposes, for example because the user disconnected, the underlying HTTP stream should be closed.
  • Completion means done. onComplete follows the last response; an error terminates the stream and nothing follows it.
  • No blocking on the caller's thread. The runtime composes these Flowables; blocking I/O belongs on an I/O scheduler or the client's own executor. The trade-offs are discussed in sync versus async LLM contracts.

LlmResponse fields the runtime reads

AccessorMeaning
content()The model output: text parts, function calls, thoughts
partial()True for a streaming chunk that is not the complete content
turnComplete()The model has finished its turn (used in live sessions)
finishReason()Why generation stopped, for example stop, max tokens, safety
errorCode(), errorMessage()An in-band failure the model reported
usageMetadata()Token counts for cost and quota tracking
interrupted()The user interrupted a live response
groundingMetadata()Sources when grounding is used
inputTranscription(), outputTranscription()Speech transcripts in live audio sessions

All accessors return Optionals; build responses with LlmResponse.builder() or LlmResponse.create(GenerateContentResponse). Errors therefore arrive on two channels. Transport failures, timeouts, authentication errors and rate limits surface as Flowable.error. Model-level refusals and blocked content can arrive as a normal response with errorCode and errorMessage set. Code that only handles one channel will mishandle the other.

Live sessions: connect and BaseLlmConnection

For bidirectional audio and video the runtime calls connect(request) once and drives the returned BaseLlmConnection, whose six methods are:

Completable sendHistory(List<Content> history); // right after connecting; model answers if last turn is the user's
Completable sendContent(Content content);       // a user turn or function responses; model answers
Completable sendRealtime(Blob blob);            // audio chunk or video frame; model decides when to answer
Flowable<LlmResponse> receive();                // all model output for the session
void close();
void close(Throwable throwable);

The model may not answer a realtime blob immediately because it performs voice activity detection. Function responses sent with sendContent must contain only function-response parts. The receive stream carries partial text, transcriptions, turnComplete and interrupted flags, which the runtime turns into events. See streaming in ADK Java for the application side.

Worked example: a resilient, measured model

A team runs a support agent on Gemini and wants three things without touching agent code: latency and token metrics per call, a retry for transient failures, and no duplicated text when a stream fails midway. Because BaseLlm is the single seam, a decorator does it.

public final class ObservedLlm extends BaseLlm {
  private static final int MAX_ATTEMPTS = 3;
  private final BaseLlm delegate;
  private final Metrics metrics;

  public ObservedLlm(BaseLlm delegate, Metrics metrics) {
    super(delegate.model());
    this.delegate = delegate;
    this.metrics = metrics;
  }

  @Override
  public Flowable<LlmResponse> generateContent(LlmRequest request, boolean stream) {
    return Flowable.defer(() -> {
      long start = System.nanoTime();
      AtomicBoolean emitted = new AtomicBoolean(false);
      return Flowable.defer(() -> delegate.generateContent(request, stream))
          .doOnNext(r -> emitted.set(true))
          // retry only if nothing reached the caller yet: a resend after a partial would duplicate text
          .retry((attempt, error) -> !emitted.get() && attempt < MAX_ATTEMPTS && isTransient(error))
          .doOnNext(r -> r.usageMetadata().ifPresent(u -> metrics.recordUsage(model(), u)))
          .doOnError(e -> metrics.recordFailure(model(), e))
          .doFinally(() -> metrics.recordLatency(model(), System.nanoTime() - start));
    });
  }

  @Override
  public BaseLlmConnection connect(LlmRequest request) {
    return delegate.connect(request);   // delegate explicitly; live sessions are not retried here
  }

  private static boolean isTransient(Throwable e) {
    // classify by your client's exception types: timeouts, 429, 503 - never 400 or auth errors
    return e instanceof java.net.SocketTimeoutException || e instanceof java.io.IOException;
  }
}

BaseLlm gemini = Gemini.builder().modelName("gemini-2.5-flash")
    .apiKey(System.getenv("GOOGLE_API_KEY")).build();
LlmAgent agent = LlmAgent.builder().name("support")
    .model(new ObservedLlm(gemini, metrics)).instruction("...").build();

The outer defer gives each subscription its own timer and emitted flag; the inner one makes each retry a fresh call. The retry predicate shown has no backoff; in production use retryWhen with a timer, and respect any retry-after hint your client exposes. Usage is recorded per response, so with streaming you should record only from the final, non-partial response or you will double count, depending on how your delegate reports usage. Because the decorator is a BaseLlm, it can also be registered under a pattern so name-based agents get it too.

Decorator or callback?

ADK also offers before-model and after-model callbacks on the agent, and the two overlap. The flow runs before-model callbacks in order; if one returns a response, the model call is skipped and that response is used, which makes callbacks the natural place for guardrails, response caching keyed on the request, and request edits such as injecting context. After-model callbacks can replace a response. The details are in ADK Java callbacks.

ConcernBetter asWhy
Policy check, block a requestBefore-model callbackCan short-circuit with a canned response; has agent context
Edit instructions or contentsBefore-model callbackReceives a builder; the request stays immutable elsewhere
Retries, timeouts, circuit breakingBaseLlm decoratorMust wrap the actual network call and its errors
Latency and token metricsBaseLlm decoratorSees every call, including ones from callbacks' own flows
Model routing or fallbackBaseLlm decorator or registryAgents keep one model reference
Redacting outputAfter-model callbackWorks on the final response regardless of model

Failure modes

SymptomCauseFix
Duplicated or garbled streamed textRetry resubscribed after partials were emittedRetry only before the first emission
Request sent twiceFlowable built eagerly, or subscribed twiceBuild inside defer; share or cache if needed
Cross-talk between sessionsPer-request state in fields of a shared registry instanceKeep adapters stateless; state in local variables
IllegalArgumentException on first messageModel name matches no registered patternRegister at startup; smoke-test every model name
New factory ignoredInstance already cached for that nameRegister before first use; restart after changes
Tool call with broken JSON argumentsAdapter emitted a function call from a partial chunkEmit function calls only in the final aggregated response
Silent empty answersIn-band errorCode ignoredCheck errorCode and finishReason as well as Flowable errors
Live session hangsconnect delegated nowhere or connection never closedImplement or explicitly reject connect; close in finally

What to do next

  1. List every model name your agents use and confirm each resolves, either through an instance or a registered, disjoint pattern.
  2. Decide your streaming mode per run and test both SSE and non-streaming paths against the same agent.
  3. Wrap production models in a decorator for metrics and retry-before-first-emission, and delegate connect explicitly.
  4. Handle both error channels: Flowable errors and errorCode or finishReason on responses.
  5. Move policy checks and response caching into before-model callbacks, and keep the model wrapper free of agent logic.
  6. Write a scripted BaseLlm for tests that returns fixed responses, including partials and an in-band error, and assert the agent's events.
Key takeaway: BaseLlm is a small abstract class with a large job: generateContent returns a cold Flowable of LlmResponses that together form one answer, and connect opens a live session. The runtime resolves it from an instance or the regex-keyed, caching LlmRegistry, passes stream only in SSE mode, and may skip it entirely when a before-model callback answers. Keep implementations stateless and thread-safe, put retries and metrics in a decorator, and handle errors on both channels.