ADK Java promises that an agent can run on Gemini, Claude or any model LangChain4j can reach by changing one line: the model passed to LlmAgent.builder().model(...). That promise is mostly true for the shape of a conversation, meaning messages, system instructions and tool calls. It is not true for generation settings, structured output, streaming and live sessions. Those either pass through untouched, are translated, or are silently dropped, depending on the adapter. An abstraction boundary is the line between what the framework guarantees for every model and what it leaves to each provider. Knowing where that line sits in your version decides whether a model swap is a config change or an incident.

This page maps that boundary as it stands in the adk-java main branch, read on 2026-10-03. It covers the request and response types that cross it, a field-by-field table of what each shipped adapter honours, how to keep provider-specific settings out of agent code, startup capability checks, and a conformance test you can run against every model you use. The BaseLlm contract itself is covered in the BaseLLM interface overview and writing an adapter in implementing a custom LLM. Adapters change between releases, so treat the tables as a method for checking your own version, not as a permanent fact.

The types that cross the boundary

The boundary has one class and two value types. BaseLlm has a model name and two abstract methods, generateContent(LlmRequest, boolean stream) returning Flowable<LlmResponse> and connect(LlmRequest) for live sessions. LlmRequest carries model(), contents() as a List<Content>, config() as an Optional<GenerateContentConfig>, liveConnectConfig() and tools() as a Map<String, BaseTool>. LlmResponse carries optional content(), partial(), turnComplete(), finishReason(), errorCode(), errorMessage(), usageMetadata(), groundingMetadata() and a few more.

Look at where Content, GenerateContentConfig and the usage type come from: com.google.genai.types, the Google Gen AI SDK. The lingua franca of the abstraction is Gemini’s own wire model. For Gemini the adapter is close to a pass-through. Every other adapter is a translator from Gemini’s vocabulary into another provider’s, and a translator can only carry what it has words for. GenerateContentConfig has fields for safety settings, thinking configuration, cached content, response schema and media resolution; nothing forces a non-Gemini adapter to read them, and nothing tells you when it does not.

What each adapter actually honours

LlmAgentinstruction, tools, configLlmRequestgenai Content + ConfigBaseLlmgenerateContent / connectGemininear pass-throughClaudetranslates a subsetLangChain4jmaps sampling knobsconfig handed to the Gen AI SDKsystem text + first Tool onlyno response schemaYour boundary layercapability profile + startup check + conformance testsThe framework guarantees the conversation shape; you guarantee everything else.
Where the request crosses the boundary. One request type, three translations of different completeness.

The table below comes from reading each adapter’s request-mapping code on main. ‘Ignored’ means the code does not read the field; the request still succeeds, which is what makes these gaps dangerous.

Request fieldGeminiClaudeLangChain4j (contrib)
contents, function calls and responsesPassed throughTranslated to messagesTranslated to chat messages
config.systemInstructionPassed throughText parts joined into the system promptTranslated
Tool declarationsPassed throughFunction declarations of the first Tool onlyTranslated to tool specifications
temperature, topP, topKPassed throughIgnoredMapped
maxOutputTokensPassed throughIgnored; constructor value, default 8192Mapped
stopSequences, frequency and presence penaltiesPassed throughIgnoredMapped
toolConfig function-calling modePassed throughIgnored; tool choice is auto with parallel tool use disabledAUTO, ANY (with allowed names) and NONE mapped
responseSchema, responseMimeTypePassed throughIgnoredNot mapped
stream = trueStreamsIgnored; one complete responseNeeds a StreamingChatModel, else an error
connect() live sessionSupportedThrows UnsupportedOperationExceptionThrows UnsupportedOperationException

Two rows deserve emphasis. A Claude-backed agent configured with temperature 0 for reproducible extraction runs at the provider’s default temperature, and an output cap set through config is replaced by the adapter’s own. And an agent that relies on responseSchema for JSON output gets free text from a non-Gemini adapter that does not map it. ADK has its own agent-level outputSchema, but it still has to reach the provider as a request field the adapter understands. Check what your version does before relying on it with a non-Gemini model.

Portable, verified and provider-specific settings

The stub this page replaces suggested exposing generation controls through the abstraction. The table shows why that is only half right: the controls are exposed, but enforcement depends on the adapter. A useful rule is to split request settings into three classes.

  • Portable. Must mean the same thing on every model you deploy: conversation contents, system instruction, tool declarations and tool results. Treat a model that cannot carry these as unsupported.
  • Portable with verification. Exist on most providers but are translated by adapters: temperature, top-p, output length, stop sequences, tool choice, structured output, streaming. Set them through GenerateContentConfig, and verify at startup that the chosen adapter honours each one you depend on.
  • Provider-specific. Have no equivalent elsewhere: Gemini safety settings, thinking budgets, cached content, media resolution, Anthropic-only beta features. Keep them out of shared agent definitions and attach them in a provider profile or a custom adapter, so they cannot leak into a request for the wrong model.

A capability layer in code

Put a thin layer between agent definitions and adapters: a capability profile per model, a factory that builds the adapter and config together, and a startup check that refuses combinations the profile cannot honour. A sketch, using real ADK and Gen AI types; the profile classes are yours:

enum Cap { SAMPLING_PARAMS, MAX_OUTPUT_TOKENS, RESPONSE_SCHEMA, STREAMING, LIVE, MULTIPLE_TOOL_OBJECTS }

record ModelProfile(String name, Set<Cap> caps, Supplier<BaseLlm> factory) {}

final class ModelProfiles {
  static final Map<String, ModelProfile> PROFILES = Map.of(
      "gemini-2.5-flash", new ModelProfile("gemini-2.5-flash", EnumSet.allOf(Cap.class),
          () -> Gemini.builder().modelName("gemini-2.5-flash").apiKey(System.getenv("GOOGLE_API_KEY")).build()),
      "claude-sonnet", new ModelProfile("claude-sonnet", EnumSet.noneOf(Cap.class),
          () -> new Claude(System.getenv("CLAUDE_MODEL"), anthropicClient(), 4096)));

  static BaseLlm checked(String name, GenerateContentConfig cfg, boolean streaming) {
    ModelProfile p = PROFILES.get(name);
    List<String> missing = new ArrayList<>();
    if ((cfg.temperature().isPresent() || cfg.topP().isPresent()) && !p.caps().contains(Cap.SAMPLING_PARAMS))
      missing.add("sampling parameters");
    if (cfg.maxOutputTokens().isPresent() && !p.caps().contains(Cap.MAX_OUTPUT_TOKENS))
      missing.add("maxOutputTokens (set it on the adapter instead)");
    if (cfg.responseSchema().isPresent() && !p.caps().contains(Cap.RESPONSE_SCHEMA))
      missing.add("responseSchema");
    if (streaming && !p.caps().contains(Cap.STREAMING))
      missing.add("streaming");
    if (!missing.isEmpty())
      throw new IllegalStateException(name + " does not honour: " + missing);
    return p.factory().get();
  }
}

LlmAgent agent = LlmAgent.builder()
    .name("extractor")
    .model(ModelProfiles.checked(modelName, cfg, false))
    .generateContentConfig(cfg)
    .instruction("Extract invoice fields.")
    .build();

Passing a BaseLlm instance rather than a string makes the factory the only place a model is constructed. String names resolve through LlmRegistry, which on main pre-registers the patterns gemini-.*, apigee/.* and gemma-.*; anything else needs LlmRegistry.registerLlm(pattern, factory) and gives you no place to run the check. Fill the capability sets by reading each adapter’s source for the version you build with, then confirm them with the conformance test below. Do not copy them from this page.

Leaks on the way back

The boundary leaks on the way back too. Agent code, callbacks and observability read fields of LlmResponse, and adapters fill different subsets. On main the Claude adapter maps token usage into usageMetadata() but does not set finishReason(), so a callback that checks for a length cut-off never fires on Claude. Grounding metadata exists only where the provider produces it. Error reporting differs as well: some failures arrive as an errorCode() on a response, others as an error signal on the Flowable, and a raw provider exception type can pass through untranslated.

Normalise at one point. An after-model callback or a wrapping BaseLlm decorator can map provider failures to your own error categories and fill a consistent set of fields your code reads, such as an output-truncated flag and token counts, defaulting where the adapter is silent. Code beyond that point should read only the normalised view. Keep the raw response available in debug logs.

There is a threading leak as well. The Claude adapter on main calls the Anthropic client synchronously inside generateContent and wraps the finished message in Flowable.just, so the blocking call happens when the method is invoked, on the caller’s thread, not when the stream is subscribed. A reactive pipeline that assumes model calls are lazy and non-blocking will stall a scheduler thread for the whole model latency. Sync versus async LLM contracts shows how to move such calls onto an I/O scheduler.

Conformance tests for every adapter

A capability table is a claim, and a conformance test checks it. Write one parameterised test that runs against every adapter you deploy, with a small paid budget, in CI before upgrades:

@ParameterizedTest
@MethodSource("models")
void honoursOutputCap(BaseLlm llm) {
  LlmRequest req = LlmRequest.builder()
      .contents(List.of(Content.fromParts(Part.fromText("Count from 1 to 500, comma separated."))))
      .config(GenerateContentConfig.builder().maxOutputTokens(20).build())
      .build();
  LlmResponse r = llm.generateContent(req, false).blockingLast();
  int out = r.usageMetadata().flatMap(u -> u.candidatesTokenCount()).orElse(-1);
  assertTrue(out > 0 && out <= 25, llm.model() + " produced " + out + " tokens");
}

Write similar probes for each capability you mark: a schema request should parse as JSON matching the schema; temperature 0 with the same prompt five times should agree on a short factual answer far more often than temperature 1.5; streaming should deliver more than one partial response for a long answer; two tool objects should both be visible to the model. Probes are statistical for sampling settings, so assert on clear differences, not exact values. A failing probe after an ADK upgrade is the signal to update the profile, which is far better than discovering the change in production output.

Failure modes

Silent config drop. The model swap works, tests on happy paths pass, and output length, determinism or format change quietly. Prevent it with the startup check and conformance tests.

Provider-only settings on the wrong model. Gemini safety settings or thinking config placed in a shared config are ignored by other adapters, or worse, a custom adapter passes them through and the provider rejects the request. Keep them in provider profiles.

Tools split across Tool objects. An adapter that reads only the first Tool’s function declarations hides the rest from the model, which then never calls them. Merge function declarations into one Tool when targeting such an adapter, and test that every tool is reachable.

Streaming assumptions. A UI expects partial responses and receives one final response after the whole latency, or a stream request errors because no streaming model was configured. Make streaming a declared capability and fall back explicitly.

Live features on the wrong model. connect() throws on adapters without a live API. Route live agents only to models whose profile includes it.

Leaky errors. Retry logic written for Gemini error shapes does not recognise another provider’s rate-limit exception and gives up, or retries non-retryable failures. Normalise errors at the boundary; model call orchestration shows where errors and retries sit in a turn.

Operations and trade-offs

Pin and review. Pin ADK and contrib versions, and when you upgrade, diff the adapters’ request-mapping code. The gaps above are the kind of thing adapters fix release by release, and a fix can change production behaviour as surely as a bug can.

Record the effective request. Log the provider-level request each adapter actually sends, at debug level with redaction, alongside the ADK request. The difference between the two is the boundary, made visible, and it shortens most model-swap investigations to minutes.

Trade-offs. A strict capability layer slows down trying a new model, because each one needs a profile and probes. Accepting the framework’s abstraction as-is is faster and fine for prototypes where only the conversation shape matters. Writing your own adapter gives full control and full maintenance cost. For production agents with format, length or determinism requirements, the profile and probe layer is cheap insurance.

What to do next

  1. List every GenerateContentConfig field and agent-level setting your agents rely on, and why.
  2. For each model you deploy, read its adapter’s request mapping in your ADK version and fill a capability profile.
  3. Construct models through one factory that checks the profile against the config at startup and fails fast.
  4. Move provider-only settings, such as safety, thinking and cached content, into per-provider profiles.
  5. Add a normalising after-model callback or decorator for finish reason, usage and errors.
  6. Write conformance probes for output cap, schema, sampling, streaming and multiple tools, and run them on every ADK upgrade.
  7. Log the effective provider request at debug level so model swaps can be diffed.
Key takeaway: ADK Java's model abstraction guarantees the conversation shape, meaning contents, system instruction and tools, but expresses everything in Gemini's genai types, and other adapters translate only part of it. On main the Claude adapter ignores sampling settings, output caps, response schema and streaming, and the LangChain4j adapter maps sampling knobs but not response schema. Classify settings as portable, verified or provider-specific, build models through a factory that checks a capability profile at startup, normalise responses and errors at one point, and run conformance probes on every upgrade.