ADK Java promises that an agent can run on Gemini, Claude or any model LangChain4j can reach by changing one line: the model passed to LlmAgent.builder().model(...). That promise is mostly true for the shape of a conversation, meaning messages, system instructions and tool calls. It is not true for generation settings, structured output, streaming and live sessions. Those either pass through untouched, are translated, or are silently dropped, depending on the adapter. An abstraction boundary is the line between what the framework guarantees for every model and what it leaves to each provider. Knowing where that line sits in your version decides whether a model swap is a config change or an incident.
This page maps that boundary as it stands in the adk-java main branch, read on 2026-10-03. It covers the request and response types that cross it, a field-by-field table of what each shipped adapter honours, how to keep provider-specific settings out of agent code, startup capability checks, and a conformance test you can run against every model you use. The BaseLlm contract itself is covered in the BaseLLM interface overview and writing an adapter in implementing a custom LLM. Adapters change between releases, so treat the tables as a method for checking your own version, not as a permanent fact.
The types that cross the boundary
The boundary has one class and two value types. BaseLlm has a model name and two abstract methods, generateContent(LlmRequest, boolean stream) returning Flowable<LlmResponse> and connect(LlmRequest) for live sessions. LlmRequest carries model(), contents() as a List<Content>, config() as an Optional<GenerateContentConfig>, liveConnectConfig() and tools() as a Map<String, BaseTool>. LlmResponse carries optional content(), partial(), turnComplete(), finishReason(), errorCode(), errorMessage(), usageMetadata(), groundingMetadata() and a few more.
Look at where Content, GenerateContentConfig and the usage type come from: com.google.genai.types, the Google Gen AI SDK. The lingua franca of the abstraction is Gemini’s own wire model. For Gemini the adapter is close to a pass-through. Every other adapter is a translator from Gemini’s vocabulary into another provider’s, and a translator can only carry what it has words for. GenerateContentConfig has fields for safety settings, thinking configuration, cached content, response schema and media resolution; nothing forces a non-Gemini adapter to read them, and nothing tells you when it does not.
What each adapter actually honours
The table below comes from reading each adapter’s request-mapping code on main. ‘Ignored’ means the code does not read the field; the request still succeeds, which is what makes these gaps dangerous.
| Request field | Gemini | Claude | LangChain4j (contrib) |
|---|---|---|---|
| contents, function calls and responses | Passed through | Translated to messages | Translated to chat messages |
| config.systemInstruction | Passed through | Text parts joined into the system prompt | Translated |
| Tool declarations | Passed through | Function declarations of the first Tool only | Translated to tool specifications |
| temperature, topP, topK | Passed through | Ignored | Mapped |
| maxOutputTokens | Passed through | Ignored; constructor value, default 8192 | Mapped |
| stopSequences, frequency and presence penalties | Passed through | Ignored | Mapped |
| toolConfig function-calling mode | Passed through | Ignored; tool choice is auto with parallel tool use disabled | AUTO, ANY (with allowed names) and NONE mapped |
| responseSchema, responseMimeType | Passed through | Ignored | Not mapped |
| stream = true | Streams | Ignored; one complete response | Needs a StreamingChatModel, else an error |
| connect() live session | Supported | Throws UnsupportedOperationException | Throws UnsupportedOperationException |
Two rows deserve emphasis. A Claude-backed agent configured with temperature 0 for reproducible extraction runs at the provider’s default temperature, and an output cap set through config is replaced by the adapter’s own. And an agent that relies on responseSchema for JSON output gets free text from a non-Gemini adapter that does not map it. ADK has its own agent-level outputSchema, but it still has to reach the provider as a request field the adapter understands. Check what your version does before relying on it with a non-Gemini model.
Portable, verified and provider-specific settings
The stub this page replaces suggested exposing generation controls through the abstraction. The table shows why that is only half right: the controls are exposed, but enforcement depends on the adapter. A useful rule is to split request settings into three classes.
- Portable. Must mean the same thing on every model you deploy: conversation contents, system instruction, tool declarations and tool results. Treat a model that cannot carry these as unsupported.
- Portable with verification. Exist on most providers but are translated by adapters: temperature, top-p, output length, stop sequences, tool choice, structured output, streaming. Set them through
GenerateContentConfig, and verify at startup that the chosen adapter honours each one you depend on. - Provider-specific. Have no equivalent elsewhere: Gemini safety settings, thinking budgets, cached content, media resolution, Anthropic-only beta features. Keep them out of shared agent definitions and attach them in a provider profile or a custom adapter, so they cannot leak into a request for the wrong model.
A capability layer in code
Put a thin layer between agent definitions and adapters: a capability profile per model, a factory that builds the adapter and config together, and a startup check that refuses combinations the profile cannot honour. A sketch, using real ADK and Gen AI types; the profile classes are yours:
enum Cap { SAMPLING_PARAMS, MAX_OUTPUT_TOKENS, RESPONSE_SCHEMA, STREAMING, LIVE, MULTIPLE_TOOL_OBJECTS }
record ModelProfile(String name, Set<Cap> caps, Supplier<BaseLlm> factory) {}
final class ModelProfiles {
static final Map<String, ModelProfile> PROFILES = Map.of(
"gemini-2.5-flash", new ModelProfile("gemini-2.5-flash", EnumSet.allOf(Cap.class),
() -> Gemini.builder().modelName("gemini-2.5-flash").apiKey(System.getenv("GOOGLE_API_KEY")).build()),
"claude-sonnet", new ModelProfile("claude-sonnet", EnumSet.noneOf(Cap.class),
() -> new Claude(System.getenv("CLAUDE_MODEL"), anthropicClient(), 4096)));
static BaseLlm checked(String name, GenerateContentConfig cfg, boolean streaming) {
ModelProfile p = PROFILES.get(name);
List<String> missing = new ArrayList<>();
if ((cfg.temperature().isPresent() || cfg.topP().isPresent()) && !p.caps().contains(Cap.SAMPLING_PARAMS))
missing.add("sampling parameters");
if (cfg.maxOutputTokens().isPresent() && !p.caps().contains(Cap.MAX_OUTPUT_TOKENS))
missing.add("maxOutputTokens (set it on the adapter instead)");
if (cfg.responseSchema().isPresent() && !p.caps().contains(Cap.RESPONSE_SCHEMA))
missing.add("responseSchema");
if (streaming && !p.caps().contains(Cap.STREAMING))
missing.add("streaming");
if (!missing.isEmpty())
throw new IllegalStateException(name + " does not honour: " + missing);
return p.factory().get();
}
}
LlmAgent agent = LlmAgent.builder()
.name("extractor")
.model(ModelProfiles.checked(modelName, cfg, false))
.generateContentConfig(cfg)
.instruction("Extract invoice fields.")
.build();Passing a BaseLlm instance rather than a string makes the factory the only place a model is constructed. String names resolve through LlmRegistry, which on main pre-registers the patterns gemini-.*, apigee/.* and gemma-.*; anything else needs LlmRegistry.registerLlm(pattern, factory) and gives you no place to run the check. Fill the capability sets by reading each adapter’s source for the version you build with, then confirm them with the conformance test below. Do not copy them from this page.
Leaks on the way back
The boundary leaks on the way back too. Agent code, callbacks and observability read fields of LlmResponse, and adapters fill different subsets. On main the Claude adapter maps token usage into usageMetadata() but does not set finishReason(), so a callback that checks for a length cut-off never fires on Claude. Grounding metadata exists only where the provider produces it. Error reporting differs as well: some failures arrive as an errorCode() on a response, others as an error signal on the Flowable, and a raw provider exception type can pass through untranslated.
Normalise at one point. An after-model callback or a wrapping BaseLlm decorator can map provider failures to your own error categories and fill a consistent set of fields your code reads, such as an output-truncated flag and token counts, defaulting where the adapter is silent. Code beyond that point should read only the normalised view. Keep the raw response available in debug logs.
There is a threading leak as well. The Claude adapter on main calls the Anthropic client synchronously inside generateContent and wraps the finished message in Flowable.just, so the blocking call happens when the method is invoked, on the caller’s thread, not when the stream is subscribed. A reactive pipeline that assumes model calls are lazy and non-blocking will stall a scheduler thread for the whole model latency. Sync versus async LLM contracts shows how to move such calls onto an I/O scheduler.
Conformance tests for every adapter
A capability table is a claim, and a conformance test checks it. Write one parameterised test that runs against every adapter you deploy, with a small paid budget, in CI before upgrades:
@ParameterizedTest
@MethodSource("models")
void honoursOutputCap(BaseLlm llm) {
LlmRequest req = LlmRequest.builder()
.contents(List.of(Content.fromParts(Part.fromText("Count from 1 to 500, comma separated."))))
.config(GenerateContentConfig.builder().maxOutputTokens(20).build())
.build();
LlmResponse r = llm.generateContent(req, false).blockingLast();
int out = r.usageMetadata().flatMap(u -> u.candidatesTokenCount()).orElse(-1);
assertTrue(out > 0 && out <= 25, llm.model() + " produced " + out + " tokens");
}Write similar probes for each capability you mark: a schema request should parse as JSON matching the schema; temperature 0 with the same prompt five times should agree on a short factual answer far more often than temperature 1.5; streaming should deliver more than one partial response for a long answer; two tool objects should both be visible to the model. Probes are statistical for sampling settings, so assert on clear differences, not exact values. A failing probe after an ADK upgrade is the signal to update the profile, which is far better than discovering the change in production output.
Failure modes
Silent config drop. The model swap works, tests on happy paths pass, and output length, determinism or format change quietly. Prevent it with the startup check and conformance tests.
Provider-only settings on the wrong model. Gemini safety settings or thinking config placed in a shared config are ignored by other adapters, or worse, a custom adapter passes them through and the provider rejects the request. Keep them in provider profiles.
Tools split across Tool objects. An adapter that reads only the first Tool’s function declarations hides the rest from the model, which then never calls them. Merge function declarations into one Tool when targeting such an adapter, and test that every tool is reachable.
Streaming assumptions. A UI expects partial responses and receives one final response after the whole latency, or a stream request errors because no streaming model was configured. Make streaming a declared capability and fall back explicitly.
Live features on the wrong model. connect() throws on adapters without a live API. Route live agents only to models whose profile includes it.
Leaky errors. Retry logic written for Gemini error shapes does not recognise another provider’s rate-limit exception and gives up, or retries non-retryable failures. Normalise errors at the boundary; model call orchestration shows where errors and retries sit in a turn.
Operations and trade-offs
Pin and review. Pin ADK and contrib versions, and when you upgrade, diff the adapters’ request-mapping code. The gaps above are the kind of thing adapters fix release by release, and a fix can change production behaviour as surely as a bug can.
Record the effective request. Log the provider-level request each adapter actually sends, at debug level with redaction, alongside the ADK request. The difference between the two is the boundary, made visible, and it shortens most model-swap investigations to minutes.
Trade-offs. A strict capability layer slows down trying a new model, because each one needs a profile and probes. Accepting the framework’s abstraction as-is is faster and fine for prototypes where only the conversation shape matters. Writing your own adapter gives full control and full maintenance cost. For production agents with format, length or determinism requirements, the profile and probe layer is cheap insurance.
What to do next
- List every
GenerateContentConfigfield and agent-level setting your agents rely on, and why. - For each model you deploy, read its adapter’s request mapping in your ADK version and fill a capability profile.
- Construct models through one factory that checks the profile against the config at startup and fails fast.
- Move provider-only settings, such as safety, thinking and cached content, into per-provider profiles.
- Add a normalising after-model callback or decorator for finish reason, usage and errors.
- Write conformance probes for output cap, schema, sampling, streaming and multiple tools, and run them on every ADK upgrade.
- Log the effective provider request at debug level so model swaps can be diffed.