ADK Java ships adapters for Gemini, for Anthropic's Claude and for models behind Apigee, and its contrib directory has LangChain4j and Spring AI integration modules. Sooner or later a team needs something else: a model served by vLLM or another server exposing an OpenAI-compatible endpoint, an internal gateway with its own authentication and audit headers, or a deterministic fake for tests. ADK supports all three through one abstraction, the BaseLlm class.

Writing the adapter is mostly translation, but the details decide whether multi-turn tool use works. This article reads the contract from the ADK source, builds an adapter for an OpenAI-compatible chat completions endpoint, traces one tool-calling turn through it, and covers registration, streaming, errors, thread safety and testing. How the agent flow calls a model, including callbacks and call budgets, is covered in model call orchestration; this page is about the adapter itself.

Advertisement

First, check whether you need one

A custom adapter is code you will maintain against two moving targets: ADK's request types and the provider's API. Before writing one, check whether the contrib LangChain4j or Spring AI modules already reach your provider, and whether your gateway can present a protocol an existing adapter speaks. Write your own when you need exact control over the wire format, headers, retries or cost accounting, or when no integration exists. The rest of this article assumes you do.

The contract

package com.google.adk.models;

public abstract class BaseLlm {
  public BaseLlm(String model) { ... }
  public String model() { ... }

  // One model turn. stream == true asks for partial responses as they arrive.
  public abstract Flowable<LlmResponse> generateContent(LlmRequest llmRequest, boolean stream);

  // Bidirectional live session (audio, realtime). Optional for most adapters.
  public abstract BaseLlmConnection connect(LlmRequest llmRequest);
}

The constructor takes the model name the agent refers to, and model() returns it. generateContent receives an LlmRequest and returns an RxJava Flowable<LlmResponse>. The request carries contents(), the conversation as a list of genai Content objects; config(), an optional GenerateContentConfig holding tool declarations and sampling settings; getSystemInstructions(), the instruction strings the flow appended; and tools(), a map of the ADK tools themselves.

A Content has a role, which ADK sets to user or model, and a list of Part objects. A part is text, a FunctionCall the model made, a FunctionResponse carrying a tool's result, or media. An LlmResponse carries one Content with role model, plus optional partial, turnComplete, finishReason, errorCode, errorMessage and usageMetadata.

connect opens a live, bidirectional session and returns a BaseLlmConnection with sendHistory, sendContent, sendRealtime, receive and close. It is used for realtime audio and streaming-input agents. The built-in Claude adapter throws UnsupportedOperationException from it, and most text-only adapters should do the same rather than half-implement it.

Where a custom BaseLlm sits: the flow builds an LlmRequest, your adapter owns the wire formatLlmAgentmodel("local/qwen")BaseLlmFlowrequest processorsLlmRegistrypattern -> factory, cachedOpenAiCompatibleLlmextends BaseLlmRequest mapperContent, Part, tools -> JSONResponse mapperJSON -> LlmResponseModel server/v1/chat/completionsgetLlm(name)generateContentHTTP POSTJSON / SSEFlowable of LlmResponseThe adapter is one shared instance per model name: keep it stateless, thread-safe and non-blocking for the caller.
The agent names a model; the registry maps the name to your adapter; the adapter translates LlmRequest to the provider's JSON and the reply back to LlmResponse.
Advertisement

The adapter skeleton

public final class OpenAiCompatibleLlm extends BaseLlm {
  private static final ObjectMapper JSON = new ObjectMapper();   // thread-safe once built
  private final HttpClient http;                                 // thread-safe, pooled
  private final URI endpoint;
  private final String apiKey;
  private final String wireModel;                                // name the server expects

  public OpenAiCompatibleLlm(String adkName, URI baseUrl, String apiKey, HttpClient http) {
    super(adkName);                                              // e.g. "local/qwen2.5-32b"
    this.wireModel = adkName.substring(adkName.indexOf('/') + 1);
    this.endpoint = baseUrl.resolve("/v1/chat/completions");
    this.apiKey = apiKey;
    this.http = http;
  }

  @Override
  public Flowable<LlmResponse> generateContent(LlmRequest req, boolean stream) {
    // Cold and lazy: one HTTP call per subscription, off the caller's thread.
    return Flowable.fromCallable(() -> call(req))
        .subscribeOn(Schedulers.io());
  }

  private LlmResponse call(LlmRequest req) throws Exception {
    ObjectNode body = RequestMapper.toChatCompletion(wireModel, req);
    HttpRequest httpReq = HttpRequest.newBuilder(endpoint)
        .timeout(Duration.ofSeconds(60))
        .header("Content-Type", "application/json")
        .header("Authorization", "Bearer " + apiKey)
        .POST(HttpRequest.BodyPublishers.ofString(JSON.writeValueAsString(body)))
        .build();
    HttpResponse<String> resp = http.send(httpReq, HttpResponse.BodyHandlers.ofString());
    if (resp.statusCode() == 429 || resp.statusCode() >= 500) {
      throw new RetryableModelException(resp.statusCode(), resp.body());
    }
    if (resp.statusCode() != 200) {
      throw new ModelRequestException(resp.statusCode(), resp.body());
    }
    return ResponseMapper.fromChatCompletion(JSON.readTree(resp.body()));
  }

  @Override
  public BaseLlmConnection connect(LlmRequest req) {
    throw new UnsupportedOperationException("Live sessions are not supported by " + model());
  }
}

Three decisions here matter more than they look. The Flowable is cold and lazy: fromCallable performs one HTTP call per subscription, so a retry operator upstream re-sends the request instead of replaying a cached failure. The blocking call runs on Schedulers.io(), so it never ties up the thread that subscribed. And the class holds no per-request state: LlmRegistry caches one instance per model name, so a single object serves every concurrent invocation in the process. HttpClient and a configured ObjectMapper are safe to share; a mutable field holding the current conversation is not.

The two exception types are your own. Separating retryable failures, such as rate limits and server errors, from permanent ones, such as a malformed request, lets a retry or circuit-breaker layer act on the type. See circuit breakers in ADK Java for the wrapping pattern.

Mapping the request

static ObjectNode toChatCompletion(String wireModel, LlmRequest req) throws JsonProcessingException {
  ObjectNode body = JSON.createObjectNode().put("model", wireModel);
  ArrayNode messages = body.putArray("messages");

  String system = String.join("\n\n", req.getSystemInstructions());
  if (!system.isBlank()) {
    messages.addObject().put("role", "system").put("content", system);
  }

  for (Content content : req.contents()) {
    String role = content.role().orElse("user");               // ADK uses "user" and "model"
    List<Part> parts = content.parts().orElse(List.of());
    ObjectNode assistant = null;
    ArrayNode calls = null;
    StringBuilder text = new StringBuilder();
    for (Part part : parts) {
      part.text().ifPresent(text::append);
      if (part.functionCall().isPresent()) {                    // model asked for a tool
        FunctionCall fc = part.functionCall().get();
        if (assistant == null) {
          assistant = JSON.createObjectNode().put("role", "assistant");
          calls = assistant.putArray("tool_calls");
        }
        ObjectNode call = calls.addObject();
        call.put("id", fc.id().orElseThrow()).put("type", "function");
        call.putObject("function")
            .put("name", fc.name().orElseThrow())
            .put("arguments", JSON.writeValueAsString(fc.args().orElse(Map.of())));
      }
      if (part.functionResponse().isPresent()) {                // tool result from ADK
        FunctionResponse fr = part.functionResponse().get();
        messages.addObject().put("role", "tool")
            .put("tool_call_id", fr.id().orElseThrow())
            .put("content", JSON.writeValueAsString(fr.response().orElse(Map.of())));
      }
    }
    if (assistant != null) {
      if (text.length() > 0) assistant.put("content", text.toString());
      messages.add(assistant);
    } else if (text.length() > 0) {
      messages.addObject().put("role", role.equals("model") ? "assistant" : "user")
          .put("content", text.toString());
    }
  }

  req.config().flatMap(GenerateContentConfig::tools).ifPresent(tools -> {
    ArrayNode out = body.putArray("tools");
    for (Tool tool : tools) {
      for (FunctionDeclaration fd : tool.functionDeclarations().orElse(List.of())) {
        ObjectNode fn = out.addObject().put("type", "function").putObject("function");
        fn.put("name", fd.name().orElseThrow());
        fd.description().ifPresent(d -> fn.put("description", d));
        // schemaToJson: your own converter from genai Schema to JSON Schema.
        // Check field names against the google-genai version you depend on.
        fd.parameters().ifPresent(s -> fn.set("parameters", schemaToJson(s)));
      }
    }
  });
  return body;
}

Roles are the first trap. ADK says model where OpenAI-style APIs say assistant, and a model turn that contains function calls must become a single assistant message with a tool_calls array, not one message per call. Each FunctionResponse becomes its own tool message carrying the id of the call it answers. Parallel tool calls therefore round-trip only if ids are preserved exactly.

System instructions arrive as a list because several processors may append them; join them into one system message. Tool declarations come from config().tools() as FunctionDeclaration objects with a name, description and parameter Schema, which you convert to the JSON Schema the provider expects. That is the same place the built-in Claude adapter reads them. Two things are deliberately not shown: media parts, which need provider-specific encoding, and sampling settings such as temperature in GenerateContentConfig, which map field by field.

Mapping the response

static LlmResponse fromChatCompletion(JsonNode json) throws Exception {
  JsonNode msg = json.path("choices").path(0).path("message");
  List<Part> parts = new ArrayList<>();
  if (msg.hasNonNull("content") && !msg.get("content").asText().isEmpty()) {
    parts.add(Part.fromText(msg.get("content").asText()));
  }
  for (JsonNode call : msg.path("tool_calls")) {
    String id = call.path("id").asText("");
    if (id.isEmpty()) id = "call-" + UUID.randomUUID();         // ADK pairs results by id
    Map<String, Object> args = JSON.readValue(
        call.path("function").path("arguments").asText("{}"), MAP_TYPE);  // a JSON string
    parts.add(Part.builder().functionCall(FunctionCall.builder()
        .id(id).name(call.path("function").path("name").asText()).args(args).build()).build());
  }
  LlmResponse.Builder out = LlmResponse.builder()
      .content(Content.builder().role("model").parts(parts).build());
  JsonNode usage = json.path("usage");
  if (!usage.isMissingNode()) {
    out.usageMetadata(GenerateContentResponseUsageMetadata.builder()
        .promptTokenCount(usage.path("prompt_tokens").asInt())
        .candidatesTokenCount(usage.path("completion_tokens").asInt())
        .totalTokenCount(usage.path("total_tokens").asInt())
        .build());
  }
  return out.build();
}

OpenAI-style APIs return tool arguments as a JSON string, while FunctionCall.args is a map, so parse it. A model that emits invalid JSON arguments should surface as an error you can see, not a silently empty map. If the provider omits call ids, generate one with your own prefix. ADK's flow copies each call's id onto the matching function response, and fills blank ids with ids prefixed adk-. The Gemini adapter strips adk- ids from outgoing requests; do not copy that step, because an OpenAI-style server needs every id echoed back. Mapping usage into usageMetadata is what makes token accounting, cost tracking and budget enforcement work downstream; an adapter that drops it makes every dashboard read zero.

Map the provider's finish reason too. A reply cut off by the token limit looks like a complete answer unless you record it, and a safety refusal should reach the caller as a clear error or finish reason rather than as an empty text part.

Worked example: one tool-calling turn

A user asks the billing agent about invoice 4411. The flow sends contents with one user text part, a system instruction, and one declaration, lookupInvoice(invoiceId). The adapter sends a system message, a user message and a tools array. The server replies with no text and one tool call, id call_7f2, arguments {"invoiceId":"4411"}. The adapter returns an LlmResponse whose model content has one function-call part with that id and a parsed argument map.

ADK sees the function call, runs the Java tool, and appends the model's function-call content and a user-role content holding a FunctionResponse to the session. The flow calls generateContent again. Now the adapter emits an assistant message with a tool_calls entry for call_7f2, followed by a tool message with tool_call_id call_7f2 and the invoice JSON. The server answers in text, the adapter returns a text part, and the turn ends. If the id had been dropped or regenerated between the two calls, strict OpenAI-compatible servers reject the second request because the tool result answers no known call. That is the most common bug in hand-written adapters.

For how Gemini itself shapes function calls, compare Gemini function calling in ADK Java.

Streaming

When stream is true, the built-in Gemini adapter emits responses marked partial(true) as chunks arrive, then one aggregated, non-partial response with the complete text and any function calls. Partial events drive incremental output to the user; the final one is what the flow acts on. To support streaming, parse the provider's server-sent events, emit a partial response per text delta, accumulate tool-call argument fragments by index until the stream ends, and emit the aggregated final response last.

Ignoring the flag and always returning one complete response is legal; the built-in Claude adapter does exactly that. The agent still works, it just shows nothing until the whole answer exists. Start there, and add streaming when users need it. See streaming in ADK Java for the event side.

Registering the model

// At startup, before any agent resolves its model.
HttpClient http = HttpClient.newBuilder().connectTimeout(Duration.ofSeconds(5)).build();
URI base = URI.create(System.getenv("LOCAL_LLM_URL"));
String key = System.getenv("LOCAL_LLM_KEY");
LlmRegistry.registerLlm("local/.*", name -> new OpenAiCompatibleLlm(name, base, key, http));

LlmAgent agent = LlmAgent.builder()
    .name("support")
    .model("local/qwen2.5-32b-instruct")        // resolved through the registry
    .instruction("Answer billing questions. Use tools for account data.")
    .tools(FunctionTool.create(BillingTools.class, "lookupInvoice"))
    .build();

// Or skip the registry and hand the agent an instance directly:
// LlmAgent.builder().model(new OpenAiCompatibleLlm("local/qwen2.5-32b-instruct", base, key, http))

LlmRegistry.registerLlm takes a regular expression and a factory. The defaults cover gemini-.*, gemma-.* and apigee/.*. Use a prefix that cannot overlap them, such as local/: the registry picks the first matching pattern from a map, so overlapping patterns make the choice unpredictable. Register during startup, before any agent resolves its model, because the agent resolves lazily and the registry caches the instance it creates. Resolving every configured model name at boot makes a missing registration fail at deploy time rather than on the first user request; the runtime boot sequence shows where.

Testing with a scripted model

final class ScriptedLlm extends BaseLlm {
  private final Queue<LlmResponse> script;
  final List<LlmRequest> seen = new CopyOnWriteArrayList<>();

  ScriptedLlm(LlmResponse... responses) {
    super("scripted");
    this.script = new ConcurrentLinkedQueue<>(List.of(responses));
  }

  @Override
  public Flowable<LlmResponse> generateContent(LlmRequest req, boolean stream) {
    return Flowable.defer(() -> {
      seen.add(req);
      LlmResponse next = script.poll();
      return next == null
          ? Flowable.error(new IllegalStateException("script exhausted"))
          : Flowable.just(next);
    });
  }

  @Override
  public BaseLlmConnection connect(LlmRequest req) {
    throw new UnsupportedOperationException();
  }
}

A scripted model makes agent tests deterministic: queue a function-call response followed by a text response, run the agent, and assert both on the final output and on the recorded requests, for example that the second request contains a function response with the right id. Test the real adapter separately with contract tests: golden JSON for request mapping, recorded provider responses for response mapping, and one live smoke test per deploy that performs a real tool-calling turn.

Failure modes

SymptomCauseFix
Second model call rejected after a tool runsCall id lost, regenerated or not echoed as tool_call_idPreserve ids exactly; generate once only if absent
Model ignores tool results or repeats callsEach call sent as its own assistant message, or roles unmappedOne assistant message per model turn with all tool_calls
Answers bleed between usersPer-request state stored in the shared adapterKeep the adapter stateless; state lives in the request
Event loop stalls under loadBlocking HTTP on the subscribing threadsubscribeOn(Schedulers.io()) or an async client
Cost dashboards read zerousageMetadata not mappedMap prompt, completion and total tokens
Wrong adapter picked for a modelOverlapping registry patternsUse a distinct prefix
Truncated answers look completeFinish reason droppedMap it and alert on token-limit stops

What to do next

  1. Check whether an existing adapter or contrib module already reaches your provider.
  2. Implement generateContent without streaming, and throw from connect.
  3. Write golden-file tests for request mapping, including a turn with two parallel tool calls and their responses.
  4. Map usage metadata and finish reasons before the first production deploy.
  5. Register under a unique prefix at startup and resolve every configured model at boot.
  6. Add a scripted fake for agent tests, then add streaming only when users need incremental output.
Key takeaway: A custom model in ADK Java is a subclass of BaseLlm that turns an LlmRequest into your provider's request and the reply back into an LlmResponse, returned as a cold Flowable. Most of the work is faithful translation: user and model roles, one assistant message per turn with every tool call, tool results paired by exact id, and usage and finish reasons carried through. Because the registry shares one instance per model name, keep the adapter stateless and run blocking I/O on the IO scheduler. Start without streaming, register under a unique prefix, and test agents with a scripted fake.