ADK for Java talks to models through one abstraction, BaseLlm. Gemini is built in, and the project's contrib directory ships a Spring AI integration whose SpringAI class implements BaseLlm by wrapping any Spring AI ChatModel. Spring AI in turn has a first-class Ollama client. Put the two together and an ADK agent, with its instructions, tools, sessions and callbacks unchanged, runs against a model on your own machine or inside your own network.

That combination is useful for three jobs: offline development without API keys or spend, air-gapped and data-residency deployments where prompts may not leave the building, and cheap deterministic-ish integration tests. It is also easy to get subtly wrong, because there are four layers between your agent and the weights and each has its own defaults. This article wires the stack explicitly, traces one tool-calling turn through every layer, and lists the failure modes that show up in practice. How Ollama itself schedules and loads models is covered in Ollama local serving; writing your own adapter instead is covered in implementing a custom LLM.

Four layers between the agent and the weights

An ADK agent on a local model: four layers, two protocolsLlmAgentinstruction, toolsRunner + flowtool loop, eventsSpringAI adapterextends BaseLlmOllamaChatModelSpring AI ChatModelOllamaApiREST clientOllama serverport 11434Model weightsloaded, kept aliveFunctionToolruns in your JVMLlmRequestPromptChatOptionsPOST /api/chatFunctionCallADK owns the tool loop: the model onlyproposes calls, the Runner executes them
The agent and Runner speak ADK types; the SpringAI adapter translates to Spring AI; OllamaChatModel speaks Ollama's REST API. Tools always execute in your JVM under ADK's control.

Read the diagram from left to right. The LlmAgent holds the instruction and the tool list. The Runner's flow turns the session history into an LlmRequest: a list of genai Content objects, the system instructions, tool declarations and a GenerateContentConfig with sampling settings. The SpringAI adapter converts that request into a Spring AI Prompt (system, user and assistant messages plus tool definitions) and calls the ChatModel. OllamaChatModel turns the prompt into JSON for Ollama's /api/chat endpoint, using its own OllamaChatOptions for the model name, context size and keep-alive. Ollama loads the weights if they are not resident and generates.

The return path mirrors it: an Ollama message with text or tool_calls becomes a Spring AI ChatResponse, which the adapter maps to an LlmResponse whose content parts are text or FunctionCall objects. The important consequence is ownership. Every setting exists in two vocabularies, ADK's and Spring AI's, and the adapter maps only some of them. The contrib README lists temperature, max output tokens, top-p and stop sequences as mapped, and says plainly that top-k, presence and frequency penalties are not. Anything Ollama-specific, such as num_ctx or keep_alive, has no ADK equivalent at all and must be set on the Spring AI side.

Dependencies and version pinning

You need three things on the classpath: ADK core, the ADK Spring AI module and Spring AI's Ollama model module, with the Spring AI version managed by its BOM. The ADK module has been published to Maven Central as com.google.adk:google-adk-spring-ai and tracks ADK core releases. Each release is built against a specific Spring AI line, and Spring AI renamed starters and options classes between 1.x and 2.x, so pin one ADK release, open its POM, and import the Spring AI BOM version it declares rather than the newest one.

<dependencyManagement>
  <dependencies>
    <dependency>
      <groupId>org.springframework.ai</groupId>
      <artifactId>spring-ai-bom</artifactId>
      <version>${spring-ai.version}</version>  <!-- the line your ADK release was built with -->
      <type>pom</type>
      <scope>import</scope>
    </dependency>
  </dependencies>
</dependencyManagement>
<dependencies>
  <dependency>
    <groupId>com.google.adk</groupId>
    <artifactId>google-adk</artifactId>
    <version>${adk.version}</version>
  </dependency>
  <dependency>
    <groupId>com.google.adk</groupId>
    <artifactId>google-adk-spring-ai</artifactId>
    <version>${adk.version}</version>
  </dependency>
  <dependency>
    <groupId>org.springframework.ai</groupId>
    <artifactId>spring-ai-starter-model-ollama</artifactId>
  </dependency>
</dependencies>

If the build resolves two Spring AI versions, the symptom is a NoSuchMethodError at the first call, not a compile error. Run mvn dependency:tree -Dincludes=org.springframework.ai once after every upgrade and make sure a single version appears.

Wiring the model explicitly

The starter can auto-configure an OllamaChatModel from properties, and the ADK module's auto-configuration can then build a SpringAI bean from it. For agent work, prefer an explicit bean. The reference documentation for property names has shifted between Spring AI versions (the 2.0 reference lists spring.ai.ollama.chat.model while its own YAML sample still nests the model under chat.options), and an explicit builder makes the settings that matter for agents visible in code review instead of resting on defaults.

@Configuration
class LocalModelConfig {

  @Bean
  OllamaChatModel ollamaChatModel(@Value("${ollama.base-url:http://localhost:11434}") String baseUrl) {
    OllamaApi api = OllamaApi.builder().baseUrl(baseUrl).build();
    return OllamaChatModel.builder()
        .ollamaApi(api)
        .options(OllamaChatOptions.builder()   // 1.x: defaultOptions(OllamaOptions...)
            .model("qwen2.5:7b")               // a model tagged for tool use
            .temperature(0.2)
            .numCtx(16384)                     // never rely on the 2048 default for agents
            .keepAlive("30m")                  // avoid cold reloads between turns
            .build())
        .build();
  }

  @Bean
  BaseLlm localLlm(OllamaChatModel ollama) {
    // ChatModel extends StreamingChatModel, so one bean serves both paths.
    return new SpringAI(ollama, ollama, "qwen2.5:7b");
  }

  @Bean
  LlmAgent supportAgent(BaseLlm localLlm) {
    return LlmAgent.builder()
        .name("support")
        .model(localLlm)                        // LlmAgent.Builder.model(BaseLlm)
        .instruction("You answer order questions. Use lookupOrder for any order id.")
        .tools(FunctionTool.create(OrderTools.class, "lookupOrder"))
        .build();
  }
}

public final class OrderTools {
  public static Map<String, Object> lookupOrder(
      @Schema(name = "orderId", description = "Numeric order id") String orderId) {
    return Map.of("orderId", orderId, "status", "SHIPPED", "carrier", "DHL");
  }
}

Two details deserve attention. The model name passed to SpringAI is what ADK reports in events and traces; the model Ollama actually runs is the one in OllamaChatOptions. Keep them identical, or your dashboards will attribute latency to a model that never served the call. And numCtx is the setting most teams forget: Spring AI documents a default of 2,048 tokens, which an agent with a long instruction, several tool schemas and a few turns of history exceeds quickly. Ollama truncates an over-long prompt rather than rejecting it, so the failure is a model that silently forgets its instruction or its tools.

Properties that still matter

Properties still matter for the parts you do not build by hand. Two blocks are relevant: Spring AI's spring.ai.ollama block for the connection and model pulling, and the ADK module's adk.spring-ai block for its defaults, validation and observability.

spring:
  ai:
    ollama:
      base-url: http://ollama.internal:11434   # default http://localhost:11434
      init:
        pull-model-strategy: when_missing      # always | when_missing | never (default never)
        timeout: 10m                            # default 5m; large models pull slowly
        max-retries: 2                          # default 0
adk:
  spring-ai:
    temperature: 0.2
    max-tokens: 1024
    top-p: 0.9
    # top-k: 40   <- accepted here, but the README says top-k is NOT mapped to Spring AI
    validation:
      enabled: true
      fail-fast: true
    observability:
      enabled: true
      metrics-enabled: true
      include-content: false                    # keep prompts out of logs by default

Use when_missing in development so a fresh laptop works on first run, and never in production: a pod that starts pulling several gigabytes on boot will fail its readiness probe, and you want images and model files to be provisioned deliberately, not by whichever replica starts first. Set include-content to false unless you have a retention policy for prompt text in logs.

One tool-calling turn, end to end

Here is one turn, traced. The user asks: Where is order 1042?

  1. The Runner appends the user message to the session and the flow builds an LlmRequest with the instruction, one user Content and a declaration for lookupOrder whose single string parameter is orderId.
  2. The adapter's message converter produces a system message and a user message, and its tool converter turns the declaration into a Spring AI tool definition carrying the JSON schema.
  3. OllamaChatModel posts to /api/chat with "tools", "options": {"num_ctx": 16384, "temperature": 0.2} and "keep_alive": "30m". Ollama returns an assistant message with a tool_calls entry: name lookupOrder, arguments {"orderId": "1042"}.
  4. The adapter maps that to an LlmResponse containing a FunctionCall part. The Runner emits it as an event, executes the Java method in your JVM, and emits a second event carrying the FunctionResponse.
  5. The flow builds a new request that now contains the call and its result, and the second round trip returns plain text: Order 1042 has shipped with DHL.

Step 4 is the one to protect. ADK must own the tool loop, because that is where its callbacks, plugins, tool confirmation and event log live. Spring AI also knows how to run tools itself (its ToolCallingChatOptions carry an internalToolExecutionEnabled flag), and if a tool were ever executed inside the Spring AI layer, ADK would see only the final text: no FunctionCall event, no beforeToolCallback, no audit record. Pin that with a test that asserts the function-call event exists, shown below, instead of trusting a default.

Choosing a local model for agent work

Agent workloads stress a local model differently from chat. The model must emit well-formed tool calls, respect a long system instruction and stay coherent across a growing history. Choose from models that Ollama's library tags for tool use, then measure on your own tool schemas; published benchmark numbers rarely transfer to a specific set of tools.

ConcernWhat to checkWhy it matters for agents
Tool-call supportModel tagged for tools; Ollama 0.2.8+ (0.4.6+ for streaming tools)Without it the request fails or the model writes JSON into prose
Context windownum_ctx set explicitly; prompt tokens per turn measuredSilent truncation drops the instruction or tool schemas first
MemoryWeights plus KV cache for the context you set fit in VRAMA larger num_ctx costs memory; spill to CPU multiplies latency
ConcurrencyHow many requests Ollama runs in parallel per modelParallel ADK agents queue behind each other
DeterminismTemperature near 0 and a fixed seed for testsFlaky tool selection makes evaluations meaningless

Testing against a real Ollama

A local model makes a realistic integration test affordable. Run Ollama in a container, pull a small tool-capable model once per CI cache, and assert on ADK's event stream rather than on exact wording.

@Test
void lookupOrderGoesThroughAdkToolLoop() {
  InMemoryRunner runner = new InMemoryRunner(supportAgent);
  Session session = runner.sessionService()
      .createSession(runner.appName(), "u1").blockingGet();

  List<Event> events = runner.runAsync("u1", session.id(),
          Content.fromParts(Part.fromText("Where is order 1042?")))
      .toList().blockingGet();

  // The model proposed the call and ADK, not Spring AI, executed it.
  assertTrue(events.stream().anyMatch(e -> !e.functionCalls().isEmpty()));
  assertTrue(events.stream().anyMatch(e -> !e.functionResponses().isEmpty()));
  String answer = events.get(events.size() - 1).stringifyContent();
  assertTrue(answer.contains("DHL") || answer.contains("shipped"));
}

Keep a second test layer that needs no model at all: a scripted fake BaseLlm that returns a fixed FunctionCall, described in testing custom LLMs. The fake proves your agent logic; the Ollama test proves the wiring and the model's tool-calling behaviour.

Failure modes

  • Silent context truncation. Long instructions plus tool schemas exceed num_ctx and the model starts ignoring rules. Log prompt token counts from the usage metadata and alert when they approach the window.
  • Model not present. With pull-model-strategy: never a missing model fails at the first request, not at startup. Add a readiness check that lists local models and fails if yours is absent.
  • Cold loads. After keep_alive expires, the next turn pays the full load time, often seconds. Raise keep-alive for interactive agents and size timeouts for the cold path.
  • Unsupported tool calling. A model without tool support returns an error or prose that looks like JSON. Treat a missing FunctionCall on a turn that clearly needed one as a model-selection problem, not a prompt problem.
  • localhost in containers. http://localhost:11434 inside a container points at the container. Configure the base URL per environment.
  • Unmapped settings. Top-k set in ADK config has no effect. Set sampling that matters on OllamaChatOptions and verify it in Ollama's request log.
  • Queueing under parallel agents. A fan-out of five sub-agents against one model instance serialises behind Ollama's parallelism limit; latency, not errors, is the symptom.

Trade-offs

Route to a local modelStrengthCost
ADK SpringAI + Spring AI OllamaSpring Boot wiring, one adapter for many providers, observability hooksTwo config vocabularies; partial option mapping
LangChain4j integrationRich agentic toolkit if you already use itAnother abstraction with its own tool loop to keep out of ADK's way
Custom BaseLlm over Ollama's APIExact control of wire format and optionsYou maintain the adapter against two moving APIs
Hosted GeminiStrongest tool use, no hardwareSpend, network egress, data leaves your boundary

A common production split is hosted models for user-facing agents and local models for development, evaluation sweeps and classification sub-agents whose prompts must stay inside the network. Because the agent only sees BaseLlm, that split is a bean swap, not a rewrite. See LangChain4j agents for the other integration route.

What to do next

  1. Pin one ADK release, import the Spring AI BOM version its POM declares and confirm a single Spring AI version in the dependency tree.
  2. Build the OllamaChatModel bean explicitly with model, numCtx, temperature and keep-alive, and use the same model name in SpringAI.
  3. Measure prompt tokens for your longest realistic turn and set num_ctx with headroom.
  4. Write the integration test that asserts a FunctionCall and a FunctionResponse event, and run it in CI against a containerised Ollama.
  5. Set pull-model-strategy: never in production and add a readiness check for the model.
  6. Read Ollama local serving before sizing hardware or concurrency.
Key takeaway: ADK's contrib SpringAI adapter lets any Spring AI ChatModel stand in for Gemini, and Spring AI's OllamaChatModel makes that a local model. Wire the Ollama bean explicitly, set the context window and keep-alive yourself, keep the model name identical on both sides, and prove with a test that ADK, not Spring AI, executes your tools. Then a local model becomes a configuration choice instead of a fork of your agent code.