ADK for Java talks to models through one abstraction, BaseLlm. Gemini is built in, and the project's contrib directory ships a Spring AI integration whose SpringAI class implements BaseLlm by wrapping any Spring AI ChatModel. Spring AI in turn has a first-class Ollama client. Put the two together and an ADK agent, with its instructions, tools, sessions and callbacks unchanged, runs against a model on your own machine or inside your own network.
That combination is useful for three jobs: offline development without API keys or spend, air-gapped and data-residency deployments where prompts may not leave the building, and cheap deterministic-ish integration tests. It is also easy to get subtly wrong, because there are four layers between your agent and the weights and each has its own defaults. This article wires the stack explicitly, traces one tool-calling turn through every layer, and lists the failure modes that show up in practice. How Ollama itself schedules and loads models is covered in Ollama local serving; writing your own adapter instead is covered in implementing a custom LLM.
Four layers between the agent and the weights
Read the diagram from left to right. The LlmAgent holds the instruction and the tool list. The Runner's flow turns the session history into an LlmRequest: a list of genai Content objects, the system instructions, tool declarations and a GenerateContentConfig with sampling settings. The SpringAI adapter converts that request into a Spring AI Prompt (system, user and assistant messages plus tool definitions) and calls the ChatModel. OllamaChatModel turns the prompt into JSON for Ollama's /api/chat endpoint, using its own OllamaChatOptions for the model name, context size and keep-alive. Ollama loads the weights if they are not resident and generates.
The return path mirrors it: an Ollama message with text or tool_calls becomes a Spring AI ChatResponse, which the adapter maps to an LlmResponse whose content parts are text or FunctionCall objects. The important consequence is ownership. Every setting exists in two vocabularies, ADK's and Spring AI's, and the adapter maps only some of them. The contrib README lists temperature, max output tokens, top-p and stop sequences as mapped, and says plainly that top-k, presence and frequency penalties are not. Anything Ollama-specific, such as num_ctx or keep_alive, has no ADK equivalent at all and must be set on the Spring AI side.
Dependencies and version pinning
You need three things on the classpath: ADK core, the ADK Spring AI module and Spring AI's Ollama model module, with the Spring AI version managed by its BOM. The ADK module has been published to Maven Central as com.google.adk:google-adk-spring-ai and tracks ADK core releases. Each release is built against a specific Spring AI line, and Spring AI renamed starters and options classes between 1.x and 2.x, so pin one ADK release, open its POM, and import the Spring AI BOM version it declares rather than the newest one.
<dependencyManagement>
<dependencies>
<dependency>
<groupId>org.springframework.ai</groupId>
<artifactId>spring-ai-bom</artifactId>
<version>${spring-ai.version}</version> <!-- the line your ADK release was built with -->
<type>pom</type>
<scope>import</scope>
</dependency>
</dependencies>
</dependencyManagement>
<dependencies>
<dependency>
<groupId>com.google.adk</groupId>
<artifactId>google-adk</artifactId>
<version>${adk.version}</version>
</dependency>
<dependency>
<groupId>com.google.adk</groupId>
<artifactId>google-adk-spring-ai</artifactId>
<version>${adk.version}</version>
</dependency>
<dependency>
<groupId>org.springframework.ai</groupId>
<artifactId>spring-ai-starter-model-ollama</artifactId>
</dependency>
</dependencies>If the build resolves two Spring AI versions, the symptom is a NoSuchMethodError at the first call, not a compile error. Run mvn dependency:tree -Dincludes=org.springframework.ai once after every upgrade and make sure a single version appears.
Wiring the model explicitly
The starter can auto-configure an OllamaChatModel from properties, and the ADK module's auto-configuration can then build a SpringAI bean from it. For agent work, prefer an explicit bean. The reference documentation for property names has shifted between Spring AI versions (the 2.0 reference lists spring.ai.ollama.chat.model while its own YAML sample still nests the model under chat.options), and an explicit builder makes the settings that matter for agents visible in code review instead of resting on defaults.
@Configuration
class LocalModelConfig {
@Bean
OllamaChatModel ollamaChatModel(@Value("${ollama.base-url:http://localhost:11434}") String baseUrl) {
OllamaApi api = OllamaApi.builder().baseUrl(baseUrl).build();
return OllamaChatModel.builder()
.ollamaApi(api)
.options(OllamaChatOptions.builder() // 1.x: defaultOptions(OllamaOptions...)
.model("qwen2.5:7b") // a model tagged for tool use
.temperature(0.2)
.numCtx(16384) // never rely on the 2048 default for agents
.keepAlive("30m") // avoid cold reloads between turns
.build())
.build();
}
@Bean
BaseLlm localLlm(OllamaChatModel ollama) {
// ChatModel extends StreamingChatModel, so one bean serves both paths.
return new SpringAI(ollama, ollama, "qwen2.5:7b");
}
@Bean
LlmAgent supportAgent(BaseLlm localLlm) {
return LlmAgent.builder()
.name("support")
.model(localLlm) // LlmAgent.Builder.model(BaseLlm)
.instruction("You answer order questions. Use lookupOrder for any order id.")
.tools(FunctionTool.create(OrderTools.class, "lookupOrder"))
.build();
}
}
public final class OrderTools {
public static Map<String, Object> lookupOrder(
@Schema(name = "orderId", description = "Numeric order id") String orderId) {
return Map.of("orderId", orderId, "status", "SHIPPED", "carrier", "DHL");
}
}Two details deserve attention. The model name passed to SpringAI is what ADK reports in events and traces; the model Ollama actually runs is the one in OllamaChatOptions. Keep them identical, or your dashboards will attribute latency to a model that never served the call. And numCtx is the setting most teams forget: Spring AI documents a default of 2,048 tokens, which an agent with a long instruction, several tool schemas and a few turns of history exceeds quickly. Ollama truncates an over-long prompt rather than rejecting it, so the failure is a model that silently forgets its instruction or its tools.
Properties that still matter
Properties still matter for the parts you do not build by hand. Two blocks are relevant: Spring AI's spring.ai.ollama block for the connection and model pulling, and the ADK module's adk.spring-ai block for its defaults, validation and observability.
spring:
ai:
ollama:
base-url: http://ollama.internal:11434 # default http://localhost:11434
init:
pull-model-strategy: when_missing # always | when_missing | never (default never)
timeout: 10m # default 5m; large models pull slowly
max-retries: 2 # default 0
adk:
spring-ai:
temperature: 0.2
max-tokens: 1024
top-p: 0.9
# top-k: 40 <- accepted here, but the README says top-k is NOT mapped to Spring AI
validation:
enabled: true
fail-fast: true
observability:
enabled: true
metrics-enabled: true
include-content: false # keep prompts out of logs by defaultUse when_missing in development so a fresh laptop works on first run, and never in production: a pod that starts pulling several gigabytes on boot will fail its readiness probe, and you want images and model files to be provisioned deliberately, not by whichever replica starts first. Set include-content to false unless you have a retention policy for prompt text in logs.
One tool-calling turn, end to end
Here is one turn, traced. The user asks: Where is order 1042?
- The Runner appends the user message to the session and the flow builds an
LlmRequestwith the instruction, one userContentand a declaration forlookupOrderwhose single string parameter isorderId. - The adapter's message converter produces a system message and a user message, and its tool converter turns the declaration into a Spring AI tool definition carrying the JSON schema.
OllamaChatModelposts to/api/chatwith"tools","options": {"num_ctx": 16384, "temperature": 0.2}and"keep_alive": "30m". Ollama returns an assistant message with atool_callsentry: namelookupOrder, arguments{"orderId": "1042"}.- The adapter maps that to an
LlmResponsecontaining aFunctionCallpart. The Runner emits it as an event, executes the Java method in your JVM, and emits a second event carrying theFunctionResponse. - The flow builds a new request that now contains the call and its result, and the second round trip returns plain text: Order 1042 has shipped with DHL.
Step 4 is the one to protect. ADK must own the tool loop, because that is where its callbacks, plugins, tool confirmation and event log live. Spring AI also knows how to run tools itself (its ToolCallingChatOptions carry an internalToolExecutionEnabled flag), and if a tool were ever executed inside the Spring AI layer, ADK would see only the final text: no FunctionCall event, no beforeToolCallback, no audit record. Pin that with a test that asserts the function-call event exists, shown below, instead of trusting a default.
Choosing a local model for agent work
Agent workloads stress a local model differently from chat. The model must emit well-formed tool calls, respect a long system instruction and stay coherent across a growing history. Choose from models that Ollama's library tags for tool use, then measure on your own tool schemas; published benchmark numbers rarely transfer to a specific set of tools.
| Concern | What to check | Why it matters for agents |
|---|---|---|
| Tool-call support | Model tagged for tools; Ollama 0.2.8+ (0.4.6+ for streaming tools) | Without it the request fails or the model writes JSON into prose |
| Context window | num_ctx set explicitly; prompt tokens per turn measured | Silent truncation drops the instruction or tool schemas first |
| Memory | Weights plus KV cache for the context you set fit in VRAM | A larger num_ctx costs memory; spill to CPU multiplies latency |
| Concurrency | How many requests Ollama runs in parallel per model | Parallel ADK agents queue behind each other |
| Determinism | Temperature near 0 and a fixed seed for tests | Flaky tool selection makes evaluations meaningless |
Testing against a real Ollama
A local model makes a realistic integration test affordable. Run Ollama in a container, pull a small tool-capable model once per CI cache, and assert on ADK's event stream rather than on exact wording.
@Test
void lookupOrderGoesThroughAdkToolLoop() {
InMemoryRunner runner = new InMemoryRunner(supportAgent);
Session session = runner.sessionService()
.createSession(runner.appName(), "u1").blockingGet();
List<Event> events = runner.runAsync("u1", session.id(),
Content.fromParts(Part.fromText("Where is order 1042?")))
.toList().blockingGet();
// The model proposed the call and ADK, not Spring AI, executed it.
assertTrue(events.stream().anyMatch(e -> !e.functionCalls().isEmpty()));
assertTrue(events.stream().anyMatch(e -> !e.functionResponses().isEmpty()));
String answer = events.get(events.size() - 1).stringifyContent();
assertTrue(answer.contains("DHL") || answer.contains("shipped"));
}Keep a second test layer that needs no model at all: a scripted fake BaseLlm that returns a fixed FunctionCall, described in testing custom LLMs. The fake proves your agent logic; the Ollama test proves the wiring and the model's tool-calling behaviour.
Failure modes
- Silent context truncation. Long instructions plus tool schemas exceed
num_ctxand the model starts ignoring rules. Log prompt token counts from the usage metadata and alert when they approach the window. - Model not present. With
pull-model-strategy: nevera missing model fails at the first request, not at startup. Add a readiness check that lists local models and fails if yours is absent. - Cold loads. After
keep_aliveexpires, the next turn pays the full load time, often seconds. Raise keep-alive for interactive agents and size timeouts for the cold path. - Unsupported tool calling. A model without tool support returns an error or prose that looks like JSON. Treat a missing
FunctionCallon a turn that clearly needed one as a model-selection problem, not a prompt problem. - localhost in containers.
http://localhost:11434inside a container points at the container. Configure the base URL per environment. - Unmapped settings. Top-k set in ADK config has no effect. Set sampling that matters on
OllamaChatOptionsand verify it in Ollama's request log. - Queueing under parallel agents. A fan-out of five sub-agents against one model instance serialises behind Ollama's parallelism limit; latency, not errors, is the symptom.
Trade-offs
| Route to a local model | Strength | Cost |
|---|---|---|
ADK SpringAI + Spring AI Ollama | Spring Boot wiring, one adapter for many providers, observability hooks | Two config vocabularies; partial option mapping |
| LangChain4j integration | Rich agentic toolkit if you already use it | Another abstraction with its own tool loop to keep out of ADK's way |
Custom BaseLlm over Ollama's API | Exact control of wire format and options | You maintain the adapter against two moving APIs |
| Hosted Gemini | Strongest tool use, no hardware | Spend, network egress, data leaves your boundary |
A common production split is hosted models for user-facing agents and local models for development, evaluation sweeps and classification sub-agents whose prompts must stay inside the network. Because the agent only sees BaseLlm, that split is a bean swap, not a rewrite. See LangChain4j agents for the other integration route.
What to do next
- Pin one ADK release, import the Spring AI BOM version its POM declares and confirm a single Spring AI version in the dependency tree.
- Build the
OllamaChatModelbean explicitly with model,numCtx, temperature and keep-alive, and use the same model name inSpringAI. - Measure prompt tokens for your longest realistic turn and set
num_ctxwith headroom. - Write the integration test that asserts a
FunctionCalland aFunctionResponseevent, and run it in CI against a containerised Ollama. - Set
pull-model-strategy: neverin production and add a readiness check for the model. - Read Ollama local serving before sizing hardware or concurrency.