An agent test that calls a real model is slow, costs money, needs credentials in CI and fails at random, because the same prompt does not always produce the same reply. The fix is the same as for any external dependency: put a fake behind the interface. In ADK Java that interface is BaseLlm, and a scripted fake of it is usually called a mock LLM.
One fact first, because the title invites a wrong assumption: ADK Java does not ship a public MockLlm class. Its own test suite uses a class called TestLlm, which lives in the repository's test sources rather than the published library. This article shows how to write your own MockLlm modelled on it, how to test agents and tools with it, and how to use the same scripts to contract-test a custom BaseLlm adapter you wrote for another provider. It ends with a checklist.
Why fake the model
Agent behaviour has two layers. The model decides what to say and which tool to call. The code around it builds the request (instructions, history, tool declarations), runs tools, feeds results back, stores session state, and turns the reply into events. Bugs in the second layer are deterministic and are exactly what unit tests catch, but only if the first layer is held still.
A mock model holds it still. It returns replies you script, in order, and records every request it receives. That lets a test ask precise questions: did the agent call the right tool with the right arguments, did the tool's result reach the second model call, did the final text reach the user, did an error from the model surface as an error. Quality questions, such as whether the real model picks the right tool, belong to evaluation, covered in the evaluation framework article.
The contract you are faking
Every model in ADK Java extends com.google.adk.models.BaseLlm, which has a model name and two abstract methods: generateContent(LlmRequest, boolean stream), returning a Flowable<LlmResponse>, and connect(LlmRequest) for live bidirectional sessions. The BaseLlm interface overview covers the contract in detail.
An LlmRequest carries contents() (the conversation so far), tools() (a map from tool name to tool), config() and getSystemInstructions(). An LlmResponse is built with LlmResponse.builder() and carries a content, optional partial and turnComplete flags, and optional errorCode (a FinishReason) and errorMessage. A non-streaming call emits one response. ADK's Gemini adapter, in streaming mode, emits chunks marked partial(true) followed by a final aggregated response with partial left unset; a mock that supports streaming should follow the same shape.
Writing MockLlm
The mock keeps a list of scripted responses and a thread-safe list of received requests. Flowable.defer makes each subscription consume exactly one script entry when it is subscribed, not when the Flowable is built. Running past the end of the script is an error with a clear message rather than a silent empty reply, because an unexpected extra model call is itself a bug worth seeing.
public final class MockLlm extends BaseLlm {
private final List<LlmResponse> script;
private final AtomicInteger next = new AtomicInteger();
private final List<LlmRequest> requests = Collections.synchronizedList(new ArrayList<>());
private MockLlm(List<LlmResponse> script) {
super("mock-llm");
this.script = List.copyOf(script);
}
public static MockLlm of(LlmResponse... replies) { return new MockLlm(List.of(replies)); }
public static LlmResponse text(String s) {
return LlmResponse.builder()
.content(Content.builder().role("model").parts(Part.fromText(s)).build())
.build();
}
public static LlmResponse call(String id, String name, Map<String, Object> args) {
Part part = Part.builder()
.functionCall(FunctionCall.builder().id(id).name(name).args(args).build())
.build();
return LlmResponse.builder()
.content(Content.builder().role("model").parts(part).build())
.build();
}
public static LlmResponse error(FinishReason.Known code, String message) {
return LlmResponse.builder().errorCode(new FinishReason(code)).errorMessage(message).build();
}
@Override
public Flowable<LlmResponse> generateContent(LlmRequest request, boolean stream) {
return Flowable.defer(() -> {
requests.add(request);
int i = next.getAndIncrement();
if (i >= script.size()) {
return Flowable.error(new AssertionError(
"MockLlm script exhausted: model call " + (i + 1) + ", " + script.size() + " scripted"));
}
return stream ? Flowable.fromIterable(chunks(script.get(i))) : Flowable.just(script.get(i));
});
}
@Override
public BaseLlmConnection connect(LlmRequest request) {
throw new UnsupportedOperationException("MockLlm does not support live sessions");
}
public List<LlmRequest> requests() {
synchronized (requests) { return List.copyOf(requests); }
}
}The chunks helper splits a text reply into word-sized partial(true) responses and appends the original response, unchanged, as the final one. Replies without text, such as function calls, pass through as a single response.
private static List<LlmResponse> chunks(LlmResponse full) {
String text = full.content().flatMap(Content::parts).orElse(List.of()).stream()
.map(p -> p.text().orElse("")).collect(Collectors.joining());
if (text.isEmpty()) return List.of(full);
List<LlmResponse> out = new ArrayList<>();
for (String piece : text.split("(?<= )")) { // keep the trailing spaces
out.add(LlmResponse.builder().partial(true)
.content(Content.builder().role("model").parts(Part.fromText(piece)).build())
.build());
}
out.add(full); // final aggregated response
return out;
}A test for a streaming UI can then assert that the concatenated partial text equals the final text, and that the UI layer ignores partial events when it persists the transcript. A scripted transport failure, as opposed to a model-level error, is a response list entry you replace with Flowable.error(new IOException(...)); a variant constructor taking a supplier, as TestLlm has, makes that easy.
Testing an agent and its tools
Wire the mock into a real LlmAgent with a real tool and run it through InMemoryRunner. Only the model is fake; the flow that turns a function call into a tool invocation and back is ADK's own code, which is what makes the test meaningful.
public final class OrderTools {
public static Map<String, Object> getOrder(
@Annotations.Schema(name = "orderId", description = "Order id") String orderId) {
return Map.of("orderId", orderId, "status", "SHIPPED");
}
}
@Test
void toolResultReachesSecondModelCall() {
MockLlm llm = MockLlm.of(
MockLlm.call("c1", "getOrder", Map.of("orderId", "A-17")),
MockLlm.text("Order A-17 has shipped."));
LlmAgent agent = LlmAgent.builder()
.name("support").model(llm)
.instruction("Answer order questions using the tools.")
.tools(FunctionTool.create(OrderTools.class, "getOrder"))
.build();
InMemoryRunner runner = new InMemoryRunner(agent);
Session s = runner.sessionService().createSession(runner.appName(), "u1").blockingGet();
List<Event> events = runner
.runAsync("u1", s.id(), Content.fromParts(Part.fromText("Where is A-17?")))
.toList().blockingGet();
String finalText = events.stream().filter(Event::finalResponse)
.flatMap(e -> e.content().flatMap(Content::parts).orElse(List.of()).stream())
.map(p -> p.text().orElse("")).collect(Collectors.joining());
assertEquals("Order A-17 has shipped.", finalText);
assertEquals(2, llm.requests().size());
assertTrue(llm.requests().get(0).tools().containsKey("getOrder"));
boolean sawResult = llm.requests().get(1).contents().stream()
.flatMap(c -> c.parts().orElse(List.of()).stream())
.anyMatch(p -> p.functionResponse().flatMap(FunctionResponse::name)
.map("getOrder"::equals).orElse(false));
assertTrue(sawResult, "second model call must carry the tool result");
}Notice the script has two entries. After the tool runs, ADK calls the model again with the function response appended, so a script containing only the function call ends in the exhaustion error on the second call. Forgetting the follow-up reply is the most common first mistake with mock models.
Contract-testing your own adapter
The second use of the same idea is testing a custom BaseLlm you wrote, such as the OpenAI-compatible adapter in the custom LLM article. Here the adapter is the code under test, so the fake moves one layer down: a local HTTP server that returns canned provider JSON. The JDK's com.sun.net.httpserver.HttpServer is enough and needs no extra dependency.
Write the checks as an abstract contract test with one subclass per implementation. Run it against MockLlm to prove the tests themselves are right, then against your adapter backed by the fake server. Useful contract cases:
- A plain text reply maps to one response with role model and the expected text.
- A provider tool call maps to a function call part with the right name, parsed arguments and a stable id.
- A conversation containing a function response serialises to the provider's tool-result message, which you check by inspecting the body the fake server received.
- System instructions from
getSystemInstructions()appear in the outgoing request exactly once. - Streaming emits
partial(true)chunks whose concatenated text equals the final response. - HTTP 429 and 5xx surface as errors the runner can see, and a malformed body does not hang the subscription.
- Nothing blocks the subscribing thread; test with a short timeout on the blocking call.
The sync versus async contracts article explains why the last case matters: an adapter that does its HTTP call during generateContent instead of on subscription blocks whichever thread built the Flowable.
Errors and a worked example
Script a model-level error with MockLlm.error(FinishReason.Known.SAFETY, "blocked") and assert what your agent does with it. In the code checked for this article, the response's error code and message are carried onto the emitted event; rather than relying on that, write the assertion against the events your test receives so an upgrade that changes the behaviour fails a test instead of production. For transport failures, assert that iterating the runner's events throws, and that any retry or fallback logic you added, such as a callback that substitutes a canned apology, produces the expected final text.
Worked example: a support agent must stop when the model refuses. The script is a single SAFETY error. The test asserts that exactly one model request was made, that no event carries a function call (so no tool ran), and that the refusal is visible on an event where your UI or callback can act on it. Three assertions, no network, and it runs in milliseconds on every commit in the CI pipeline. If you add an apology callback, add a fourth assertion on its text.
@Test
void refusalIsSurfacedAndNoToolRuns() {
MockLlm llm = MockLlm.of(MockLlm.error(FinishReason.Known.SAFETY, "blocked"));
// supportAgent(llm): the LlmAgent builder from the previous test, extracted into a helper.
// runOnce(agent, text): new InMemoryRunner, createSession, runAsync(...).toList().blockingGet().
LlmAgent agent = supportAgent(llm);
List<Event> events = runOnce(agent, "Cancel every order in the system");
assertEquals(1, llm.requests().size());
assertTrue(events.stream().allMatch(e -> e.functionCalls().isEmpty()));
assertTrue(events.stream().anyMatch(e -> e.errorCode().isPresent()),
"pin this: the refusal must be visible on an event");
}If the last assertion ever fails after an upgrade, that is the test doing its job: the error now travels a different way, and your apology callback probably needs to change with it.
Failure modes
- Asserting on prompt text verbatim. Requests include ADK-generated instructions that change between releases. Assert on what you own: your instruction is present, your tools are declared, your tool result is in the history.
- Shared mocks across tests. The script index and request log are state; build a new mock per test.
- Mock drift. A mock that returns shapes the real adapter never produces, such as a function call without an id, makes tests pass for code that fails in production. Keep the mock's builders aligned with recorded real responses.
- Over-mocking. Replacing tools as well as the model leaves nothing real to test. Fake the model and external services, but run your own tools.
- Eager subscriptions. Consuming the script when the Flowable is created rather than subscribed misattributes replies when ADK retries or resubscribes.
Trade-offs
Scripted mocks give speed and determinism at the cost of realism: they prove the plumbing, not the judgement. Recorded-replay fakes, which return real responses captured once from the provider, are more realistic but go stale when prompts change. Evaluation against the live model measures quality but is slow, costly and noisy. A healthy suite has many mock-based unit tests, a handful of replay tests for each adapter, and a scheduled evaluation run, with an LLM judge where exact matching does not work.
What to do next
- Add a MockLlm class to your test sources, modelled on ADK's TestLlm, with text, call and error builders.
- Make generateContent lazy with Flowable.defer and fail loudly when the script runs out.
- Write one runner-level test per tool that scripts the call and the follow-up reply and asserts the tool result reached the second request.
- Add tests for model errors and transport failures that assert on the events your users see.
- Turn your custom adapter's checks into an abstract contract test and run it against MockLlm and a local fake HTTP server.
- Run these tests on every commit, and keep live-model evaluation as a separate scheduled job.
- Re-run the suite on each ADK upgrade before changing production code.