An agent looks hard to unit test because its central component, the model, is non-deterministic. Most of an ADK Java agent, however, is ordinary deterministic code: tool methods, argument validation, callbacks that allow or block, state updates, the declarations ADK sends to the model, and the shape of the agent tree. Those parts can and should be tested the same way as any Java code, quickly, on every build and without a network.

This page is about that deterministic layer. Faking the model itself is covered in Testing Custom LLMs with MockLLM, and measuring answer quality against a real model is the job of the evaluation framework. Here the goal is a suite of a few hundred tests that runs in seconds and catches the bugs that actually break agents in production: a tool that accepts a bad argument, a guard callback with an off-by-one limit, a renamed parameter that silently changes the model's schema, a sub-agent that was dropped from the tree.

Four layers, sorted by what a test needs

Unit tests for an ADK agent: four layers, one fakeEvals (not unit tests)real model, scored, separate suite4. Runner tests with a scripted modelwiring: call, callback, result, state3. Golden structure testsdeclarations, agent tree, instructions2. Callback decisions as pure functionsallow, block, rewrite1. Tool logic as plain Javavalidation, business rules, error envelopesFewer, slower tests toward the top; most tests at the bottom run in milliseconds with no ADK typesOnly layer 4 constructs an agent; only evals call a real model
Most tests sit at the bottom and touch no ADK types; only runner tests build an agent, and only evals call a real model.

Sort every check you want into one of four layers by asking what it needs in order to run. If it needs only inputs and outputs, it is layer 1 or 2. If it needs ADK to describe your code, but not to execute it, it is layer 3. If it needs ADK's flow, where a function call becomes a tool invocation and the result goes back to the model, it is layer 4, and it needs a scripted model. If it needs the real model's judgement, it is not a unit test, and it belongs in an eval suite with tolerances.

The ratio matters more than the tooling. A healthy agent project has many layer 1 and 2 tests, a handful of layer 3 golden files, and perhaps one layer 4 test per tool and per callback. If you find yourself writing a runner test to check a validation rule, the rule is in the wrong place.

Layer 1: tool logic as plain Java

The design move that makes tools testable is to split each one into a pure core and a thin adapter. The core takes plain values and returns a plain result. The adapter is the method ADK registers; it reads anything it needs from ToolContext, calls the core, and writes state. Take a refund tool with a per-order limit:

public record RefundDecision(String status, String code, long approvedCents) {
  Map<String, Object> toMap() {
    return Map.of("status", status, "code", code, "approved_cents", approvedCents);
  }
}

public final class RefundPolicy {
  static final long AUTO_LIMIT_CENTS = 50_00;

  public static RefundDecision decide(String orderId, long amountCents, long alreadyRefunded,
                                      long orderTotalCents) {
    if (orderId == null || !orderId.matches("ORD-\\d{5}"))
      return new RefundDecision("error", "INVALID_ORDER_ID", 0);
    if (amountCents <= 0)
      return new RefundDecision("error", "NON_POSITIVE_AMOUNT", 0);
    if (alreadyRefunded + amountCents > orderTotalCents)
      return new RefundDecision("error", "EXCEEDS_ORDER_TOTAL", 0);
    if (amountCents > AUTO_LIMIT_CENTS)
      return new RefundDecision("needs_approval", "OVER_AUTO_LIMIT", 0);
    return new RefundDecision("ok", "APPROVED", amountCents);
  }
}

Every rule is now one assertion away. Parameterised tests are the natural fit, because the interesting cases are boundaries:

@ParameterizedTest
@CsvSource({
  "ORD-00001, 5000, 0, 9000, ok,             APPROVED",
  "ORD-00001, 5001, 0, 9000, needs_approval, OVER_AUTO_LIMIT",
  "ORD-00001, 1000, 8500, 9000, error,       EXCEEDS_ORDER_TOTAL",
  "ORD-1,     1000, 0, 9000, error,          INVALID_ORDER_ID",
  "ORD-00001, 0,    0, 9000, error,          NON_POSITIVE_AMOUNT",
})
void refundRules(String id, long amt, long prior, long total, String status, String code) {
  RefundDecision d = RefundPolicy.decide(id, amt, prior, total);
  assertEquals(status, d.status());
  assertEquals(code, d.code());
}

Two habits pay off here. Assert on machine-readable codes, not on message text, because the codes are the contract the model and your error handling act on (see wrapping tool errors for the model). And test that the core never throws for bad model input: a model can and will send a null, an empty string or a number in the wrong unit, and every one of those should come back as an error result rather than an exception.

Layer 2: callbacks as pure decisions

Callbacks follow the same pattern. A before-tool callback in ADK Java has this synchronous form:

Optional<Map<String, Object>> call(InvocationContext invocationContext, BaseTool baseTool,
                                   Map<String, Object> input, ToolContext toolContext);

Returning a map short-circuits the tool and uses the map as its result; returning empty lets the tool run. You do not want to construct those four arguments in a unit test, and you do not need to. Put the decision in a static method over plain values, and make the registered lambda a one-line adapter:

public final class ToolGuard {
  static final Set<String> WRITE_TOOLS = Set.of("issueRefund", "cancelOrder");

  public static Optional<Map<String, Object>> decide(String tool, Map<String, Object> args,
                                                     Map<String, Object> state) {
    if (WRITE_TOOLS.contains(tool) && !Boolean.TRUE.equals(state.get("user:verified")))
      return Optional.of(Map.of("status", "error", "code", "IDENTITY_NOT_VERIFIED"));
    Object n = state.getOrDefault("temp:write_calls", 0);
    if (WRITE_TOOLS.contains(tool) && n instanceof Integer i && i >= 3)
      return Optional.of(Map.of("status", "error", "code", "WRITE_BUDGET_EXHAUSTED"));
    return Optional.empty();
  }
}

// registration: the only line that touches ADK types
agentBuilder.beforeToolCallbackSync((inv, tool, args, ctx) ->
    ToolGuard.decide(tool.name(), args, ctx.state()));

Now the guard has ordinary tests: a write tool with no verification is blocked, a read tool is never blocked, the third write passes and the fourth does not, and a state value of the wrong type does not crash. The single wiring line is covered once, in layer 4. The same split works for before-model callbacks that rewrite requests and for after-tool callbacks that redact results: test the rewrite as a function from one value to another. Callback ordering and composition are described in the callback architecture.

Layer 3: golden declarations and agent trees

Some bugs are not in your logic but in what ADK derives from it. Rename a Java parameter without the schema annotation, compile without -parameters, or edit a description, and the function declaration the model sees changes, and tool selection quietly shifts. A golden test makes that change visible in review. BaseTool.declaration() returns an Optional<FunctionDeclaration>, and the genai types serialise with toJson():

@Test
void toolDeclarationsMatchGolden() throws IOException {
  List<BaseTool> tools = List.of(
      FunctionTool.create(OrderTools.class, "issueRefund"),
      FunctionTool.create(OrderTools.class, "getOrder"));
  String actual = tools.stream()
      .map(t -> t.declaration().orElseThrow().toJson())
      .map(GoldenJson::canonical)                 // sort keys, fixed indentation
      .collect(Collectors.joining("\n"));
  Path golden = Path.of("src/test/resources/golden/order_tools.json");
  if (Boolean.getBoolean("updateGolden")) Files.writeString(golden, actual);
  assertEquals(Files.readString(golden), actual,
      "tool schema changed; rerun with -DupdateGolden=true if intended");
}

Canonicalise before comparing (sorted keys, stable whitespace), or the test fails on serialiser noise. Apply the same idea to the agent tree: walk it from the root, record each agent's name, type, model name and tool names, and compare with a golden file. Add two invariants to that walk: names are unique, because transfer resolves agents by name, and no description is empty, because a parent model routes on descriptions. Instructions that are templates with state placeholders deserve a test that renders them with representative state and asserts that no placeholder survives unresolved.

Layer 4: runner tests that prove wiring

Layer 4 proves the wiring: the guard is registered, ADK calls it, and its result reaches the model. Use a scripted model such as the MockLlm from the MockLLM article, an InMemoryRunner, and a state seeded at session creation. The model asks for a refund, the guard blocks it, and the test checks what ADK sent back on the second model call:

@Test
void unverifiedRefundIsBlockedBeforeTheToolRuns() {
  AtomicInteger toolRuns = new AtomicInteger();
  OrderTools.onIssueRefund = toolRuns::incrementAndGet;     // test-only counter; prefer an injected fake client
  MockLlm llm = MockLlm.of(
      MockLlm.call("c1", "issueRefund", Map.of("orderId", "ORD-00042", "amountCents", 1200)),
      MockLlm.text("I need to verify your identity first."));
  LlmAgent agent = LlmAgent.builder()
      .name("support").model(llm).instruction("Help with orders.")
      .tools(FunctionTool.create(OrderTools.class, "issueRefund"))
      .beforeToolCallbackSync((inv, tool, args, ctx) ->
          ToolGuard.decide(tool.name(), args, ctx.state()))
      .build();
  InMemoryRunner runner = new InMemoryRunner(agent);
  Session s = runner.sessionService()
      .createSession(runner.appName(), "u1", Map.of("user:verified", false), null)
      .blockingGet();
  runner.runAsync("u1", s.id(), Content.fromParts(Part.fromText("Refund ORD-00042")))
      .toList().blockingGet();

  assertEquals(0, toolRuns.get());
  String sentBack = llm.requests().get(1).contents().toString();
  assertTrue(sentBack.contains("IDENTITY_NOT_VERIFIED"));
}

Note what this test does not assert: the final wording. The scripted reply is fixed, so asserting on it only proves that the mock works. Assert on effects: how many times the tool ran, what the model was sent, and what ended up in session state. Keep the script one entry per model call, including the reply after the tool result; a short script fails with an exhaustion error, which is the most common first mistake.

Worked example: changing a refund limit

Here is how the layers divide a real change. A product manager asks for refunds of up to 75.00 to be approved automatically, instead of 50.00. The change is one constant. The tests that fail, and what each one tells you, are:

  • Layer 1: the 5001-cent row in refundRules now returns ok. You update the boundary rows to 7500 and 7501, and the test documents the new rule.
  • Layer 3: nothing fails, which is correct, because the schema did not change. If the PR had also renamed amountCents to amount, the golden diff would show the model-facing change, and a reviewer would ask whether the unit changed too.
  • Layer 4: nothing fails, because the wiring is unchanged.
  • Evals: the nightly suite may show the agent offering refunds it previously escalated. That is a product question, not a unit test failure.

The whole unit suite still runs in a few seconds, so it can gate every commit. The eval suite, with its variance and cost, gates releases instead, as described in integrating evals into CI.

Failure modes

FailureCauseRemedy
Tests pass, agent breaks in stagingLogic lives in the tool adapter or a lambda, untestedMove it into a pure core; keep adapters one call long
Flaky unit suiteA test calls a real model or the networkScripted model only; real models belong in evals
Golden test fails on every runNon-canonical JSON or unordered mapsCanonicalise keys and whitespace before comparing
Mock script exhaustedForgot the model reply after a tool resultOne script entry per model call
State-dependent bug slips throughTests always seed the happy-path stateParameterise over missing, wrong-type and boundary state
Tests assert on final textScripted text proves nothingAssert on tool runs, model requests and state

Trade-offs

Splitting every tool into a core and an adapter adds a type or two per tool, and some teams find the indirection fussy for trivial tools. A reasonable rule: if a tool has a branch, it gets a core. Golden tests catch real schema drift but create review churn; keep them to declarations and the tree, not to whole prompts, which change too often to be useful as goldens. Runner tests with scripted models are realistic about ADK's flow but encode an assumed conversation. If the real model would never make that call sequence, the test is proving a path nobody takes. Use production traces, as in writing eval test cases, to choose which sequences to script.

What to do next

  1. List every tool and callback in your agent, and mark each one that contains a branch.
  2. For each marked tool, extract a pure core that returns a record with a status and a code, and write parameterised boundary tests.
  3. Rewrite each callback as a static decision over name, arguments and state, and register it with a one-line lambda.
  4. Add a golden test for tool declarations and one for the agent tree, with unique-name and non-empty-description invariants.
  5. Write one runner test per callback and per write tool, using a scripted model, and assert on effects rather than text.
  6. Make the unit suite the commit gate. Move anything that touches a real model into the eval suite.
  7. When a production incident traces to deterministic code, add the failing case at the lowest layer that can express it.
Key takeaway: Most of an agent is deterministic, so test it like ordinary Java. Put tool rules in pure cores with boundary tests, make callbacks static decisions over plain values, golden-test the declarations and agent tree that ADK derives, and keep a small set of scripted-model runner tests to prove the wiring. Assert on effects, never on scripted text, and leave model judgement to evals.