A unit test proves that one tool parses a response. An evaluation proves that the model makes good decisions. Neither proves that the pieces of an ADK Java agent work together: that the Runner persists what your tool wrote to state, that a second replica can resume the conversation, that a plugin sees every tool call, that the router really hands off to the right sub-agent, and that a slow downstream service produces an error the model can read instead of a hung request. Those are integration bugs, and they are the ones that reach production because each component passes its own tests.
This article builds an integration harness for ADK Java agents. The model is scripted so tests are deterministic, using the MockLlm built in the custom LLM testing article; it is not reprinted here. Everything else is real, except the far side of the network, which is a fake HTTP server that can be told to fail. You will see the harness, the fake, state and reload tests, plugins, transfer, a worked example, and where these tests sit next to unit tests and evaluations.
What an integration test is for an agent
An integration test earns its cost by crossing a seam that a unit test stubs out. For an agent the seams are concrete:
| Seam | Typical bug | What the test asserts |
|---|---|---|
| Tool to downstream HTTP | No timeout, wrong path, error body thrown as an exception | Tool returns an error map within a deadline; request log shows the right call |
| Tool to session state | State written under the wrong key or scope, or not at all | Reloaded session holds the expected keys and values |
| Runner to session service | Events not persisted, history lost on another replica | A second runner over the same store continues the conversation |
| Plugins and callbacks | Audit or policy plugin skipped for some tools | Plugin log has one entry per tool call, in order |
| Agent to sub-agent | Transfer target misnamed, wrong agent answers | Event authors follow the expected path |
| Model request assembly | Tool missing from the request, result not fed back | Recorded model requests contain the tool and its result |
Notice what is not on the list: whether the answer is good. That belongs to evaluation. An integration test with a scripted model checks plumbing, so its assertions are exact and it never flakes.
The harness
Build the runner the way production does, through Runner.builder(), rather than InMemoryRunner. The builder lets you pass the session service and plugins you want to test, and it requires an artifact service, so the harness supplies an in-memory one.
final class AgentHarness implements AutoCloseable {
final FakeDownstream downstream; // JDK HttpServer, below
final BaseSessionService sessions;
final MockLlm llm;
final Runner runner;
final List<String> pluginLog = new CopyOnWriteArrayList<>();
AgentHarness(BaseSessionService sessions, MockLlm llm) throws IOException {
this.downstream = FakeDownstream.start();
this.sessions = sessions;
this.llm = llm;
BaseAgent root = Agents.build(llm, downstream.baseUrl()); // the SAME factory production uses
this.runner = Runner.builder()
.agent(root).appName("travel")
.artifactService(new InMemoryArtifactService())
.sessionService(sessions)
.plugins(new AuditPlugin(pluginLog))
.build();
}
List<Event> turn(String userId, String sessionId, String text) {
return runner.runAsync(userId, sessionId, Content.fromParts(Part.fromText(text)))
.toList().timeout(10, TimeUnit.SECONDS).blockingGet();
}
@Override public void close() { downstream.stop(); }
}Two choices matter. First, the agent graph comes from the same factory method production calls, with the model and base URL as parameters. If the test assembles its own graph, it tests a copy. Second, every blocking call has a timeout. A test that hangs on a stuck subscription tells you nothing and blocks the build.
Fake the downstream, not the tool
Fake the downstream service, not the tool. A stubbed tool skips the HTTP client, the serialisation and the error mapping, which is where the bugs live. The JDK's com.sun.net.httpserver.HttpServer is enough and adds no dependency. Give it a queue of canned behaviours so each test can script success, a 500, a slow reply or a malformed body, and record every request it receives.
final class FakeDownstream {
record Reply(int status, String body, long delayMs) {}
private final HttpServer server;
final Deque<Reply> script = new ConcurrentLinkedDeque<>();
final List<String> requests = new CopyOnWriteArrayList<>();
static FakeDownstream start() throws IOException { return new FakeDownstream(); }
private FakeDownstream() throws IOException {
server = HttpServer.create(new InetSocketAddress("127.0.0.1", 0), 0); // port 0: no clashes
server.createContext("/", ex -> {
requests.add(ex.getRequestMethod() + " " + ex.getRequestURI());
Reply r = Optional.ofNullable(script.poll()).orElse(new Reply(500, "{\"error\":\"unscripted\"}", 0));
try { Thread.sleep(r.delayMs()); } catch (InterruptedException ignored) {}
byte[] b = r.body().getBytes(StandardCharsets.UTF_8);
ex.sendResponseHeaders(r.status(), b.length);
try (OutputStream os = ex.getResponseBody()) { os.write(b); }
});
server.start();
}
String baseUrl() { return "http://127.0.0.1:" + server.getAddress().getPort(); }
void stop() { server.stop(0); }
}The tool is an instance with its client and base URL injected, registered with FunctionTool.create(instance, "searchFlights"). Static tool methods, the shortest path in examples, make this injection awkward and tempt people to read a global URL, which then leaks between tests.
State and the toolContext name trap
Tools write session state through ToolContext, and that is a common place for silent failure. FunctionTool injects the context into the parameter whose @Schema name, or else reflected parameter name, is toolContext. Reflected names are only real if the code was compiled with javac -parameters; without it the parameter is called arg0, the context is not injected, and the tool may fail or treat the context as an ordinary argument. Annotate the parameter explicitly so the build flag does not matter, and let an integration test prove the write lands.
public Map<String, Object> holdSeat(
@Schema(name = "flightId", description = "Flight to hold") String flightId,
@Schema(name = "toolContext") ToolContext toolContext) {
Map<String, Object> hold = client.hold(flightId); // POST /holds via the fake
toolContext.state().put("held_flight", flightId); // session scope: no prefix
toolContext.state().put("hold_id", hold.get("holdId")); // confirmHold reads this
toolContext.state().put("temp:last_hold_raw", hold); // temp: is not persisted
return hold;
}Assert against a session read back from the service, not the object you passed in, because the point is persistence. Read it with sessions.getSession(app, user, id, Optional.empty()).blockingGet() and check that held_flight is present and that the temp: key is not. Keys prefixed user: and app: are shared across a user's sessions or the whole app, so tests that use them need a fresh user id or app name per test, or they pass alone and fail in a suite.
Persist, reload, and a second replica
The most valuable test in the harness simulates a second replica. Run turn one through one runner, then build a new runner over the same session service and run turn two. In production the second turn often lands on a different instance, and anything held in memory outside the session, such as a field on a tool object, a cache in a callback or a static map, is gone.
@Test
void secondReplicaContinuesTheBooking() throws IOException {
BaseSessionService store = sessionService(); // in-memory here, container below
MockLlm llm = MockLlm.of(
MockLlm.call("c1", "holdSeat", Map.of("flightId", "LH-401")),
MockLlm.text("Seat held on LH-401."),
MockLlm.call("c2", "confirmHold", Map.of()), // reads held_flight from state
MockLlm.text("Booked."));
try (AgentHarness a = new AgentHarness(store, llm);
AgentHarness b = new AgentHarness(store, llm)) { // shared script, shared store
a.downstream.script.add(new FakeDownstream.Reply(200, "{\"holdId\":\"H1\"}", 0));
b.downstream.script.add(new FakeDownstream.Reply(200, "{\"status\":\"BOOKED\"}", 0));
Session s = store.createSession("travel", "u-" + UUID.randomUUID()).blockingGet();
a.turn(s.userId(), s.id(), "Hold LH-401");
b.turn(s.userId(), s.id(), "Book it");
assertEquals(List.of("POST /holds/LH-401"), a.downstream.requests);
assertEquals(List.of("POST /bookings/H1"), b.downstream.requests);
}
}Run the same abstract test class against every BaseSessionService you deploy: one subclass returns new InMemorySessionService(), another returns your database-backed service pointed at a throwaway container. The in-memory run is fast and runs on every commit; the container run proves serialisation of real event and state types, which is where JSON round-trip bugs hide. The runner streams partial events but does not persist them, so assert on persisted history only for final events.
Plugins and agent transfer in the loop
Plugins are attached to the runner, so they apply to every agent and tool in the graph, and they are easy to lose when someone builds a runner by hand. An audit plugin in the harness is both a test subject and a probe. Its beforeToolCallback(BaseTool, Map, ToolContext) returns an empty Maybe to let the tool run; a non-empty map would replace the tool's result. Assert one log entry per tool call, in order, and assert that a policy plugin's override actually prevents the downstream request, which the fake server's log shows.
Agent transfer is scripted the same way as a tool call. ADK adds a transfer_to_agent tool with an agent_name argument to agents that have transfer targets, so the scripted router reply is MockLlm.call("t1", "transfer_to_agent", Map.of("agent_name", "bookings")), followed by whatever the bookings agent should say. Assert on event.author(): the final response must come from bookings. Because a single scripted model serves every agent in the graph, the script reads as one conversation across agents; give each sub-agent its own mock if that becomes hard to follow. The callbacks article covers callback ordering, and the tool dispatch article explains how function calls become tool invocations.
Worked example: a hold that hung the turn
Worked example: a travel agent with a router and a bookings sub-agent. The test scripts five model replies: transfer to bookings, call searchFlights, answer, call holdSeat, answer. The fake server is scripted with a search result and then a hold that takes eight seconds.
The first run hung for the full ten-second test timeout. The HTTP client had been built with no request timeout, so holdSeat blocked the turn. The fix was a three-second timeout and a catch that returns {"status":"error","reason":"downstream_timeout"}. The test now asserts that the turn completes in under four seconds, that the model request after the hold call contains the error map as the function response, and that held_flight was not written. A second test scripts a 500 with an HTML body and asserts the same shape, which caught a JSON parse exception escaping the tool. Neither bug was visible to unit tests, which mocked the client, or to evaluations, which ran against a healthy staging service.
Failure modes
- Testing a hand-built graph. The test passes and production differs. Build agents from the production factory.
- Shared scoped state.
user:andapp:keys leak between tests that reuse ids. Use random ids. - Fixed ports. Parallel test forks collide. Bind port 0 and read the port back.
- Unscripted calls returning success. A fake that defaults to 200 hides unexpected requests. Default to an error and assert the script is drained.
- Asserting on ADK's prompt text. Framework instructions change between releases. Assert on your tools, your state and your request log.
- No timeouts. A stuck
Flowableblocks the build. Bound every blocking call.
Trade-offs
Integration tests cost more than unit tests: a fake server, a runner per test and, for the container variant, seconds of start-up. Keep them to the seams in the table, perhaps twenty to forty tests for a mid-sized agent, and keep logic tests at unit level. They are deterministic because the model is scripted, which means they cannot tell you the model will choose the right tool; the eval-in-CI article covers that layer with recorded responses. Run the in-memory suite on every commit and the container suite on merge, both inside the CI pipeline.
What to do next
- Move agent construction into one factory that takes the model and downstream base URLs as parameters, and use it in production and tests.
- Write
FakeDownstreamwith a scripted reply queue, a request log and an error default. - Annotate every
ToolContextparameter with@Schema(name = "toolContext")and add a test that a state write survives a reload. - Add the two-runner test so state held outside the session is caught.
- Make the session tests abstract and add a subclass for your production session store in a container.
- Add fault tests for timeout, 500 with a non-JSON body and 429, each asserting an error map reaches the next model request.
- Script one transfer and assert event authors.