Continuous integration for an ordinary Java service answers one question: did this change break anything? For an agent built with the Agent Development Kit for Java the question splits in two. Some of the system is ordinary code, tools, configuration, wiring, and can be tested deterministically on every commit. The rest is behaviour that emerges from a model reading your instructions, which is nondeterministic, costs money per run, needs credentials, and can change without any commit at all when a provider updates a model.
A pipeline that treats both halves the same fails in one of two ways. Either it calls the real model on every pull request and becomes slow, flaky and expensive, or it never calls the model and ships prompt regressions. This article builds a pipeline that separates them: a fast deterministic lane driven by a scripted model, a merge lane that builds and smoke-tests the real artifact, and a scheduled live-evaluation lane whose gate is statistical rather than a single pass or fail. Code uses ADK's documented runner and model APIs and plain JUnit 5; nothing here depends on an ADK evaluation tool for Java.
What can break, and where each failure should be caught
List the ways an ADK Java agent breaks before designing the pipeline, because each belongs to a different lane.
| Failure | Example | Cheapest place to catch it |
|---|---|---|
| Tool code bug | Order lookup throws on a lowercase id | Tool unit test, every PR |
| Tool declaration drift | Parameter names arrive as arg0 because a compiler flag was dropped | Agent test with a scripted model, every PR |
| Wiring error | Agent references a model name no registry pattern matches | Boot test, every PR |
| Adapter mapping bug | Tool-call ids lost when translating a provider response | Contract test with recorded JSON, every PR |
| Credential or network problem | Expired key, blocked egress from the cluster | Live smoke test, merge and deploy |
| Behaviour regression | A prompt edit makes the agent skip a required tool | Live eval suite, nightly and pre-release |
| Upstream model change | Provider updates the model behind an alias | Scheduled live eval with no code change |
Build configuration that tests depend on
One compiler setting is load-bearing. ADK builds a tool's declaration by reflecting over the Java method, and without the -parameters flag, reflection reports parameter names as arg0, arg1 and so on. The model then sees meaningless argument names, unless each has @Schema(name = ...), and a ToolContext parameter, which ADK recognises by the name toolContext, is no longer injected. The function tool guide explains the mechanism. Put the flag in the build, and add a test that fails if it disappears.
Split tests into tagged groups so each lane runs only what it should. Surefire's excludedGroups and groups settings work with JUnit 5 tags.
<properties>
<adk.version><!-- pin an exact released version --></adk.version>
</properties>
<dependency>
<groupId>com.google.adk</groupId>
<artifactId>google-adk</artifactId>
<version>${adk.version}</version>
</dependency>
<plugin>
<artifactId>maven-compiler-plugin</artifactId>
<configuration>
<release>21</release>
<parameters>true</parameters> <!-- tool parameter names survive reflection -->
</configuration>
</plugin>
<plugin>
<artifactId>maven-surefire-plugin</artifactId>
<configuration>
<excludedGroups>live</excludedGroups> <!-- default build never calls a real model -->
</configuration>
</plugin>
<profiles>
<profile>
<id>live</id>
<build><plugins><plugin>
<artifactId>maven-surefire-plugin</artifactId>
<configuration>
<groups>live</groups>
<excludedGroups combine.self="override"/>
</configuration>
</plugin></plugins></build>
</profile>
</profiles>Pin the ADK version exactly and update it through a pull request, never through a version range. ADK Java moves quickly, and a minor version can change event shapes or callback order; you want that change to arrive with a green or red build attached to it.
Lane one: tools are just methods
Function tools in ADK Java are static methods wrapped with FunctionTool.create(Class, String). Test the method directly with ordinary JUnit: valid input, malformed input, a downstream timeout, and the shape of the returned map. Because the model reads the result map, its keys are part of the contract; assert them, not only the happy-path values. Tests at this level run in milliseconds and catch most real bugs.
Add one reflection test per tool that asserts the parameter names, so a missing compiler flag fails loudly instead of degrading the agent.
@Test
void toolParameterNamesSurviveCompilation() throws Exception {
Method m = OrderTools.class.getMethod("getOrderStatus", String.class);
assertEquals("orderId", m.getParameters()[0].getName(),
"compile with -parameters, or tool arguments arrive as arg0");
}
Lane one: agent tests with a scripted model
The key to a deterministic agent test is replacing the model, not the agent. ADK models are subclasses of BaseLlm whose generateContent returns an RxJava Flowable<LlmResponse>. A scripted subclass returns queued responses in order and records every request it receives, as described in implementing a custom LLM. The agent, its instruction, its tools and the runner are all real.
final class ScriptedLlm extends BaseLlm {
private final Queue<LlmResponse> script;
final List<LlmRequest> seen = new CopyOnWriteArrayList<>();
ScriptedLlm(LlmResponse... responses) {
super("scripted");
this.script = new ConcurrentLinkedQueue<>(List.of(responses));
}
@Override
public Flowable<LlmResponse> generateContent(LlmRequest req, boolean stream) {
return Flowable.defer(() -> {
seen.add(req);
LlmResponse next = script.poll();
return next == null
? Flowable.error(new IllegalStateException("script exhausted"))
: Flowable.just(next);
});
}
@Override
public BaseLlmConnection connect(LlmRequest req) {
throw new UnsupportedOperationException();
}
}
static LlmResponse modelSays(Part part) {
return LlmResponse.builder()
.content(Content.builder().role("model").parts(List.of(part)).build())
.build();
}
@Test
void orderQuestionCallsToolThenAnswers() {
ScriptedLlm model = new ScriptedLlm(
modelSays(Part.builder().functionCall(FunctionCall.builder()
.id("call-1").name("getOrderStatus").args(Map.of("orderId", "A-1001")).build()).build()),
modelSays(Part.fromText("Order A-1001 shipped yesterday.")));
LlmAgent agent = LlmAgent.builder()
.name("orders")
.model(model)
.instruction("Use getOrderStatus for any question about an order.")
.tools(FunctionTool.create(OrderTools.class, "getOrderStatus"))
.build();
InMemoryRunner runner = new InMemoryRunner(agent);
Session session = runner.sessionService()
.createSession(runner.appName(), "test-user").blockingGet();
List<Event> events = runner.runAsync(session.userId(), session.id(),
Content.fromParts(Part.fromText("Where is A-1001?")), RunConfig.builder().build())
.toList().blockingGet();
String answer = events.stream().filter(Event::finalResponse)
.map(Event::stringifyContent).reduce("", String::concat);
assertTrue(answer.contains("shipped"));
// The second model call must carry the tool result back, paired to the call.
boolean resultSent = model.seen.get(1).contents().stream()
.flatMap(c -> c.parts().orElse(List.of()).stream())
.anyMatch(p -> p.functionResponse().isPresent());
assertTrue(resultSent, "tool result was not returned to the model");
}What this test proves is narrow and valuable: that the declared tool can be called by name with these arguments, that the runner executes it, that its result flows back into the next model request, and that the final event carries the answer. It says nothing about whether a real model would choose the tool; that belongs to the live lane. Write scripted tests for every tool, for sub-agent transfers, for the error path where a tool returns an error status, and for callbacks that block or rewrite requests.
Assert on recorded requests as well as outputs. The system instruction, the tool declarations and the number of model calls are all visible in seen, and a change that silently drops a tool from the declaration list will show up there before any user notices.
Lane one: boot and contract tests
A boot test starts the application's real wiring with test configuration and resolves every model name the agents reference through the registry, failing the build if any name does not match a registered pattern. Without it, a misspelled model name or a missing custom-adapter registration fails on the first user request in production. The runtime boot sequence shows where resolution happens and why it is lazy by default.
Contract tests cover any custom BaseLlm adapter or external tool client. Record real provider responses once, including a turn with two parallel tool calls, a refusal, a rate-limit error and a truncated response, and replay them through the adapter's response mapper with an HTTP stub such as WireMock. Assert that ids, roles, finish reasons and usage metadata survive translation. Re-record deliberately when the provider changes its format, and review the diff like code.
Lane two: build the artifact once and smoke-test it
On merge, re-run lane one on the merged commit, build the container image, record its digest, and run one live smoke test against that exact image: a single real model call that performs one tool-calling turn and checks the answer contains a known fact from the tool. The smoke test is about plumbing, credentials, egress, region and model name, not quality. Run it again after each deployment, against the deployed environment, as described in Kubernetes deployment.
Model credentials never belong in the pull-request lane. On GitHub, workflows triggered by pull requests from forks do not receive repository secrets, which is the behaviour you want: lane one must pass with no key at all. Keep keys in a protected environment that only runs on main, scheduled or release triggers, and read them from the environment as configuration management recommends. The Gemini integration reads its key from the GOOGLE_API_KEY environment variable.
name: agent-ci
on:
pull_request:
push:
branches: [main]
workflow_dispatch: # manual pre-release run
schedule:
- cron: "17 3 * * *" # nightly live eval, even with no commits
jobs:
verify:
runs-on: ubuntu-latest
timeout-minutes: 20
steps:
- uses: actions/checkout@v4
- uses: actions/setup-java@v4
with: { distribution: temurin, java-version: "21", cache: maven }
- run: mvn -B -ntp verify # lane one: no secrets, no live model
live-eval:
if: github.event_name == 'schedule' || github.event_name == 'workflow_dispatch'
needs: verify
runs-on: ubuntu-latest
environment: live-model # protected environment holds the key
concurrency: live-eval # never two paid runs at once
timeout-minutes: 40
steps:
- uses: actions/checkout@v4
- uses: actions/setup-java@v4
with: { distribution: temurin, java-version: "21", cache: maven }
- run: mvn -B -ntp -Plive verify
env:
GOOGLE_API_KEY: ${{ secrets.GOOGLE_API_KEY }}
- uses: actions/upload-artifact@v4
if: always()
with: { name: eval-report, path: target/eval-report.json }
Lane three: a live eval with a statistical gate
The live suite is a set of cases, each a user message plus checks on the outcome: which tools were called with which arguments, whether the final answer contains or omits specific facts, whether a refusal happened when it should. Tool-call checks are the most reliable, because they are structured; text checks should be narrow. Each case runs several times, because one sample of a nondeterministic system tells you little.
Do not gate on the raw pass rate. A suite of 40 cases run 3 times each gives 120 trials, and a pass rate of 92.5 percent on one night and 89 percent the next may be noise. Gate on a lower confidence bound instead, which accounts for sample size. The Wilson score interval is a good default for proportions.
static double wilsonLower(int passes, int n, double z) {
double p = (double) passes / n;
double denom = 1 + z * z / n;
double centre = p + z * z / (2.0 * n);
double margin = z * Math.sqrt(p * (1 - p) / n + z * z / (4.0 * n * n));
return (centre - margin) / denom;
}
// 111 of 120 trials passed: raw 0.925, 95% lower bound about 0.864
boolean release = wilsonLower(111, 120, 1.96) >= 0.85;Worked through: with 111 passes out of 120, the raw rate is 0.925 and the 95 percent Wilson lower bound is about 0.864, which clears a floor of 0.85. If the same rate came from only 40 trials, 37 passes, the bound would be about 0.80 and the gate would fail, correctly, because the evidence is thinner. Keep a few must-pass cases outside the statistical gate, such as never calling a refund tool without confirmation, and fail the run if any trial of those fails.
Pin everything the result depends on and write it into the report: model name and version, ADK version, instruction text hash, tool declaration hash, temperature and the case-file commit. A nightly run whose score drops with no change in any of those is evidence of an upstream model change, which is exactly what the scheduled trigger exists to detect.
Failure modes
- Live calls in the PR lane. Builds become flaky, slow and expensive, and developers learn to re-run until green. Keep lane one hermetic.
- Over-scripted tests. Asserting on exact prompt text makes every instruction tweak break dozens of tests. Assert on structure: tools declared, results returned, calls counted.
- Gate on raw pass rate. Small suites flip between pass and fail on noise; teams then lower the threshold until it means nothing.
- Unpinned model aliases. An alias that tracks the latest model makes yesterday's eval irrelevant to today's production.
- Eval cases leaking into prompts. When failures are fixed by pasting the case into the instruction, the suite measures memorisation. Keep a held-out set that prompt authors do not see.
- Budget runaway. A loop in an agent under test can make hundreds of model calls. Set a per-run call budget and a job timeout.
Trade-offs
| Decision | Option A | Option B |
|---|---|---|
| Model in PR tests | Scripted fake: fast, free, deterministic, says nothing about quality | Real model: catches behaviour changes, slow, costly, flaky |
| Eval frequency | Nightly: catches upstream drift, costs every day | Per release: cheaper, drift found late |
| Trials per case | One: cheap, gate dominated by noise | Three to five: tighter bounds, proportionally more cost |
| Gate | Hard block on release: safe, can stall shipping on noise | Alert and human review: flexible, depends on discipline |
What to do next
- Turn on the compiler's parameters flag and add the reflection test that fails if it is ever dropped.
- Write plain JUnit tests for every tool, including the error map the model sees.
- Add a scripted model and one agent test per tool that asserts both the final answer and the tool result in the second request.
- Add a boot test that resolves every configured model name, and contract tests from recorded responses for any custom adapter.
- Move model credentials into a protected environment and confirm the pull-request lane passes with no key.
- Build a live eval suite with at least three trials per case, gate on a Wilson lower bound, and run it nightly with every version pinned in the report.