Continuous integration for an ordinary Java service answers one question: did this change break anything? For an agent built with the Agent Development Kit for Java the question splits in two. Some of the system is ordinary code, tools, configuration, wiring, and can be tested deterministically on every commit. The rest is behaviour that emerges from a model reading your instructions, which is nondeterministic, costs money per run, needs credentials, and can change without any commit at all when a provider updates a model.

A pipeline that treats both halves the same fails in one of two ways. Either it calls the real model on every pull request and becomes slow, flaky and expensive, or it never calls the model and ships prompt regressions. This article builds a pipeline that separates them: a fast deterministic lane driven by a scripted model, a merge lane that builds and smoke-tests the real artifact, and a scheduled live-evaluation lane whose gate is statistical rather than a single pass or fail. Code uses ADK's documented runner and model APIs and plain JUnit 5; nothing here depends on an ADK evaluation tool for Java.

Advertisement

What can break, and where each failure should be caught

List the ways an ADK Java agent breaks before designing the pipeline, because each belongs to a different lane.

FailureExampleCheapest place to catch it
Tool code bugOrder lookup throws on a lowercase idTool unit test, every PR
Tool declaration driftParameter names arrive as arg0 because a compiler flag was droppedAgent test with a scripted model, every PR
Wiring errorAgent references a model name no registry pattern matchesBoot test, every PR
Adapter mapping bugTool-call ids lost when translating a provider responseContract test with recorded JSON, every PR
Credential or network problemExpired key, blocked egress from the clusterLive smoke test, merge and deploy
Behaviour regressionA prompt edit makes the agent skip a required toolLive eval suite, nightly and pre-release
Upstream model changeProvider updates the model behind an aliasScheduled live eval with no code change
Three lanes: cheap and deterministic on every change, expensive and statistical on a scheduleEvery pull request (no secrets, no network to the model)Compile-parameters onTool unit testsplain JUnitAgent testsScriptedLlmBoot testmodels resolveContractrecorded JSONMerge to mainRe-run PR laneon merged codeBuild imagedigest recordedLive smoke1 real tool turnNightly and before release (secrets in a protected environment)Live eval suitecases x trialsScore + boundWilson lower boundGatebound vs floorReportartifact + trendrelease candidateA failure in the top lane is a code bug. A failure in the bottom lane may be the model, the prompt or the data,so it opens an investigation rather than silently blocking every developer.Lane contents are a recommended structure, not an ADK feature
The pipeline has three lanes. Only the bottom lane spends money or needs model credentials, and only it has a statistical gate.

Build configuration that tests depend on

One compiler setting is load-bearing. ADK builds a tool's declaration by reflecting over the Java method, and without the -parameters flag, reflection reports parameter names as arg0, arg1 and so on. The model then sees meaningless argument names, unless each has @Schema(name = ...), and a ToolContext parameter, which ADK recognises by the name toolContext, is no longer injected. The function tool guide explains the mechanism. Put the flag in the build, and add a test that fails if it disappears.

Split tests into tagged groups so each lane runs only what it should. Surefire's excludedGroups and groups settings work with JUnit 5 tags.

<properties>
  <adk.version><!-- pin an exact released version --></adk.version>
</properties>

<dependency>
  <groupId>com.google.adk</groupId>
  <artifactId>google-adk</artifactId>
  <version>${adk.version}</version>
</dependency>

<plugin>
  <artifactId>maven-compiler-plugin</artifactId>
  <configuration>
    <release>21</release>
    <parameters>true</parameters>          <!-- tool parameter names survive reflection -->
  </configuration>
</plugin>

<plugin>
  <artifactId>maven-surefire-plugin</artifactId>
  <configuration>
    <excludedGroups>live</excludedGroups>  <!-- default build never calls a real model -->
  </configuration>
</plugin>

<profiles>
  <profile>
    <id>live</id>
    <build><plugins><plugin>
      <artifactId>maven-surefire-plugin</artifactId>
      <configuration>
        <groups>live</groups>
        <excludedGroups combine.self="override"/>
      </configuration>
    </plugin></plugins></build>
  </profile>
</profiles>

Pin the ADK version exactly and update it through a pull request, never through a version range. ADK Java moves quickly, and a minor version can change event shapes or callback order; you want that change to arrive with a green or red build attached to it.

Advertisement

Lane one: tools are just methods

Function tools in ADK Java are static methods wrapped with FunctionTool.create(Class, String). Test the method directly with ordinary JUnit: valid input, malformed input, a downstream timeout, and the shape of the returned map. Because the model reads the result map, its keys are part of the contract; assert them, not only the happy-path values. Tests at this level run in milliseconds and catch most real bugs.

Add one reflection test per tool that asserts the parameter names, so a missing compiler flag fails loudly instead of degrading the agent.

@Test
void toolParameterNamesSurviveCompilation() throws Exception {
  Method m = OrderTools.class.getMethod("getOrderStatus", String.class);
  assertEquals("orderId", m.getParameters()[0].getName(),
      "compile with -parameters, or tool arguments arrive as arg0");
}

Lane one: agent tests with a scripted model

The key to a deterministic agent test is replacing the model, not the agent. ADK models are subclasses of BaseLlm whose generateContent returns an RxJava Flowable<LlmResponse>. A scripted subclass returns queued responses in order and records every request it receives, as described in implementing a custom LLM. The agent, its instruction, its tools and the runner are all real.

final class ScriptedLlm extends BaseLlm {
  private final Queue<LlmResponse> script;
  final List<LlmRequest> seen = new CopyOnWriteArrayList<>();

  ScriptedLlm(LlmResponse... responses) {
    super("scripted");
    this.script = new ConcurrentLinkedQueue<>(List.of(responses));
  }

  @Override
  public Flowable<LlmResponse> generateContent(LlmRequest req, boolean stream) {
    return Flowable.defer(() -> {
      seen.add(req);
      LlmResponse next = script.poll();
      return next == null
          ? Flowable.error(new IllegalStateException("script exhausted"))
          : Flowable.just(next);
    });
  }

  @Override
  public BaseLlmConnection connect(LlmRequest req) {
    throw new UnsupportedOperationException();
  }
}

static LlmResponse modelSays(Part part) {
  return LlmResponse.builder()
      .content(Content.builder().role("model").parts(List.of(part)).build())
      .build();
}

@Test
void orderQuestionCallsToolThenAnswers() {
  ScriptedLlm model = new ScriptedLlm(
      modelSays(Part.builder().functionCall(FunctionCall.builder()
          .id("call-1").name("getOrderStatus").args(Map.of("orderId", "A-1001")).build()).build()),
      modelSays(Part.fromText("Order A-1001 shipped yesterday.")));

  LlmAgent agent = LlmAgent.builder()
      .name("orders")
      .model(model)
      .instruction("Use getOrderStatus for any question about an order.")
      .tools(FunctionTool.create(OrderTools.class, "getOrderStatus"))
      .build();

  InMemoryRunner runner = new InMemoryRunner(agent);
  Session session = runner.sessionService()
      .createSession(runner.appName(), "test-user").blockingGet();

  List<Event> events = runner.runAsync(session.userId(), session.id(),
          Content.fromParts(Part.fromText("Where is A-1001?")), RunConfig.builder().build())
      .toList().blockingGet();

  String answer = events.stream().filter(Event::finalResponse)
      .map(Event::stringifyContent).reduce("", String::concat);
  assertTrue(answer.contains("shipped"));

  // The second model call must carry the tool result back, paired to the call.
  boolean resultSent = model.seen.get(1).contents().stream()
      .flatMap(c -> c.parts().orElse(List.of()).stream())
      .anyMatch(p -> p.functionResponse().isPresent());
  assertTrue(resultSent, "tool result was not returned to the model");
}

What this test proves is narrow and valuable: that the declared tool can be called by name with these arguments, that the runner executes it, that its result flows back into the next model request, and that the final event carries the answer. It says nothing about whether a real model would choose the tool; that belongs to the live lane. Write scripted tests for every tool, for sub-agent transfers, for the error path where a tool returns an error status, and for callbacks that block or rewrite requests.

Assert on recorded requests as well as outputs. The system instruction, the tool declarations and the number of model calls are all visible in seen, and a change that silently drops a tool from the declaration list will show up there before any user notices.

Lane one: boot and contract tests

A boot test starts the application's real wiring with test configuration and resolves every model name the agents reference through the registry, failing the build if any name does not match a registered pattern. Without it, a misspelled model name or a missing custom-adapter registration fails on the first user request in production. The runtime boot sequence shows where resolution happens and why it is lazy by default.

Contract tests cover any custom BaseLlm adapter or external tool client. Record real provider responses once, including a turn with two parallel tool calls, a refusal, a rate-limit error and a truncated response, and replay them through the adapter's response mapper with an HTTP stub such as WireMock. Assert that ids, roles, finish reasons and usage metadata survive translation. Re-record deliberately when the provider changes its format, and review the diff like code.

Lane two: build the artifact once and smoke-test it

On merge, re-run lane one on the merged commit, build the container image, record its digest, and run one live smoke test against that exact image: a single real model call that performs one tool-calling turn and checks the answer contains a known fact from the tool. The smoke test is about plumbing, credentials, egress, region and model name, not quality. Run it again after each deployment, against the deployed environment, as described in Kubernetes deployment.

Model credentials never belong in the pull-request lane. On GitHub, workflows triggered by pull requests from forks do not receive repository secrets, which is the behaviour you want: lane one must pass with no key at all. Keep keys in a protected environment that only runs on main, scheduled or release triggers, and read them from the environment as configuration management recommends. The Gemini integration reads its key from the GOOGLE_API_KEY environment variable.

name: agent-ci
on:
  pull_request:
  push:
    branches: [main]
  workflow_dispatch:              # manual pre-release run
  schedule:
    - cron: "17 3 * * *"          # nightly live eval, even with no commits

jobs:
  verify:
    runs-on: ubuntu-latest
    timeout-minutes: 20
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-java@v4
        with: { distribution: temurin, java-version: "21", cache: maven }
      - run: mvn -B -ntp verify     # lane one: no secrets, no live model

  live-eval:
    if: github.event_name == 'schedule' || github.event_name == 'workflow_dispatch'
    needs: verify
    runs-on: ubuntu-latest
    environment: live-model         # protected environment holds the key
    concurrency: live-eval          # never two paid runs at once
    timeout-minutes: 40
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-java@v4
        with: { distribution: temurin, java-version: "21", cache: maven }
      - run: mvn -B -ntp -Plive verify
        env:
          GOOGLE_API_KEY: ${{ secrets.GOOGLE_API_KEY }}
      - uses: actions/upload-artifact@v4
        if: always()
        with: { name: eval-report, path: target/eval-report.json }

Lane three: a live eval with a statistical gate

The live suite is a set of cases, each a user message plus checks on the outcome: which tools were called with which arguments, whether the final answer contains or omits specific facts, whether a refusal happened when it should. Tool-call checks are the most reliable, because they are structured; text checks should be narrow. Each case runs several times, because one sample of a nondeterministic system tells you little.

Do not gate on the raw pass rate. A suite of 40 cases run 3 times each gives 120 trials, and a pass rate of 92.5 percent on one night and 89 percent the next may be noise. Gate on a lower confidence bound instead, which accounts for sample size. The Wilson score interval is a good default for proportions.

static double wilsonLower(int passes, int n, double z) {
  double p = (double) passes / n;
  double denom = 1 + z * z / n;
  double centre = p + z * z / (2.0 * n);
  double margin = z * Math.sqrt(p * (1 - p) / n + z * z / (4.0 * n * n));
  return (centre - margin) / denom;
}

// 111 of 120 trials passed: raw 0.925, 95% lower bound about 0.864
boolean release = wilsonLower(111, 120, 1.96) >= 0.85;

Worked through: with 111 passes out of 120, the raw rate is 0.925 and the 95 percent Wilson lower bound is about 0.864, which clears a floor of 0.85. If the same rate came from only 40 trials, 37 passes, the bound would be about 0.80 and the gate would fail, correctly, because the evidence is thinner. Keep a few must-pass cases outside the statistical gate, such as never calling a refund tool without confirmation, and fail the run if any trial of those fails.

Pin everything the result depends on and write it into the report: model name and version, ADK version, instruction text hash, tool declaration hash, temperature and the case-file commit. A nightly run whose score drops with no change in any of those is evidence of an upstream model change, which is exactly what the scheduled trigger exists to detect.

Failure modes

  • Live calls in the PR lane. Builds become flaky, slow and expensive, and developers learn to re-run until green. Keep lane one hermetic.
  • Over-scripted tests. Asserting on exact prompt text makes every instruction tweak break dozens of tests. Assert on structure: tools declared, results returned, calls counted.
  • Gate on raw pass rate. Small suites flip between pass and fail on noise; teams then lower the threshold until it means nothing.
  • Unpinned model aliases. An alias that tracks the latest model makes yesterday's eval irrelevant to today's production.
  • Eval cases leaking into prompts. When failures are fixed by pasting the case into the instruction, the suite measures memorisation. Keep a held-out set that prompt authors do not see.
  • Budget runaway. A loop in an agent under test can make hundreds of model calls. Set a per-run call budget and a job timeout.

Trade-offs

DecisionOption AOption B
Model in PR testsScripted fake: fast, free, deterministic, says nothing about qualityReal model: catches behaviour changes, slow, costly, flaky
Eval frequencyNightly: catches upstream drift, costs every dayPer release: cheaper, drift found late
Trials per caseOne: cheap, gate dominated by noiseThree to five: tighter bounds, proportionally more cost
GateHard block on release: safe, can stall shipping on noiseAlert and human review: flexible, depends on discipline

What to do next

  1. Turn on the compiler's parameters flag and add the reflection test that fails if it is ever dropped.
  2. Write plain JUnit tests for every tool, including the error map the model sees.
  3. Add a scripted model and one agent test per tool that asserts both the final answer and the tool result in the second request.
  4. Add a boot test that resolves every configured model name, and contract tests from recorded responses for any custom adapter.
  5. Move model credentials into a protected environment and confirm the pull-request lane passes with no key.
  6. Build a live eval suite with at least three trials per case, gate on a Wilson lower bound, and run it nightly with every version pinned in the report.
Key takeaway: CI for an ADK Java agent works when deterministic and statistical checks live in different lanes. Every pull request runs hermetic tests: tools as plain methods, agents driven by a scripted BaseLlm through the real runner, a boot test that resolves every model, and contract tests from recorded responses. Merges build one image and smoke-test it against the real model. A nightly and pre-release live eval runs each case several times, gates on a lower confidence bound rather than a raw pass rate, and pins every version so a score drop points at its cause.