Wiring evals into continuous integration means a prompt edit, a tool change or a dependency bump cannot merge while it quietly breaks a conversation you care about. For an agent built on the Agent Development Kit for Java, that runs into three practical problems. Real model calls are slow, cost money and give different answers on each run. Pull requests from forks get no secrets, so they cannot call the model at all. And a single pass or fail number hides the thing reviewers need: which cases changed compared with main.
This article solves those three problems with plumbing rather than new scoring. It assumes you already have an eval-set format, a replay harness on InMemoryRunner and scorers, as built in ADK Java Evaluation Framework, in depth, and the lane structure from ADK Java CI, in depth. Here we add four pieces: a cassette model that records and replays model responses, a stable request key so replays actually hit, JUnit and Maven wiring that produces reports CI understands, and a baseline gate that blocks on regressions rather than on absolute scores. As those pages note, ADK Java does not ship an evaluator in its core module, so everything here uses the documented model, runner and JSON classes; check the release notes in case that changes.
What CI needs from an eval run
A CI system needs four things from an eval run: the same result every time for the same code, so a red build means the code changed; no credentials on untrusted branches; machine-readable output, meaning JUnit XML, a JSON report and a Markdown summary; and a reference point, because a score of 0.86 means nothing until you know it was 0.91 yesterday.
The design that meets all four splits the model out of the loop. On pull requests the agent talks to a cassette: a directory of recorded model responses keyed by a hash of the request. Your agent code, instructions, tools and scorers all run for real; only the model is replayed. On a schedule, or by manual dispatch, the same harness runs against the real model, either refreshing the cassettes or measuring live behaviour. Both lanes write the same report, and a gate compares it with a baseline file kept in the repository.
A cassette model that records and replays
ADK Java models extend BaseLlm, whose generateContent(LlmRequest, boolean) returns a Flowable<LlmResponse>. A wrapper with the same type can sit between the agent and the real model. Both LlmRequest and LlmResponse carry Jackson annotations, and JsonBaseModel.getMapper() exposes the mapper ADK itself uses, so recording needs no custom serialiser. The live model is supplied lazily so replay mode never constructs a client or looks for an API key.
final class CassetteLlm extends BaseLlm {
enum Mode { REPLAY, RECORD, LIVE }
private static final ObjectMapper M = JsonBaseModel.getMapper();
private static final TypeReference<List<LlmResponse>> LIST = new TypeReference<>() {};
private final Supplier<BaseLlm> live;
private final Path dir;
private final Mode mode;
CassetteLlm(String name, Supplier<BaseLlm> live, Path dir, Mode mode) {
super(name);
this.live = live; this.dir = dir; this.mode = mode;
}
@Override public Flowable<LlmResponse> generateContent(LlmRequest req, boolean stream) {
if (mode == Mode.LIVE) return live.get().generateContent(req, false);
String key = RequestKey.of(req);
Path file = dir.resolve(key + ".json");
if (mode == Mode.REPLAY) {
if (!Files.exists(file)) {
return Flowable.error(new IllegalStateException(
"cassette miss " + key + ": re-record with EVAL_MODEL_MODE=RECORD"));
}
return Flowable.defer(() -> Flowable.fromIterable(M.readValue(file.toFile(), LIST)));
}
return live.get().generateContent(req, false).toList()
.doOnSuccess(rs -> M.writerWithDefaultPrettyPrinter().writeValue(file.toFile(), rs))
.flatMapPublisher(Flowable::fromIterable);
}
@Override public BaseLlmConnection connect(LlmRequest req) {
throw new UnsupportedOperationException("live bidi sessions are not recorded");
}
}Streaming is forced off when recording, so replays do not depend on the provider's chunking. A miss is an error, not a fallback to the live model, which would make pull requests need credentials again.
Making the request key stable
A cassette is only as good as its key. Hash the serialised request naively and you will find that almost nothing replays, for three reasons. JSON object key order is not guaranteed to be stable across library versions. Function call ids are generated per run, and from the second model call of a turn onward they sit inside the conversation history in functionCall and functionResponse parts. And anything you inject into instructions at run time, such as today's date or a request id, changes the request on every run.
The first two are fixed in the key function: sort keys recursively and drop only the id field inside those two part types. The third is fixed in the agent: make injected values come from session state, and give the eval cases a fixed state, which the eval-set format already supports.
final class RequestKey {
private static final ObjectMapper M = JsonBaseModel.getMapper();
static String of(LlmRequest req) {
JsonNode canon = canon(M.valueToTree(req), "");
// Guava is already on the classpath through ADK
return Hashing.sha256().hashString(canon.toString(), StandardCharsets.UTF_8)
.toString().substring(0, 24);
}
private static JsonNode canon(JsonNode n, String field) {
if (n.isObject()) {
ObjectNode out = M.createObjectNode();
List<String> names = new ArrayList<>();
n.fieldNames().forEachRemaining(names::add);
Collections.sort(names);
boolean callPart = field.equals("functionCall") || field.equals("functionResponse");
for (String k : names) {
if (callPart && k.equals("id")) continue; // regenerated every run
out.set(k, canon(n.get(k), k));
}
return out;
}
if (n.isArray()) {
ArrayNode out = M.createArrayNode();
n.forEach(x -> out.add(canon(x, field)));
return out;
}
return n;
}
}Commit a test for the key itself: build the same agent twice, run one case in RECORD mode against a scripted model, then run it in REPLAY mode and assert zero misses. That test catches the day an ADK upgrade adds a new volatile field to the request, which otherwise shows up as every cassette missing at once.
Running cases as JUnit tests
Run the eval cases as JUnit dynamic tests so that each case appears by name in the CI test tab, and tag them so the normal build skips them. The test records its result in a shared report and still asserts, so a developer running it locally sees failures in the IDE. In CI the build is told not to stop on test failures, because the decision belongs to the baseline gate, which can tell a regression from a case that was already failing.
@Tag("eval")
class SupportAgentEvalTest {
static final EvalReport REPORT = new EvalReport();
@TestFactory
Stream<DynamicTest> supportCases() throws IOException {
var mode = CassetteLlm.Mode.valueOf(System.getenv().getOrDefault("EVAL_MODEL_MODE", "REPLAY"));
BaseLlm model = new CassetteLlm("gemini-under-test", Models::production,
Path.of("src/test/resources/cassettes/support"), mode);
BaseAgent agent = SupportAgent.build(model);
return EvalSetLoader.load(Path.of("src/test/resources/evals/support.evalset.json")).stream()
.map(ec -> DynamicTest.dynamicTest(ec.id(), () -> {
CaseResult r = Scorers.score(ec, Harness.replay(agent, ec));
REPORT.add(r);
assertTrue(r.passed(), r::explain);
}));
}
@AfterAll
static void write() throws IOException {
REPORT.writeJson(Path.of("target/eval-report.json"));
}
}In Maven, set Surefire's <excludedGroups>${eval.excluded}</excludedGroups> with the property defaulting to eval, and add an evals profile that empties that property and sets <groups>eval</groups> and <testFailureIgnore>true</testFailureIgnore>. Hard-coding the exclusion would survive the profile merge, and exclusion wins, so nothing would run. Surefire still writes target/surefire-reports/TEST-*.xml, so the CI test view works unchanged.
The baseline gate
The baseline is a small JSON file in the repository, evals/baseline.json, holding each case's pass flag and scores from the last accepted run. Keeping it in git rather than in a CI artifact has two advantages: artifacts expire, and a change to the baseline becomes a reviewable diff in the same pull request that changed behaviour. The gate applies three rules.
- Regression: a case that passed in the baseline and fails now blocks the merge.
- New failure: a case that is not in the baseline and fails blocks, unless it is listed as known failing, which makes adding a test for an unfixed bug possible without breaking the build.
- Drift: the mean trajectory or response score dropping by more than a set margin blocks, even if no single case flipped. This catches slow erosion that per-case thresholds miss.
public final class BaselineGate {
static final double MAX_MEAN_DROP = 0.02;
public static void main(String[] args) throws IOException {
var base = EvalReport.read(Path.of("evals/baseline.json"));
var now = EvalReport.read(Path.of("target/eval-report.json"));
List<String> blocking = new ArrayList<>(), notes = new ArrayList<>();
for (CaseResult r : now.cases()) {
CaseResult b = base.get(r.id());
if (b == null && !r.passed() && !base.knownFailing().contains(r.id())) blocking.add("new failure: " + r.id());
else if (b != null && b.passed() && !r.passed()) blocking.add("regression: " + r.id() + " " + r.explain());
else if (b != null && !b.passed() && r.passed()) notes.add("now passing: " + r.id());
}
double drop = base.meanTrajectory() - now.meanTrajectory();
if (drop > MAX_MEAN_DROP) blocking.add(String.format("mean trajectory fell by %.3f", drop));
Files.writeString(Path.of("target/eval-summary.md"), Summary.markdown(base, now, blocking, notes));
if (!blocking.isEmpty()) System.exit(1);
}
}Updating the baseline is a deliberate act, never automatic. When cases improve, the author regenerates the file with a flag of your own, such as -Dupdate.baseline=true, and commits it. A pipeline that rewrites the baseline on every green build converts every regression it misses into the new normal.
The workflow and fork pull requests
The workflow below has two jobs. The pull request job uses the pull_request trigger, which gives fork branches no repository secrets; that is fine because replay needs none. Do not switch to pull_request_target to get secrets: it runs with the base repository's credentials, and checking out and building the fork's code in that context hands your model key to whoever opened the pull request. The scheduled or dispatched job uses a protected environment for the key, and either measures live behaviour or uploads re-recorded cassettes as an artifact for a maintainer to commit.
name: agent-evals
on:
pull_request:
schedule:
- cron: "0 3 * * *"
workflow_dispatch:
inputs:
mode: { type: choice, options: [LIVE, RECORD], default: LIVE }
jobs:
replay:
if: github.event_name == 'pull_request'
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-java@v4
with: { distribution: temurin, java-version: "21", cache: maven }
- run: mvn -B -Pevals verify
env: { EVAL_MODEL_MODE: REPLAY }
- run: mvn -B -q exec:java -Dexec.classpathScope=test -Dexec.mainClass=evals.BaselineGate
- if: always()
run: cat target/eval-summary.md >> "$GITHUB_STEP_SUMMARY"
- if: always()
uses: actions/upload-artifact@v4
with: { name: eval-report, path: target/eval-report.json }
live:
if: github.event_name != 'pull_request'
runs-on: ubuntu-latest
environment: model-evals
steps:
- uses: actions/checkout@v4
- uses: actions/setup-java@v4
with: { distribution: temurin, java-version: "21", cache: maven }
- run: mvn -B -Pevals verify
env:
EVAL_MODEL_MODE: ${{ inputs.mode || 'LIVE' }}
GOOGLE_API_KEY: ${{ secrets.GOOGLE_API_KEY }}
- if: inputs.mode == 'RECORD'
uses: actions/upload-artifact@v4
with: { name: cassettes, path: src/test/resources/cassettes }
- run: mvn -B -q exec:java -Dexec.classpathScope=test -Dexec.mainClass=evals.BaselineGateThe summary goes to $GITHUB_STEP_SUMMARY instead of a pull request comment because the token on a fork pull request is read-only and cannot comment. If the suite grows past a few minutes, shard it with a job matrix over eval-set files and merge the per-shard reports before the gate; sharding by file keeps each cassette directory owned by one shard.
Worked example: a tool description change
Take a support agent with 40 cases. On Monday a developer tightens the refund tool's description to say amounts are in cents. The tool declaration lives in the request config, so the request hash changes for the 17 cases routed to the billing sub-agent that offers the refund tool, and the pull request lane fails with 17 cassette misses. That is the system telling the truth: the model has not seen this prompt. A maintainer dispatches RECORD mode on the branch and commits the 17 uploaded cassette files.
The re-run replays cleanly and the gate reports one regression: the case where the customer says "refund 20 dollars" now calls the tool with amount=20 instead of 2000, because the instruction still talks about dollars. The developer fixes the instruction, re-records, and the gate goes green with a note that two previously failing cases now pass. They update the baseline in the same pull request, so the reviewer sees cassette, baseline and code diffs together.
Failure modes
| Failure | Symptom | Fix |
|---|---|---|
| Volatile field in the request | Every case misses after an upgrade or on every run | Extend the key canonicaliser; keep the round-trip key test |
| Silent live fallback on miss | Fork PRs fail on missing key; costs appear on PR builds | Make a miss an error in REPLAY mode |
| Stale cassettes | Replays pass while live behaviour has drifted | Nightly LIVE run against the same baseline; refresh on a schedule |
| Recorded secrets or PII | Customer data in git through cassettes | Use synthetic eval data; scan cassettes in the pre-commit hook |
pull_request_target for secrets | Untrusted code runs with your model key | Keep PRs on replay; credentials only in a protected environment |
| Auto-updated baseline | Regressions become the new normal | Baseline changes only through reviewed commits |
| Build stops before the gate | Red build with no summary | Ignore test failures in the eval profile; let the gate decide |
Trade-offs
Replay makes the pull request lane fast, free and deterministic, but it only proves your code handles the responses the model gave when they were recorded. It cannot show how a new prompt is received; that is why a miss must force a re-record rather than reuse an old answer. Recording outside pull requests keeps credentials safe but adds a manual step when prompts change. Live nightly runs catch provider-side change but cost money and need repeats and a statistical threshold. If you need a judge model for response quality, its calls can go through the same cassette, which makes an LLM-as-judge scorer replayable as well.
What to do next
- Wrap your production model in a cassette model with REPLAY, RECORD and LIVE modes, supplied lazily so replay needs no credentials.
- Write the request key with recursive key sorting and function-call id stripping, plus a round-trip test that asserts zero misses.
- Move run-time values such as dates out of instructions and into session state fixed by each eval case.
- Tag eval tests, add an
evalsprofile that ignores test failures, and write a JSON report after all cases. - Commit
evals/baseline.jsonand a gate with regression, new-failure and drift rules, writing a Markdown summary. - Add the two-job workflow: replay on pull requests, live or record on a schedule behind a protected environment. Compare it with the deployment flow in the ADK Java CI/CD pipeline.