A traditional service changes when its code changes. An agent changes when any of six things change: its Java code, its instruction prompts, its tool declarations, the model version behind it, the ADK release it runs on, and the knowledge it retrieves from. Some of those change without a commit at all, as when a provider retires a model or the retrieval index is rebuilt overnight. A CI/CD pipeline that runs the same steps for every change is therefore either too slow and expensive, because it pays for a live evaluation on a logging refactor, or too weak, because it skips evaluation on the prompt edit that actually changes behaviour.

This article is the map and the router. The individual gates have their own pages: ADK Java CI covers deterministic tests with a scripted model, integrating evals into CI covers recorded cassettes, the CI/CD pipeline article covers the release manifest and promotion by image digest, and the deploy pipeline covers ramps and rollback. Here you will build the step that connects them: fingerprint the agent, diff the fingerprint against production, and derive the set of gates the change must pass.

Six surfaces that change an agent

Start from first principles: what can change the agent's behaviour, how likely is each kind of change to break it, and what is the cheapest test that would notice?

SurfaceChanges whenWhat can breakCheapest gate that catches it
Java codea committool logic, state handling, wiringunit and contract tests with a scripted model
Instruction promptsa commit to prompt filestool choice, tone, refusals, formatevaluation; cassettes no longer match, so live sample
Tool declarationsmethod signature, description or schema editswhen and how the model calls toolscontract test plus evaluation
Model versionconfig change or provider retirementalmost anythingfull evaluation and a longer canary
ADK versiondependency bumprequest assembly, callbacks, session handlingfull test suite plus replay evaluation
Knowledge snapshotindex rebuild, often scheduledanswers that depend on retrieved factslive evaluation on retrieval cases

Two observations follow. First, the expensive gates (live evaluation, extended canary) are only necessary for a minority of changes; most commits touch code only. Second, two rows can change with no commit, so the pipeline needs triggers other than pull requests.

Architecture: plan first, then gate

The pipeline adds a planning step at the front. It computes a fingerprint of the candidate, loads the fingerprint recorded when the current production release was promoted, diffs the two, and emits a gate plan. Every later job reads the plan and runs or skips itself.

Change-aware agent CI/CD: fingerprint, diff, then run only the gates the change needsPull request / mergeNightly scheduledrift checkModel or KB changescheduled or manualFingerprintmodel, ADK, code, prompts, tools, KBProduction fingerprintfrom the release recordDiff and planfail closed if unsureUnit + contractReplay evalLive evalSmoke on artifactCanary: standard or extendedPromote digest, record new production fingerprintCheap gates always run. Expensive gates run when the diff touches the surface they protect.Red marks the gate that spends model tokens and needs a protected secret.
Triggers feed a fingerprint, the diff against production yields a plan, and the plan selects gates. Promotion records the new production fingerprint.

The production fingerprint lives next to the release manifest described in the CI/CD pipeline article; it is a hash summary of the same surfaces. Recording it at promotion time, rather than recomputing it from the main branch, matters: main may contain merged-but-unreleased changes, and you want to compare against what users are actually talking to.

Fingerprinting what the model sees

Fingerprint what the model actually sees, not the source files that produce it. A tool's declaration is generated from the Java method by ADK, so renaming a parameter or editing a description changes the declaration even though no prompt file changed. In ADK Java every tool extends BaseTool, whose declaration() returns an Optional<FunctionDeclaration>; the GenAI types serialise with toJson(). Hash a canonical form, with object keys sorted, so that a change in serialisation order does not look like a behaviour change.

public record AgentFingerprint(
    String model, String adkVersion, String code,
    Map<String, String> prompts, Map<String, String> tools, String knowledge) {

  private static final ObjectMapper JSON = new ObjectMapper();

  /** Hash each tool's declaration exactly as the model will receive it. */
  static Map<String, String> toolHashes(List<BaseTool> tools) throws Exception {
    Map<String, String> out = new TreeMap<>();
    for (BaseTool t : tools) {
      String decl = t.declaration()
          .map(d -> d.toJson())
          .orElse("{\"name\":\"" + t.name() + "\"}");    // built-in tools may have none
      out.put(t.name(), sha256(canonical(JSON.readTree(decl))));
    }
    return out;
  }

  /** Recursively sort object keys so the hash ignores serialisation order. */
  static String canonical(JsonNode n) throws Exception {
    if (n.isObject()) {
      TreeMap<String, JsonNode> sorted = new TreeMap<>();
      n.fields().forEachRemaining(e -> sorted.put(e.getKey(), e.getValue()));
      StringBuilder sb = new StringBuilder("{");
      for (var e : sorted.entrySet()) {
        if (sb.length() > 1) sb.append(',');
        sb.append(JSON.writeValueAsString(e.getKey())).append(':').append(canonical(e.getValue()));
      }
      return sb.append('}').toString();
    }
    if (n.isArray()) {
      StringBuilder sb = new StringBuilder("[");
      for (JsonNode x : n) { if (sb.length() > 1) sb.append(','); sb.append(canonical(x)); }
      return sb.append(']').toString();
    }
    return JSON.writeValueAsString(n);
  }

  static String sha256(String s) throws Exception {
    return HexFormat.of().formatHex(
        MessageDigest.getInstance("SHA-256").digest(s.getBytes(StandardCharsets.UTF_8)));
  }
}

Build the tool list with the same factory the application uses, so the fingerprint cannot drift from production wiring. Fill the other fields from what you already have: the model id from configuration, the ADK version from the build, prompt hashes from the prompt resources, a hash of the compiled classes excluding tests for code, and the retrieval snapshot id for knowledge. Compile with -parameters as the CI article explains, or tool parameter names in the declaration become arg0 and arg1.

From diff to gate plan

The classifier is deliberately simple and deliberately pessimistic. Cheap gates always run. Each surface that changed adds the gates that protect it. If there is no production fingerprint to compare against, or anything about the diff is unclear, it returns every gate.

enum Gate { UNIT, CONTRACT, REPLAY_EVAL, LIVE_EVAL, SMOKE, CANARY_STANDARD, CANARY_EXTENDED }

static EnumSet<Gate> plan(AgentFingerprint prod, AgentFingerprint cand) {
  EnumSet<Gate> g = EnumSet.of(Gate.UNIT, Gate.CONTRACT, Gate.SMOKE);   // always
  if (prod == null) return EnumSet.allOf(Gate.class);                  // unknown base: fail closed
  if (!cand.code().equals(prod.code()))              g.add(Gate.REPLAY_EVAL);
  if (!cand.prompts().equals(prod.prompts())
      || !cand.tools().equals(prod.tools())) {
    g.add(Gate.REPLAY_EVAL);                         // cassettes keyed on the request will miss ...
    g.add(Gate.LIVE_EVAL);                           // ... so a live sample is needed too
  }
  if (!cand.knowledge().equals(prod.knowledge()))    g.add(Gate.LIVE_EVAL);
  if (!cand.model().equals(prod.model())
      || !cand.adkVersion().equals(prod.adkVersion())) {
    g.add(Gate.REPLAY_EVAL);
    g.add(Gate.LIVE_EVAL);
    g.add(Gate.CANARY_EXTENDED);
  }
  if (!g.contains(Gate.CANARY_EXTENDED)) g.add(Gate.CANARY_STANDARD);
  return g;
}

Why does a prompt or tool change force a live evaluation when a code change only needs replay? Replay cassettes are keyed on the request sent to the model. A code change that leaves the request identical replays cleanly, and a mismatch is itself a signal. A prompt or tool change alters every request, so every cassette misses; the only way to learn how the model reacts is to ask it, on a sample, and then re-record.

Wiring the plan into the workflow

In GitHub Actions the plan becomes a job output, a JSON array, and each expensive job guards itself with contains(fromJSON(...), 'GATE'). The scripted-model tests run unconditionally because they cost nothing.

name: agent-cicd
on:
  pull_request:
  push:
    branches: [main]
  schedule:
    - cron: "40 4 * * *"            # drift check: production fingerprint, no new code

jobs:
  plan:
    runs-on: ubuntu-latest
    outputs:
      gates: ${{ steps.plan.outputs.gates }}
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-java@v4
        with: { distribution: temurin, java-version: "21", cache: maven }
      - id: plan
        run: |
          mvn -B -ntp -q exec:java -Dexec.mainClass=ci.PlanGates \
            -Dexec.args="--prod ci/prod-fingerprint.json --event ${{ github.event_name }}" \
            > plan.json
          echo "gates=$(cat plan.json)" >> "$GITHUB_OUTPUT"   # e.g. ["UNIT","CONTRACT","SMOKE","REPLAY_EVAL"]

  verify:
    needs: plan
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-java@v4
        with: { distribution: temurin, java-version: "21", cache: maven }
      - run: mvn -B -ntp verify                       # UNIT and CONTRACT, no secrets

  replay-eval:
    needs: [plan, verify]
    if: contains(fromJSON(needs.plan.outputs.gates), 'REPLAY_EVAL')
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-java@v4
        with: { distribution: temurin, java-version: "21", cache: maven }
      - run: mvn -B -ntp -Preplay verify             # cassettes, no live model

  live-eval:
    needs: [plan, verify]
    if: contains(fromJSON(needs.plan.outputs.gates), 'LIVE_EVAL')
    environment: live-model                          # protected secret, approval if you want it
    concurrency: live-eval
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-java@v4
        with: { distribution: temurin, java-version: "21", cache: maven }
      - run: mvn -B -ntp -Plive verify
        env:
          GOOGLE_API_KEY: ${{ secrets.GOOGLE_API_KEY }}

Two details keep this safe. Required status checks must point at jobs that always report, because a skipped job counts as successful for branch protection; make the merge gate a final summary job that fails if any planned gate did not succeed. And the live-evaluation job uses a protected environment, so pull requests from forks cannot reach the model key.

Triggers that are not commits

Three triggers have nothing to do with commits.

  • Nightly drift check. Run the live evaluation against the production fingerprint with no changes. If scores fall, something outside your repository moved: a model served behind a moving alias, a tool's backend, or the knowledge index. Pinning exact model versions instead of aliases removes one cause.
  • Model lifecycle. Providers announce retirement dates for model versions. Put them in a calendar that opens a change proposing the successor version weeks ahead, so the full evaluation and extended canary run without deadline pressure.
  • Knowledge rebuilds. When the retrieval index is rebuilt, run the retrieval evaluation subset before the new snapshot id becomes the production one, exactly as for a code release.

Worked example: three pull requests

Three pull requests arrive on the same day.

Pull requestDiffPlanRough cost
A: structured logging in a toolcode only; declarations unchangedunit, contract, replay, smoke, standard canaryminutes; no tokens
B: clarify the refund tool descriptiontoolsadds live evaluationone sampled live run
C: move to a new model versionmodeleverything, extended canaryfull live run, a day of canary

Pull request A is the common case and finishes in minutes. Pull request B looks trivial, one sentence in the method's @Schema(description = ...) annotation, but it changes how the model decides to call the tool, and the fingerprint catches it because it hashes the generated declaration. Pull request C pays the full price, which it should.

To size the budget, multiply cases, trials and tokens. As an assumed example, 150 cases × 3 trials × 8,000 tokens per run is 3.6 million tokens per live evaluation. Running it on every pull request at 40 pull requests a week is 144 million tokens; running it only on the roughly one in five that touch prompts, tools, model or knowledge, plus seven nightly drift checks, is 15 runs, or 54 million. Use your own token counts and prices; the ratio, not the number, is the point.

Failure modes

  • Fingerprinting source instead of output. Hashing prompt files misses instructions assembled at runtime, for example by an instruction provider that reads state. Hash the template and test the assembly in contract tests.
  • Tools added outside the factory. A toolset or sub-agent registered elsewhere is invisible to the fingerprint. Build the fingerprint from the root agent's full tool and sub-agent tree.
  • Comparing against main, not production. Unreleased changes on main disappear from the diff and ship without their gates.
  • Fail-open classifier. A parsing error that yields an empty plan would skip every expensive gate. Unknown means all gates.
  • Skipped jobs satisfying branch protection. Use a summary job as the required check.
  • Re-recording cassettes in the same run that evaluates. The new recordings then certify themselves. Re-record only after the live gate passes.

Trade-offs

Change-aware gating saves time and tokens but adds a component that can be wrong. A pipeline that always runs everything is simpler to trust, and for a small team with a few changes a week it may be cheaper than maintaining a classifier. The classifier pays off when change volume is high, live evaluations are expensive, or when evaluation queues delay ordinary fixes. Whatever you choose, keep the fail-closed rule, and periodically run the full plan on a random sample of code-only changes to check that the classifier is not missing a surface. For rollout itself, ADK Java Canary Deployment covers session-pinned routing.

What to do next

  1. List your agent's behavioural surfaces; add any beyond the six here, such as sub-agent definitions or callback configuration.
  2. Implement the fingerprint from the application's own agent factory and print it in CI.
  3. Record the fingerprint at promotion, next to the release manifest.
  4. Add the planning job and guard each expensive job with the plan, plus a summary job as the required check.
  5. Add the nightly drift run and a model-retirement calendar.
  6. After a month, compare token spend and pipeline time against the old pipeline, and audit a sample of code-only changes with the full plan.
Key takeaway: An agent changes through code, prompts, tool declarations, model, ADK version and knowledge, and only some of those changes need expensive gates. Fingerprint what the model actually sees, diff it against the fingerprint recorded at the last promotion, run cheap gates always and expensive ones when their surface changed, fail closed when unsure, and add nightly and model-lifecycle triggers for the changes that arrive without a commit.