CI/CD Pipeline for ADK Java, in depth: a release manifest for agents, build once, staged evaluation, session-sticky canaries and rollback

By Sandeep Belgavi · 2026-10-03 · Category: ADK for Java
Advertisement
Commitcode, prompts, configVerifyunit, scripted modelPackage onceimage digest, SBOMRelease manifestdigest, model, prompt hashStagingdeploy digest, replay evalCanarysticky sessions, 5 percentProductionsame digest, full trafficpromote by digestgategateGateseval non-inferiority, tool errors, cost, latencyRollbackredeploy previous manifest, not a rebuildEverything that changes agent behaviour travels in one manifest and is promoted unchanged from staging to production.
The delivery half of an ADK Java pipeline. Test lanes run in Verify; this article covers what happens after the artefact exists.

A conventional Java service changes behaviour when its code changes. An ADK Java agent changes behaviour when its code changes, when its instruction text changes, when the model behind a model id is upgraded, when a tool's declaration is reworded, and when configuration such as temperature or a retrieval index moves. A CI/CD pipeline that only versions the jar will happily promote a build whose behaviour nobody tested, and roll back to a build that no longer behaves the way it did last week.

This article designs the delivery side of the pipeline for ADK Java agents: what a release contains, how to build it once and reproducibly, how to promote the exact same artefact through staging and a canary, which metrics should gate each step, how to keep conversation state compatible across versions, and how to roll back. The test lanes that run before packaging, including deterministic agent tests with a scripted model and a statistically gated live evaluation, are covered in ADK Java CI; here we assume they exist and build on them.

What an agent release actually contains

Start by listing everything that can change what the agent says or does, because each item must be versioned and travel with the release:

ComponentExampleHow it drifts if unpinned
Application codeagent graph, tool implementationsrarely: it is already in git
Instruction textsystem instruction, few-shot examplesedited in a console or config store outside review
Model identifiera specific model version vs a floating aliasprovider updates the alias under you
Tool declarationsnames, descriptions, parameter schemasa reworded description changes tool choice
Generation configtemperature, max output tokens, safety settingsenvironment-specific overrides
Dependenciesthe ADK library and its transitive treeversion ranges resolve differently
External dataretrieval index, reference documentsre-indexed without a version

The rule that follows is simple: anything in this table that lives outside the image must be referenced from the release by an immutable identifier. Keep instruction text as files in the repository, loaded from the classpath, so they are reviewed like code. Pin model ids to the most specific version your provider offers. Refer to retrieval indexes by snapshot name. Then compute a release manifest at build time that records all of it.

Advertisement

The release manifest

The manifest is a small JSON file baked into the image and also stored next to the image in your registry or release system. It answers the question every incident starts with: what exactly was running? A build step can generate it from the repository:

{
  "service": "support-agent",
  "gitSha": "9c41e0d",
  "imageDigest": "sha256:<filled in by the package job>",
  "adkVersion": "<value of the adk.version property>",
  "model": "<pinned model version id>",
  "promptSha256": {
    "prompts/support/instruction.md": "e3b0c442...",
    "prompts/support/examples.md": "5f70bf18..."
  },
  "toolSchemaSha256": "a54d88e0...",
  "retrievalIndex": "kb-snapshot-2026-09-28",
  "evalReport": "<id of the staging eval run>"
}

At startup the service reads the same manifest and builds the agent from it, so the running agent cannot quietly differ from the record. A minimal factory looks like this; the builder calls mirror the ones used in the CI article, and the manifest record is your own class:

public final class AgentFactory {
    public static LlmAgent fromManifest(ReleaseManifest m, BaseTool... tools)
            throws IOException, NoSuchAlgorithmException {
        byte[] bytes;
        try (InputStream in = AgentFactory.class.getResourceAsStream("/prompts/support/instruction.md")) {
            if (in == null) throw new IllegalStateException("instruction resource missing");
            bytes = in.readAllBytes();
        }
        String instruction = new String(bytes, StandardCharsets.UTF_8);
        String actual = HexFormat.of().formatHex(MessageDigest.getInstance("SHA-256").digest(bytes));
        String expected = m.promptSha256().get("prompts/support/instruction.md");
        if (!actual.equals(expected)) {
            // the image and its manifest disagree: refuse to start rather than drift
            throw new IllegalStateException("instruction hash mismatch: " + actual);
        }
        return LlmAgent.builder()
            .name("support_agent")
            .model(m.model())
            .instruction(instruction)
            .tools(tools)
            .build();
    }
}

Two lines carry the design. The hash check turns a mismatched image into a failed readiness probe instead of a silent behaviour change. The model id comes from the manifest, not from an environment variable, so an operator cannot change the model of a running release without producing a new release. Environment-specific values that genuinely differ, such as endpoints and quotas, still come from configuration, as described in environment and configuration management, but behaviour-defining values do not.

A reproducible build, promoted by digest

Build once means the bytes you tested are the bytes you run. That needs a reproducible build. In Maven, pin the ADK version through a property, ban version ranges and snapshot dependencies in release builds, and set a fixed output timestamp so jar entries do not embed build time:

<properties>
  <adk.version><!-- pin an exact released version --></adk.version>
  <project.build.outputTimestamp>2026-01-01T00:00:00Z</project.build.outputTimestamp>
</properties>

<dependencies>
  <dependency>
    <groupId>com.google.adk</groupId>
    <artifactId>google-adk</artifactId>
    <version>${adk.version}</version>
  </dependency>
</dependencies>

Containerise with a tool that produces deterministic layers, such as Jib, which builds images straight from Maven without a Dockerfile and records the pushed image digest in the build output directory. From that point on, the pipeline refers to the image only by digest, never by a mutable tag. Generate a software bill of materials and sign the digest so the cluster can refuse unsigned images. A useful sanity check is to build the same commit twice in CI and compare digests; if they differ, something non-deterministic, often a timestamp or an unordered resource, has crept into the build and your rollback guarantees are weaker than you think.

The workflow, stage by stage

Here is the shape of a GitHub Actions workflow implementing these stages. Cloud authentication uses OIDC federation, so the repository holds no long-lived cloud keys; the exact auth step depends on your cloud and is shown as a placeholder. Production is a protected environment, which gives you required reviewers and an audit trail for every promotion:

name: agent-release
on:
  push:
    branches: [main]
concurrency:
  group: agent-release
  cancel-in-progress: false        # never cancel a half-finished promotion
permissions:
  contents: read
  id-token: write                  # OIDC token for cloud federation

jobs:
  verify:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-java@v4
        with: { distribution: temurin, java-version: '21', cache: maven }
      - run: mvn -B verify          # unit, scripted-model and contract tests

  package:
    needs: verify
    runs-on: ubuntu-latest
    outputs:
      digest: ${{ steps.push.outputs.digest }}
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-java@v4
        with: { distribution: temurin, java-version: '21', cache: maven }
      # - authenticate to the registry with your cloud's OIDC action here
      - id: push
        run: |
          mvn -B -DskipTests package jib:build -Dimage=$REGISTRY/support-agent:$GITHUB_SHA
          echo "digest=$(cat target/jib-image.digest)" >> "$GITHUB_OUTPUT"
      - run: ./ci/write-manifest.sh "$(cat target/jib-image.digest)"

  staging:
    needs: package
    environment: staging
    runs-on: ubuntu-latest
    steps:
      - run: ./ci/deploy.sh staging "${{ needs.package.outputs.digest }}"
      - run: ./ci/replay-eval.sh staging --baseline production

  production:
    needs: [package, staging]
    environment: production          # required reviewers configured on the environment
    runs-on: ubuntu-latest
    steps:
      - run: ./ci/deploy.sh production "${{ needs.package.outputs.digest }}" --canary 5
      - run: ./ci/watch-canary.sh --minutes 60
      - run: ./ci/deploy.sh production "${{ needs.package.outputs.digest }}" --full

The deploy and evaluation scripts are yours; the important property is that the staging and production jobs receive the digest from the package job and never rebuild. concurrency with cancellation disabled prevents two pushes from interleaving promotions, which otherwise leads to a canary of one build being promoted with the manifest of another.

Gates: paired evaluation and agent-aware canaries

Staging runs the candidate against a fixed replay set: a few hundred recorded or curated conversations, re-executed against the real model, with each answer scored by deterministic checks where possible and by a calibrated judge where not; the LLM-as-judge scorer article covers building that judge. The gate is relative, not absolute. Run the current production manifest against the same set in the same job, and require the candidate to be non-inferior within a margin you chose in advance. Absolute thresholds drift with the model and the dataset; a paired comparison on identical inputs does not.

The canary then watches live traffic, and agents need different signals from web services. Route by session, not by request: an agent conversation spans many turns, and switching versions mid-session mixes two instructions and two tool sets in one transcript. Gate on metrics that reflect agent behaviour:

SignalWhy it mattersTypical gate
Tool call error ratea reworded schema causes malformed argumentsno worse than baseline plus a small margin
Turns per resolved sessionloops and confusion show up as extra turnswithin 10 percent of baseline
Tokens and cost per sessionprompt growth and retries cost moneywithin budget; alert on step changes
Escalation or handoff rateusers giving up on the agentno significant increase
p95 end-to-end latencymore tool calls, longer outputswithin SLO
Safety and policy blocksnew instruction triggers more refusalsno significant increase

Sixty minutes at five percent is enough for high-volume agents and far too little for low-volume ones; compute how many sessions you need to detect the regression size you care about, and extend the canary window to reach it rather than promoting on noise.

Session state compatibility and rollback

Agents carry state between turns, in session state and often in persistent memory, and the canary means two versions read and write it at once. Treat session state like a database schema. Use expand and contract: a release may add new state keys and must tolerate missing ones, and may stop writing an old key only after every running version has stopped reading it. Never rename a key in one release. Store a schema version in the session so a reader can detect what wrote it.

Rollback is the reason the manifest exists. Rolling back means redeploying the previous manifest's image digest, not rebuilding the previous commit, which may resolve different dependencies today. Keep the instruction, model id and tool schemas of a release together; rolling back code while keeping a new instruction is a new, untested release. One trap is specific to model-backed systems: providers retire model versions, so an old manifest can become undeployable. Track retirement dates for every model id in a live manifest and requalify a replacement before the date, not after. For the cluster mechanics of draining sessions during a rollout, see ADK Java on Kubernetes.

Failure modes and trade-offs

The trade-off is speed against confidence. Replay evaluations against a real model cost money and minutes on every merge; many teams run a small smoke set per merge and the full set nightly and before production promotion. Manifest discipline adds friction to quick prompt tweaks, and that friction is the point: a prompt edit is a behaviour change and deserves the same path as a code change.

What to do next

  1. List every behaviour-defining input of your agent using the table above, and move instruction text into the repository if it lives anywhere else.
  2. Pin the ADK version through a property, pin the model id to a specific version, and set a fixed Maven output timestamp; build the same commit twice and compare digests.
  3. Generate a release manifest at build time and add the startup hash check so a mismatched image fails readiness.
  4. Change every deploy step to take an image digest from the package job; remove any rebuild from staging or production jobs.
  5. Add a paired replay evaluation in staging against the current production manifest, with a pre-agreed non-inferiority margin.
  6. Configure a session-sticky canary gated on tool errors, turns, cost, escalations and latency, sized by the number of sessions needed.
  7. Adopt expand-and-contract for session state, and rehearse a rollback to the previous manifest in staging every month.
Key takeaway: An agent's behaviour is defined by code, instruction text, model id, tool schemas and configuration together, so the pipeline must version and promote all of them as one unit. Record them in a release manifest, build once and promote by image digest, gate staging on a paired evaluation against production, run session-sticky canaries on agent metrics, keep session state backward compatible, and roll back by redeploying the previous manifest rather than rebuilding.