A conventional Java service changes behaviour when its code changes. An ADK Java agent changes behaviour when its code changes, when its instruction text changes, when the model behind a model id is upgraded, when a tool's declaration is reworded, and when configuration such as temperature or a retrieval index moves. A CI/CD pipeline that only versions the jar will happily promote a build whose behaviour nobody tested, and roll back to a build that no longer behaves the way it did last week.
This article designs the delivery side of the pipeline for ADK Java agents: what a release contains, how to build it once and reproducibly, how to promote the exact same artefact through staging and a canary, which metrics should gate each step, how to keep conversation state compatible across versions, and how to roll back. The test lanes that run before packaging, including deterministic agent tests with a scripted model and a statistically gated live evaluation, are covered in ADK Java CI; here we assume they exist and build on them.
Start by listing everything that can change what the agent says or does, because each item must be versioned and travel with the release:
| Component | Example | How it drifts if unpinned |
|---|---|---|
| Application code | agent graph, tool implementations | rarely: it is already in git |
| Instruction text | system instruction, few-shot examples | edited in a console or config store outside review |
| Model identifier | a specific model version vs a floating alias | provider updates the alias under you |
| Tool declarations | names, descriptions, parameter schemas | a reworded description changes tool choice |
| Generation config | temperature, max output tokens, safety settings | environment-specific overrides |
| Dependencies | the ADK library and its transitive tree | version ranges resolve differently |
| External data | retrieval index, reference documents | re-indexed without a version |
The rule that follows is simple: anything in this table that lives outside the image must be referenced from the release by an immutable identifier. Keep instruction text as files in the repository, loaded from the classpath, so they are reviewed like code. Pin model ids to the most specific version your provider offers. Refer to retrieval indexes by snapshot name. Then compute a release manifest at build time that records all of it.
The manifest is a small JSON file baked into the image and also stored next to the image in your registry or release system. It answers the question every incident starts with: what exactly was running? A build step can generate it from the repository:
{
"service": "support-agent",
"gitSha": "9c41e0d",
"imageDigest": "sha256:<filled in by the package job>",
"adkVersion": "<value of the adk.version property>",
"model": "<pinned model version id>",
"promptSha256": {
"prompts/support/instruction.md": "e3b0c442...",
"prompts/support/examples.md": "5f70bf18..."
},
"toolSchemaSha256": "a54d88e0...",
"retrievalIndex": "kb-snapshot-2026-09-28",
"evalReport": "<id of the staging eval run>"
}At startup the service reads the same manifest and builds the agent from it, so the running agent cannot quietly differ from the record. A minimal factory looks like this; the builder calls mirror the ones used in the CI article, and the manifest record is your own class:
public final class AgentFactory {
public static LlmAgent fromManifest(ReleaseManifest m, BaseTool... tools)
throws IOException, NoSuchAlgorithmException {
byte[] bytes;
try (InputStream in = AgentFactory.class.getResourceAsStream("/prompts/support/instruction.md")) {
if (in == null) throw new IllegalStateException("instruction resource missing");
bytes = in.readAllBytes();
}
String instruction = new String(bytes, StandardCharsets.UTF_8);
String actual = HexFormat.of().formatHex(MessageDigest.getInstance("SHA-256").digest(bytes));
String expected = m.promptSha256().get("prompts/support/instruction.md");
if (!actual.equals(expected)) {
// the image and its manifest disagree: refuse to start rather than drift
throw new IllegalStateException("instruction hash mismatch: " + actual);
}
return LlmAgent.builder()
.name("support_agent")
.model(m.model())
.instruction(instruction)
.tools(tools)
.build();
}
}Two lines carry the design. The hash check turns a mismatched image into a failed readiness probe instead of a silent behaviour change. The model id comes from the manifest, not from an environment variable, so an operator cannot change the model of a running release without producing a new release. Environment-specific values that genuinely differ, such as endpoints and quotas, still come from configuration, as described in environment and configuration management, but behaviour-defining values do not.
Build once means the bytes you tested are the bytes you run. That needs a reproducible build. In Maven, pin the ADK version through a property, ban version ranges and snapshot dependencies in release builds, and set a fixed output timestamp so jar entries do not embed build time:
<properties>
<adk.version><!-- pin an exact released version --></adk.version>
<project.build.outputTimestamp>2026-01-01T00:00:00Z</project.build.outputTimestamp>
</properties>
<dependencies>
<dependency>
<groupId>com.google.adk</groupId>
<artifactId>google-adk</artifactId>
<version>${adk.version}</version>
</dependency>
</dependencies>Containerise with a tool that produces deterministic layers, such as Jib, which builds images straight from Maven without a Dockerfile and records the pushed image digest in the build output directory. From that point on, the pipeline refers to the image only by digest, never by a mutable tag. Generate a software bill of materials and sign the digest so the cluster can refuse unsigned images. A useful sanity check is to build the same commit twice in CI and compare digests; if they differ, something non-deterministic, often a timestamp or an unordered resource, has crept into the build and your rollback guarantees are weaker than you think.
Here is the shape of a GitHub Actions workflow implementing these stages. Cloud authentication uses OIDC federation, so the repository holds no long-lived cloud keys; the exact auth step depends on your cloud and is shown as a placeholder. Production is a protected environment, which gives you required reviewers and an audit trail for every promotion:
name: agent-release
on:
push:
branches: [main]
concurrency:
group: agent-release
cancel-in-progress: false # never cancel a half-finished promotion
permissions:
contents: read
id-token: write # OIDC token for cloud federation
jobs:
verify:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-java@v4
with: { distribution: temurin, java-version: '21', cache: maven }
- run: mvn -B verify # unit, scripted-model and contract tests
package:
needs: verify
runs-on: ubuntu-latest
outputs:
digest: ${{ steps.push.outputs.digest }}
steps:
- uses: actions/checkout@v4
- uses: actions/setup-java@v4
with: { distribution: temurin, java-version: '21', cache: maven }
# - authenticate to the registry with your cloud's OIDC action here
- id: push
run: |
mvn -B -DskipTests package jib:build -Dimage=$REGISTRY/support-agent:$GITHUB_SHA
echo "digest=$(cat target/jib-image.digest)" >> "$GITHUB_OUTPUT"
- run: ./ci/write-manifest.sh "$(cat target/jib-image.digest)"
staging:
needs: package
environment: staging
runs-on: ubuntu-latest
steps:
- run: ./ci/deploy.sh staging "${{ needs.package.outputs.digest }}"
- run: ./ci/replay-eval.sh staging --baseline production
production:
needs: [package, staging]
environment: production # required reviewers configured on the environment
runs-on: ubuntu-latest
steps:
- run: ./ci/deploy.sh production "${{ needs.package.outputs.digest }}" --canary 5
- run: ./ci/watch-canary.sh --minutes 60
- run: ./ci/deploy.sh production "${{ needs.package.outputs.digest }}" --fullThe deploy and evaluation scripts are yours; the important property is that the staging and production jobs receive the digest from the package job and never rebuild. concurrency with cancellation disabled prevents two pushes from interleaving promotions, which otherwise leads to a canary of one build being promoted with the manifest of another.
Staging runs the candidate against a fixed replay set: a few hundred recorded or curated conversations, re-executed against the real model, with each answer scored by deterministic checks where possible and by a calibrated judge where not; the LLM-as-judge scorer article covers building that judge. The gate is relative, not absolute. Run the current production manifest against the same set in the same job, and require the candidate to be non-inferior within a margin you chose in advance. Absolute thresholds drift with the model and the dataset; a paired comparison on identical inputs does not.
The canary then watches live traffic, and agents need different signals from web services. Route by session, not by request: an agent conversation spans many turns, and switching versions mid-session mixes two instructions and two tool sets in one transcript. Gate on metrics that reflect agent behaviour:
| Signal | Why it matters | Typical gate |
|---|---|---|
| Tool call error rate | a reworded schema causes malformed arguments | no worse than baseline plus a small margin |
| Turns per resolved session | loops and confusion show up as extra turns | within 10 percent of baseline |
| Tokens and cost per session | prompt growth and retries cost money | within budget; alert on step changes |
| Escalation or handoff rate | users giving up on the agent | no significant increase |
| p95 end-to-end latency | more tool calls, longer outputs | within SLO |
| Safety and policy blocks | new instruction triggers more refusals | no significant increase |
Sixty minutes at five percent is enough for high-volume agents and far too little for low-volume ones; compute how many sessions you need to detect the regression size you care about, and extend the canary window to reach it rather than promoting on noise.
Agents carry state between turns, in session state and often in persistent memory, and the canary means two versions read and write it at once. Treat session state like a database schema. Use expand and contract: a release may add new state keys and must tolerate missing ones, and may stop writing an old key only after every running version has stopped reading it. Never rename a key in one release. Store a schema version in the session so a reader can detect what wrote it.
Rollback is the reason the manifest exists. Rolling back means redeploying the previous manifest's image digest, not rebuilding the previous commit, which may resolve different dependencies today. Keep the instruction, model id and tool schemas of a release together; rolling back code while keeping a new instruction is a new, untested release. One trap is specific to model-backed systems: providers retire model versions, so an old manifest can become undeployable. Track retirement dates for every model id in a live manifest and requalify a replacement before the date, not after. For the cluster mechanics of draining sessions during a rollout, see ADK Java on Kubernetes.
The trade-off is speed against confidence. Replay evaluations against a real model cost money and minutes on every merge; many teams run a small smoke set per merge and the full set nightly and before production promotion. Manifest discipline adds friction to quick prompt tweaks, and that friction is the point: a prompt edit is a behaviour change and deserves the same path as a code change.