An ADK Java agent can pass every evaluation in CI and still fail the moment it takes real traffic. The model endpoint in production has a different quota. A secret was rotated. The new revision starts slowly because the image compiles at boot, or a tool call that worked against a staging backend returns 403 against production. None of this is visible until the release meets real users, so the deploy stage, the part between "the image is good" and "users are on it", decides whether those failures reach 5 percent of conversations for ten minutes or all of them for an afternoon.
This page is about that stage alone. Building, evaluating and promoting images are covered in the companion pages linked at the end. Here we treat the deploy as a state machine on Cloud Run, the target the ADK documentation uses for Java agents. It covers what goes in, how the first deploy differs from every later one, how to smoke-test a revision before it takes traffic, how a ramp decides whether to continue, and how rollback works, including what happens to live sessions.
The deploy as a state machine
A deploy has a small number of states, and each transition should be one command whose result you can check. Before the deploy starts, the image is verified: built once, evaluated, and addressed by digest. It is then deployed as a dark revision, which exists and has its own URL but takes no user traffic. A smoked revision has completed a real agent turn through that URL. In canary, it takes a small share of traffic while the controller compares it with the current revision. A ramp raises that share in steps. A promoted revision takes everything, and the previous one stays deployed so rollback is instant. Any failed check moves to rolled back.
Two pieces surround the machine. A deploy lock ensures only one rollout touches a service at a time, because two pipelines ramping different revisions produce traffic splits nobody intended. An audit record captures every transition with the digest, revision name, traffic split and verdict, so the question "what was serving at 14:02?" has an answer.
Inputs: one digest, pinned configuration
The deploy stage should accept exactly three inputs and nothing else: an image digest, an environment name and a commit SHA. Everything else comes from versioned files, never from the machine running the pipeline. For an ADK Java agent the revision-level configuration typically includes:
| Setting | Where it lives | Why it matters at deploy time |
|---|---|---|
| GOOGLE_CLOUD_PROJECT, GOOGLE_CLOUD_LOCATION | env/<env>.yaml | which project's model quota the agent spends |
| GOOGLE_GENAI_USE_ENTERPRISE=True | env/<env>.yaml | current ADK docs use it to route the GenAI client to Google Cloud rather than an API key (older guides used GOOGLE_GENAI_USE_VERTEXAI) |
| Model id, prompt or instruction version | env/<env>.yaml | a model alias change is a behaviour change; pin it per revision |
| Tool credentials | Secret Manager, --set-secrets | pin a secret version, not latest, so rollback restores the old credential |
| Runtime service account | deploy flag | least privilege; the agent's tools run with this identity |
| Concurrency, min instances, CPU | deploy flag | agent turns are long; low concurrency limits head-of-line blocking |
The pinning rule matters for rollback. A Cloud Run revision is immutable, so it keeps the image, environment and secret references it was created with. If those references are pinned, routing traffic back to the old revision restores the old behaviour exactly. If the old revision reads secret version latest, rollback brings back the old code with the new credential, which is a combination nobody tested.
Bootstrap, and the dark deploy
The first deploy of a service is different from all later ones, and pipelines that ignore this fail on day one. A dark deploy uses --no-traffic, which only makes sense when an existing revision can keep the traffic. When the service does not exist yet, there is nothing to hold it, so the first revision ends up serving whatever arrives. Give the pipeline an explicit bootstrap path: detect that the service is new, deploy without the flag, and keep the service private (no unauthenticated access, no public routing) until the smoke test passes.
The second trap comes from the gcloud reference itself. Deploying with --no-traffic moves any traffic assigned to LATEST onto the specific revision LATEST pointed to, and from then on the latest revision no longer receives traffic automatically. That is the behaviour you want for a controlled rollout, but it changes the service's routing model. A teammate who later runs a plain gcloud run deploy by hand will deploy a revision that takes no traffic, and may conclude the deploy is broken. Write it down, or end each promotion with an explicit update-traffic --to-latest if your team prefers the default model.
#!/usr/bin/env bash
# deploy.sh ENV DIGEST GIT_SHA -- dark-deploy one verified image
set -euo pipefail
ENV=$1 DIGEST=$2 SHA=${3:0:7}
PROJECT=${PROJECT:?set PROJECT}
SVC=support-agent REGION=us-central1
IMAGE="$REGION-docker.pkg.dev/$PROJECT/agents/support-agent@$DIGEST"
TAG="cand-$SHA"
FLAGS=(--image "$IMAGE" --region "$REGION" --revision-suffix "$SHA" --tag "$TAG"
--service-account "agent-runtime@$PROJECT.iam.gserviceaccount.com"
--no-allow-unauthenticated --env-vars-file "env/$ENV.yaml"
--set-secrets "SEARCH_API_KEY=search-api-key:7"
--concurrency 8 --min-instances 1)
if gcloud run services describe "$SVC" --region "$REGION" >/dev/null 2>&1; then
# remember what serves now, so rollback has a concrete target
PREV=$(gcloud run services describe "$SVC" --region "$REGION" --format=json \
| jq -r '.status.traffic[] | select(.percent==100) | .revisionName')
echo "$PREV" > prev_revision.txt
FLAGS+=(--no-traffic)
else
echo "bootstrap: first revision takes all traffic; service stays private until smoked"
fi
gcloud run deploy "$SVC" "${FLAGS[@]}"
gcloud run services describe "$SVC" --region "$REGION" --format=json \
| jq -r --arg t "$TAG" '.status.traffic[] | select(.tag==$t) | .url' > candidate_url.txt
Smoke-testing a real turn before traffic
A health check that returns 200 proves the JVM started, not that the agent works. The smoke test should complete one real turn through the candidate's tag URL: list the apps, create a session, send a message, and check the final answer. The endpoints below are the ones the ADK Java web server exposes in the Cloud Run guide (/list-apps, the sessions path and /run_sse). If you wrap the Runner in your own controller, point the same test at your routes.
public final class SmokeTest {
private static final HttpClient HTTP = HttpClient.newBuilder()
.connectTimeout(Duration.ofSeconds(10)).build();
public static void main(String[] args) throws Exception {
String base = Files.readString(Path.of("candidate_url.txt")).strip();
String token = System.getenv("ID_TOKEN"); // identity token for the deployer
String app = "support_agent", user = "smoke", session = "smoke-" + System.currentTimeMillis();
expect(call(base + "/list-apps", token, null), b -> b.contains(app), "app not loaded");
expect(call(base + "/apps/" + app + "/users/" + user + "/sessions/" + session, token, "{}"),
b -> true, "session create failed");
String turn = """
{"app_name":"%s","user_id":"%s","session_id":"%s",
"new_message":{"role":"user","parts":[{"text":"What is your refund window?"}]},
"streaming":false}""".formatted(app, user, session);
expect(call(base + "/run_sse", token, turn),
b -> b.contains("30 days"), "final answer missing expected fact");
}
static HttpResponse<String> call(String url, String token, String json) throws Exception {
var req = HttpRequest.newBuilder(URI.create(url)).timeout(Duration.ofSeconds(60))
.header("Authorization", "Bearer " + token).header("Content-Type", "application/json");
req = json == null ? req.GET() : req.POST(HttpRequest.BodyPublishers.ofString(json));
return HTTP.send(req.build(), HttpResponse.BodyHandlers.ofString());
}
static void expect(HttpResponse<String> r, Predicate<String> ok, String why) {
if (r.statusCode() / 100 != 2 || !ok.test(r.body()))
throw new IllegalStateException(why + ": HTTP " + r.statusCode());
}
}Pick a smoke question that touches a tool and a fact that only the production backend knows, so a wrong secret or a missing IAM grant fails here and not in the canary. Keep the check to a stable substring; the full quality judgment belongs to the evaluation stage, which already ran.
Ramps that decide with agent signals
Once smoked, the candidate takes 5 percent through gcloud run services update-traffic --to-tags cand-SHA=5. The controller then waits a soak period and compares the candidate with the baseline revision over the same window. HTTP-level signals are necessary but not sufficient, because agents fail politely: a turn can return 200 with an empty final response, or the model can stop calling a tool it used to call. So the server should log, per turn, the revision (Cloud Run sets K_REVISION in the container environment), whether a final response was produced, and tool errors. The decision then fits in a few lines:
record Window(long requests, long http5xx, double p95Ms, long turns, long emptyFinals, long toolErrors) {}
enum Verdict { PROMOTE, HOLD, ROLLBACK }
static double rate(long n, long d) { return d == 0 ? 0 : (double) n / d; }
static Verdict decide(Window cand, Window base) {
if (cand.requests() < 200) return Verdict.HOLD; // not enough evidence yet
if (rate(cand.http5xx(), cand.requests())
> Math.max(0.01, 2 * rate(base.http5xx(), base.requests()))) return Verdict.ROLLBACK;
if (cand.p95Ms() > 1.3 * base.p95Ms()) return Verdict.ROLLBACK;
if (rate(cand.emptyFinals(), cand.turns())
> rate(base.emptyFinals(), base.turns()) + 0.02) return Verdict.ROLLBACK;
if (rate(cand.toolErrors(), cand.turns())
> rate(base.toolErrors(), base.turns()) + 0.02) return Verdict.ROLLBACK;
return Verdict.PROMOTE; // advance one ramp step
}HOLD matters as much as the other two verdicts. At 5 percent of a quiet service, ten minutes may carry only a few dozen turns, and a verdict from that sample is a coin toss. Either extend the soak until the minimum sample arrives or send synthetic traffic to the tag URL. The thresholds above are starting points: derive yours from the variance you see between two healthy revisions.
Worked example: a ramp that rolls back
Suppose revision support-agent-a41c9e2 serves 100 percent and the candidate is support-agent-7f3b210, which moves the instruction to a new prompt version and upgrades the ADK dependency. The dark deploy creates the revision and tag cand-7f3b210, and the smoke test passes in 9 seconds. At 5 percent over 15 minutes, the candidate serves 412 requests with one 5xx (0.24 percent against a baseline of 0.20), a p95 of 6.1 s against 5.8 s, and 2.1 percent empty finals against 1.9. The verdict is PROMOTE, so it moves to 25 percent.
At 25 percent the candidate serves 2,050 turns with a tool error rate of 4.6 percent against a baseline of 1.2. The new prompt makes the model pass dates as free text, and the order-lookup tool rejects them. The verdict is ROLLBACK. The controller runs update-traffic --to-revisions support-agent-a41c9e2=100, and the split is back within seconds because the old revision never stopped running. Total exposure was about 2,500 turns, roughly 0.3 percent of a day, and the audit record links the failure to a specific digest and prompt version. The fix goes into the evaluation set as a new case, so the next build fails before it reaches deploy.
Rollback, sessions and side effects
Rollback on Cloud Run only changes routing, which is why it is fast. Three things make it less simple than it sounds for agents.
- Session state. With
InMemorySessionService, conversations live in instance memory, so any traffic shift strands every conversation mid-flight on instances that stop receiving traffic. Production agents should keep sessions in a persistent session service, so either revision can continue a conversation. Then the new revision must not write state the old one cannot read. Treat state keys like a database schema: add before you use, and remove only after a full release cycle. - External side effects. Rolling back code does not undo tickets created or emails sent by the bad revision. Log tool calls with the revision name so you can find and repair them.
- Retention. Keep at least the previous promoted revision deployed (it scales to zero if you allow it) and its image in Artifact Registry. Cleanup policies that delete untagged images can remove the image your rollback target refers to.
Failure modes
- Slow cold start eats the canary. The Dockerfile in the ADK quick start runs
mvn compile exec:javaat container start, which is fine for a first deploy and slow for every scale-out. Ship a prebuilt jar, and set min instances so the canary is not measured on cold starts. - Two rollouts at once. Without a lock, a hotfix pipeline and a feature pipeline overwrite each other's splits. Take the lock before the dark deploy and release it after promote or rollback.
- Smoke test against the wrong URL. Testing the service URL instead of the tag URL tests the old revision. Read the URL from the traffic block for your tag, as deploy.sh does.
- Model drift without a deploy. An unpinned model alias changes behaviour with no new revision, so the pipeline has nothing to roll back. Pin model versions in the env file.
- Redeploying the same commit. Revision names must be unique, so a retry or a config-only change with the same --revision-suffix fails. Add a deploy counter or timestamp to the suffix.
- Tag sprawl. Every candidate tag gets its own URL. Remove old tags with
update-traffic --remove-tagsso stale revisions are not reachable indefinitely.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| gcloud script plus a small controller | transparent, easy to debug, no new service | you own locking, audit and retries |
| Google Cloud Deploy with a canary strategy | managed rollouts, approvals and history | another system to learn; agent-level verdicts still need your metrics |
| Fast ramp (5 to 100 in one step) | short exposure to mixed revisions | a defect that needs volume to show reaches everyone |
| Slow ramp with long soaks | catches low-rate failures | longer mixed-revision period, slower hotfixes |
What to do next
- Write env/staging.yaml and env/prod.yaml, and pin the model version and every secret version in them.
- Add the bootstrap branch and the PREV capture to your deploy script, then test both paths on a scratch service.
- Turn the smoke test into a pipeline step whose question needs a tool and a production-only fact.
- Log K_REVISION, empty-final and tool-error flags for every turn, and build the Window query over them.
- Run a rehearsed rollback in staging and time it, including session continuity.
- Keep learning: the full CI/CD pipeline and release manifest, a production server around Runner on Cloud Run, evaluation in CI, canary releases for agents and deploying on Kubernetes.