An ADK Java agent can pass every evaluation in CI and still fail the moment it takes real traffic. The model endpoint in production has a different quota. A secret was rotated. The new revision starts slowly because the image compiles at boot, or a tool call that worked against a staging backend returns 403 against production. None of this is visible until the release meets real users, so the deploy stage, the part between "the image is good" and "users are on it", decides whether those failures reach 5 percent of conversations for ten minutes or all of them for an afternoon.

This page is about that stage alone. Building, evaluating and promoting images are covered in the companion pages linked at the end. Here we treat the deploy as a state machine on Cloud Run, the target the ADK documentation uses for Java agents. It covers what goes in, how the first deploy differs from every later one, how to smoke-test a revision before it takes traffic, how a ramp decides whether to continue, and how rollback works, including what happens to live sessions.

The deploy as a state machine

The deploy stage as a state machine: every arrow is a command, every state is observableVerified imagedigest + env fileDark revision--no-traffic --tagSmokedreal turn via tag URLCanary 5%soak, compare to baseRamp 25% / 50%same verdict each stepPromoted 100%previous kept warmRolled back--to-revisions PREV=100smoke failsverdict ROLLBACKDeploy lock + audit recordone rollout per service, every step loggedSignals keyed by K_REVISION5xx, p95, empty finals, tool errorsA deploy that cannot say which state it is in, and how it got there, cannot be rolled back with confidence.
Figure: the deploy controller moves a verified image through dark, smoked, canary and promoted states; any failed check jumps to rollback.

A deploy has a small number of states, and each transition should be one command whose result you can check. Before the deploy starts, the image is verified: built once, evaluated, and addressed by digest. It is then deployed as a dark revision, which exists and has its own URL but takes no user traffic. A smoked revision has completed a real agent turn through that URL. In canary, it takes a small share of traffic while the controller compares it with the current revision. A ramp raises that share in steps. A promoted revision takes everything, and the previous one stays deployed so rollback is instant. Any failed check moves to rolled back.

Two pieces surround the machine. A deploy lock ensures only one rollout touches a service at a time, because two pipelines ramping different revisions produce traffic splits nobody intended. An audit record captures every transition with the digest, revision name, traffic split and verdict, so the question "what was serving at 14:02?" has an answer.

Inputs: one digest, pinned configuration

The deploy stage should accept exactly three inputs and nothing else: an image digest, an environment name and a commit SHA. Everything else comes from versioned files, never from the machine running the pipeline. For an ADK Java agent the revision-level configuration typically includes:

SettingWhere it livesWhy it matters at deploy time
GOOGLE_CLOUD_PROJECT, GOOGLE_CLOUD_LOCATIONenv/<env>.yamlwhich project's model quota the agent spends
GOOGLE_GENAI_USE_ENTERPRISE=Trueenv/<env>.yamlcurrent ADK docs use it to route the GenAI client to Google Cloud rather than an API key (older guides used GOOGLE_GENAI_USE_VERTEXAI)
Model id, prompt or instruction versionenv/<env>.yamla model alias change is a behaviour change; pin it per revision
Tool credentialsSecret Manager, --set-secretspin a secret version, not latest, so rollback restores the old credential
Runtime service accountdeploy flagleast privilege; the agent's tools run with this identity
Concurrency, min instances, CPUdeploy flagagent turns are long; low concurrency limits head-of-line blocking

The pinning rule matters for rollback. A Cloud Run revision is immutable, so it keeps the image, environment and secret references it was created with. If those references are pinned, routing traffic back to the old revision restores the old behaviour exactly. If the old revision reads secret version latest, rollback brings back the old code with the new credential, which is a combination nobody tested.

Bootstrap, and the dark deploy

The first deploy of a service is different from all later ones, and pipelines that ignore this fail on day one. A dark deploy uses --no-traffic, which only makes sense when an existing revision can keep the traffic. When the service does not exist yet, there is nothing to hold it, so the first revision ends up serving whatever arrives. Give the pipeline an explicit bootstrap path: detect that the service is new, deploy without the flag, and keep the service private (no unauthenticated access, no public routing) until the smoke test passes.

The second trap comes from the gcloud reference itself. Deploying with --no-traffic moves any traffic assigned to LATEST onto the specific revision LATEST pointed to, and from then on the latest revision no longer receives traffic automatically. That is the behaviour you want for a controlled rollout, but it changes the service's routing model. A teammate who later runs a plain gcloud run deploy by hand will deploy a revision that takes no traffic, and may conclude the deploy is broken. Write it down, or end each promotion with an explicit update-traffic --to-latest if your team prefers the default model.

#!/usr/bin/env bash
# deploy.sh ENV DIGEST GIT_SHA  -- dark-deploy one verified image
set -euo pipefail
ENV=$1 DIGEST=$2 SHA=${3:0:7}
PROJECT=${PROJECT:?set PROJECT}
SVC=support-agent REGION=us-central1
IMAGE="$REGION-docker.pkg.dev/$PROJECT/agents/support-agent@$DIGEST"
TAG="cand-$SHA"

FLAGS=(--image "$IMAGE" --region "$REGION" --revision-suffix "$SHA" --tag "$TAG"
       --service-account "agent-runtime@$PROJECT.iam.gserviceaccount.com"
       --no-allow-unauthenticated --env-vars-file "env/$ENV.yaml"
       --set-secrets "SEARCH_API_KEY=search-api-key:7"
       --concurrency 8 --min-instances 1)

if gcloud run services describe "$SVC" --region "$REGION" >/dev/null 2>&1; then
  # remember what serves now, so rollback has a concrete target
  PREV=$(gcloud run services describe "$SVC" --region "$REGION" --format=json \
         | jq -r '.status.traffic[] | select(.percent==100) | .revisionName')
  echo "$PREV" > prev_revision.txt
  FLAGS+=(--no-traffic)
else
  echo "bootstrap: first revision takes all traffic; service stays private until smoked"
fi

gcloud run deploy "$SVC" "${FLAGS[@]}"
gcloud run services describe "$SVC" --region "$REGION" --format=json \
  | jq -r --arg t "$TAG" '.status.traffic[] | select(.tag==$t) | .url' > candidate_url.txt

Smoke-testing a real turn before traffic

A health check that returns 200 proves the JVM started, not that the agent works. The smoke test should complete one real turn through the candidate's tag URL: list the apps, create a session, send a message, and check the final answer. The endpoints below are the ones the ADK Java web server exposes in the Cloud Run guide (/list-apps, the sessions path and /run_sse). If you wrap the Runner in your own controller, point the same test at your routes.

public final class SmokeTest {
  private static final HttpClient HTTP = HttpClient.newBuilder()
      .connectTimeout(Duration.ofSeconds(10)).build();

  public static void main(String[] args) throws Exception {
    String base = Files.readString(Path.of("candidate_url.txt")).strip();
    String token = System.getenv("ID_TOKEN");          // identity token for the deployer
    String app = "support_agent", user = "smoke", session = "smoke-" + System.currentTimeMillis();

    expect(call(base + "/list-apps", token, null), b -> b.contains(app), "app not loaded");
    expect(call(base + "/apps/" + app + "/users/" + user + "/sessions/" + session, token, "{}"),
        b -> true, "session create failed");
    String turn = """
        {"app_name":"%s","user_id":"%s","session_id":"%s",
         "new_message":{"role":"user","parts":[{"text":"What is your refund window?"}]},
         "streaming":false}""".formatted(app, user, session);
    expect(call(base + "/run_sse", token, turn),
        b -> b.contains("30 days"), "final answer missing expected fact");
  }

  static HttpResponse<String> call(String url, String token, String json) throws Exception {
    var req = HttpRequest.newBuilder(URI.create(url)).timeout(Duration.ofSeconds(60))
        .header("Authorization", "Bearer " + token).header("Content-Type", "application/json");
    req = json == null ? req.GET() : req.POST(HttpRequest.BodyPublishers.ofString(json));
    return HTTP.send(req.build(), HttpResponse.BodyHandlers.ofString());
  }

  static void expect(HttpResponse<String> r, Predicate<String> ok, String why) {
    if (r.statusCode() / 100 != 2 || !ok.test(r.body()))
      throw new IllegalStateException(why + ": HTTP " + r.statusCode());
  }
}

Pick a smoke question that touches a tool and a fact that only the production backend knows, so a wrong secret or a missing IAM grant fails here and not in the canary. Keep the check to a stable substring; the full quality judgment belongs to the evaluation stage, which already ran.

Ramps that decide with agent signals

Once smoked, the candidate takes 5 percent through gcloud run services update-traffic --to-tags cand-SHA=5. The controller then waits a soak period and compares the candidate with the baseline revision over the same window. HTTP-level signals are necessary but not sufficient, because agents fail politely: a turn can return 200 with an empty final response, or the model can stop calling a tool it used to call. So the server should log, per turn, the revision (Cloud Run sets K_REVISION in the container environment), whether a final response was produced, and tool errors. The decision then fits in a few lines:

record Window(long requests, long http5xx, double p95Ms, long turns, long emptyFinals, long toolErrors) {}
enum Verdict { PROMOTE, HOLD, ROLLBACK }

static double rate(long n, long d) { return d == 0 ? 0 : (double) n / d; }

static Verdict decide(Window cand, Window base) {
  if (cand.requests() < 200) return Verdict.HOLD;                  // not enough evidence yet
  if (rate(cand.http5xx(), cand.requests())
      > Math.max(0.01, 2 * rate(base.http5xx(), base.requests()))) return Verdict.ROLLBACK;
  if (cand.p95Ms() > 1.3 * base.p95Ms()) return Verdict.ROLLBACK;
  if (rate(cand.emptyFinals(), cand.turns())
      > rate(base.emptyFinals(), base.turns()) + 0.02) return Verdict.ROLLBACK;
  if (rate(cand.toolErrors(), cand.turns())
      > rate(base.toolErrors(), base.turns()) + 0.02) return Verdict.ROLLBACK;
  return Verdict.PROMOTE;                                          // advance one ramp step
}

HOLD matters as much as the other two verdicts. At 5 percent of a quiet service, ten minutes may carry only a few dozen turns, and a verdict from that sample is a coin toss. Either extend the soak until the minimum sample arrives or send synthetic traffic to the tag URL. The thresholds above are starting points: derive yours from the variance you see between two healthy revisions.

Worked example: a ramp that rolls back

Suppose revision support-agent-a41c9e2 serves 100 percent and the candidate is support-agent-7f3b210, which moves the instruction to a new prompt version and upgrades the ADK dependency. The dark deploy creates the revision and tag cand-7f3b210, and the smoke test passes in 9 seconds. At 5 percent over 15 minutes, the candidate serves 412 requests with one 5xx (0.24 percent against a baseline of 0.20), a p95 of 6.1 s against 5.8 s, and 2.1 percent empty finals against 1.9. The verdict is PROMOTE, so it moves to 25 percent.

At 25 percent the candidate serves 2,050 turns with a tool error rate of 4.6 percent against a baseline of 1.2. The new prompt makes the model pass dates as free text, and the order-lookup tool rejects them. The verdict is ROLLBACK. The controller runs update-traffic --to-revisions support-agent-a41c9e2=100, and the split is back within seconds because the old revision never stopped running. Total exposure was about 2,500 turns, roughly 0.3 percent of a day, and the audit record links the failure to a specific digest and prompt version. The fix goes into the evaluation set as a new case, so the next build fails before it reaches deploy.

Rollback, sessions and side effects

Rollback on Cloud Run only changes routing, which is why it is fast. Three things make it less simple than it sounds for agents.

  • Session state. With InMemorySessionService, conversations live in instance memory, so any traffic shift strands every conversation mid-flight on instances that stop receiving traffic. Production agents should keep sessions in a persistent session service, so either revision can continue a conversation. Then the new revision must not write state the old one cannot read. Treat state keys like a database schema: add before you use, and remove only after a full release cycle.
  • External side effects. Rolling back code does not undo tickets created or emails sent by the bad revision. Log tool calls with the revision name so you can find and repair them.
  • Retention. Keep at least the previous promoted revision deployed (it scales to zero if you allow it) and its image in Artifact Registry. Cleanup policies that delete untagged images can remove the image your rollback target refers to.

Failure modes

  • Slow cold start eats the canary. The Dockerfile in the ADK quick start runs mvn compile exec:java at container start, which is fine for a first deploy and slow for every scale-out. Ship a prebuilt jar, and set min instances so the canary is not measured on cold starts.
  • Two rollouts at once. Without a lock, a hotfix pipeline and a feature pipeline overwrite each other's splits. Take the lock before the dark deploy and release it after promote or rollback.
  • Smoke test against the wrong URL. Testing the service URL instead of the tag URL tests the old revision. Read the URL from the traffic block for your tag, as deploy.sh does.
  • Model drift without a deploy. An unpinned model alias changes behaviour with no new revision, so the pipeline has nothing to roll back. Pin model versions in the env file.
  • Redeploying the same commit. Revision names must be unique, so a retry or a config-only change with the same --revision-suffix fails. Add a deploy counter or timestamp to the suffix.
  • Tag sprawl. Every candidate tag gets its own URL. Remove old tags with update-traffic --remove-tags so stale revisions are not reachable indefinitely.

Trade-offs

ChoiceGainCost
gcloud script plus a small controllertransparent, easy to debug, no new serviceyou own locking, audit and retries
Google Cloud Deploy with a canary strategymanaged rollouts, approvals and historyanother system to learn; agent-level verdicts still need your metrics
Fast ramp (5 to 100 in one step)short exposure to mixed revisionsa defect that needs volume to show reaches everyone
Slow ramp with long soakscatches low-rate failureslonger mixed-revision period, slower hotfixes

What to do next

  1. Write env/staging.yaml and env/prod.yaml, and pin the model version and every secret version in them.
  2. Add the bootstrap branch and the PREV capture to your deploy script, then test both paths on a scratch service.
  3. Turn the smoke test into a pipeline step whose question needs a tool and a production-only fact.
  4. Log K_REVISION, empty-final and tool-error flags for every turn, and build the Window query over them.
  5. Run a rehearsed rollback in staging and time it, including session continuity.
  6. Keep learning: the full CI/CD pipeline and release manifest, a production server around Runner on Cloud Run, evaluation in CI, canary releases for agents and deploying on Kubernetes.
Key takeaway: Deploying an agent is a short state machine: a verified digest becomes a dark revision, completes a real turn through its tag URL, and ramps only while agent-level signals match the baseline. Pin the model, secrets and configuration so rollback restores real behaviour. Handle the first deploy and the --no-traffic routing change explicitly. Keep sessions persistent and backward-compatible so a traffic shift never strands a conversation.