Cloud Run runs a container behind an HTTPS endpoint, scales it with traffic, and can scale it to zero when nobody calls. For an agent built with the Agent Development Kit for Java that is an attractive home: no cluster to run, per-request billing, IAM on the endpoint, and a short path from a Maven project to a URL. It also has sharp edges that matter more for agents than for ordinary web services, because agent requests are long, stateful across turns, and dominated by model latency.

This article walks through both deployment routes: the one ADK's documentation describes, which gets you to a working URL fastest, and a production route with a small server you own. It covers the build, the deploy command, identity, where sessions must live, concurrency and timeouts, shutdown, a worked example, failure modes and a checklist. The general packaging and JVM sizing advice is in the ADK Java deployment guide; if you are choosing between Cloud Run and a cluster, compare this page with running ADK Java on Kubernetes.

What Cloud Run gives an agent, and what it does not

Cloud Run has a short contract with your container. It must listen on 0.0.0.0 on the port in the PORT environment variable, 8080 by default. The filesystem is in memory, so anything written to it consumes instance memory and disappears when the instance stops. Instances are started and stopped by the platform; when one is stopped it receives SIGTERM and has 10 seconds before SIGKILL. Requests are routed to any healthy instance, and with the gcloud CLI each instance accepts up to 80 concurrent requests per vCPU by default, configurable up to 1,000.

Map that onto an agent and three consequences follow. First, a conversation's session cannot live in instance memory, because turn two may land on a different instance than turn one, and either may be stopped between turns. Second, an instance can serve many conversations at once, because nearly all of an agent request is spent waiting for the model; CPU is rarely the limit, memory and model quota usually are. Third, long agent runs collide with the request timeout and the shutdown window, so you must decide what happens to a run that is cut off.

The documented route

ADK's documentation deploys Java agents with gcloud run deploy and a Dockerfile. The agent class must expose the root agent as a public static final field of type BaseAgent, which is how ADK's bundled web server discovers it. The project declares the google-adk and google-adk-dev artifacts (the documentation pins version 1.11.0 at the time of writing; use the version you tested) and configures exec-maven-plugin with com.google.adk.web.AdkWebServer as its main class.

package agents.capitalagent;

import com.google.adk.agents.BaseAgent;
import com.google.adk.agents.LlmAgent;
import com.google.adk.tools.Annotations.Schema;
import com.google.adk.tools.FunctionTool;
import java.util.Map;

public class CapitalAgent {
  public static final BaseAgent ROOT_AGENT = LlmAgent.builder()
      .name("capital_agent")
      .model("gemini-flash-latest")
      .description("Answers questions about capital cities.")
      .instruction("Use the getCapitalCity tool, then answer in one sentence.")
      .tools(FunctionTool.create(CapitalAgent.class, "getCapitalCity"))
      .build();

  public static Map<String, Object> getCapitalCity(
      @Schema(name = "country", description = "Country name") String country) {
    Map<String, String> capitals = Map.of("canada", "Ottawa", "japan", "Tokyo");
    return Map.of("result", capitals.getOrDefault(country.toLowerCase(), "unknown"));
  }
}

The documented Dockerfile copies the POM and sources into a Maven image and, at container start, runs mvn compile exec:java with the web server as the main class, passing --server.port from PORT and --adk.agents.source-dir=target. The deploy command builds the image from source with Cloud Build and creates the service:

export GOOGLE_CLOUD_PROJECT=my-project
export GOOGLE_CLOUD_LOCATION=us-central1
export GOOGLE_GENAI_USE_ENTERPRISE=True

gcloud run deploy capital-agent-service \
  --source . \
  --region $GOOGLE_CLOUD_LOCATION \
  --project $GOOGLE_CLOUD_PROJECT \
  --no-allow-unauthenticated \
  --set-env-vars="GOOGLE_CLOUD_PROJECT=$GOOGLE_CLOUD_PROJECT,GOOGLE_CLOUD_LOCATION=$GOOGLE_CLOUD_LOCATION,GOOGLE_GENAI_USE_ENTERPRISE=$GOOGLE_GENAI_USE_ENTERPRISE"

The documentation's example uses --allow-unauthenticated so you can test from a browser; the version above keeps the endpoint private, which is what you want for anything beyond a throwaway demo. ADK renamed the variable that routes model calls through Google Cloud: GOOGLE_GENAI_USE_ENTERPRISE was previously GOOGLE_GENAI_USE_VERTEXAI, and the documentation notes that older ADK versions only understand the old name. If your agent cannot reach the model after deployment, check this first.

The bundled server exposes the same API as local development: /list-apps, session creation at /apps/{app}/users/{user}/sessions/{session}, and /run_sse to run a turn. That makes it ideal for a demo or an internal tool. It is a poor production shape for two reasons. Running Maven at container start downloads and compiles on every cold start, which can take far longer than the request a user is waiting on. And the development server is built for development: it exposes every agent and its dev endpoints rather than the one narrow API your clients need.

A production server around Runner

The production route is a small server you own, wrapped around ADK's Runner. You choose the session service, expose one endpoint, and control shutdown. The example below uses the JDK's built-in HTTP server with a virtual-thread executor (Java 21), so each request blocks cheaply on the model call; a Spring Boot or Micronaut controller works the same way. The request carries the user and session IDs in headers and the message as the body, to keep the code short.

package com.example;

import com.google.adk.events.Event;
import com.google.adk.runner.Runner;
import com.google.adk.sessions.BaseSessionService;
import com.google.adk.sessions.InMemorySessionService;
import com.google.adk.sessions.Session;
import com.google.adk.sessions.VertexAiSessionService;
import com.google.genai.types.Content;
import com.google.genai.types.Part;
import com.sun.net.httpserver.HttpExchange;
import com.sun.net.httpserver.HttpServer;
import java.net.InetSocketAddress;
import java.nio.charset.StandardCharsets;
import java.util.Map;
import java.util.Optional;
import java.util.concurrent.Executors;

public final class AgentServer {
  static final String APP = System.getenv().getOrDefault("ADK_APP_NAME", "capital_agent");

  public static void main(String[] args) throws Exception {
    BaseSessionService sessions = "vertex".equals(System.getenv("SESSION_BACKEND"))
        ? new VertexAiSessionService() : new InMemorySessionService();
    Runner runner = Runner.builder()
        .agent(agents.capitalagent.CapitalAgent.ROOT_AGENT)
        .appName(APP).sessionService(sessions).build();

    int port = Integer.parseInt(System.getenv().getOrDefault("PORT", "8080"));
    HttpServer server = HttpServer.create(new InetSocketAddress("0.0.0.0", port), 0);
    server.setExecutor(Executors.newVirtualThreadPerTaskExecutor());
    server.createContext("/healthz", ex -> reply(ex, 200, "ok"));
    server.createContext("/chat", ex -> {
      String user = ex.getRequestHeaders().getFirst("X-User-Id");
      String sid = ex.getRequestHeaders().getFirst("X-Session-Id");   // absent on turn one
      if (user == null) { reply(ex, 400, "missing user"); return; }
      String msg = new String(ex.getRequestBody().readAllBytes(), StandardCharsets.UTF_8);
      try {
        Session s = (sid == null)
            ? sessions.createSession(APP, user, Map.of(), null).blockingGet()  // backend picks id
            : sessions.getSession(APP, user, sid, Optional.empty()).blockingGet();
        if (s == null) { reply(ex, 404, "unknown session"); return; }
        StringBuilder out = new StringBuilder();
        runner.runAsync(user, s.id(), Content.fromParts(Part.fromText(msg)))
            .blockingForEach(ev -> { if (ev.finalResponse()) out.append(ev.stringifyContent()); });
        ex.getResponseHeaders().set("X-Session-Id", s.id());           // caller sends it back
        reply(ex, 200, out.toString());
      } catch (RuntimeException e) {
        reply(ex, 502, "agent error");          // log e with the session id, never the prompt
      }
    });
    server.start();
    Runtime.getRuntime().addShutdownHook(new Thread(() -> server.stop(8)));
  }

  static void reply(HttpExchange ex, int code, String body) throws java.io.IOException {
    byte[] b = body.getBytes(StandardCharsets.UTF_8);
    ex.sendResponseHeaders(code, b.length);
    try (var os = ex.getResponseBody()) { os.write(b); }
  }
}

Two details matter. The shutdown hook gives in-flight requests up to 8 seconds, inside Cloud Run's 10-second window, and leaves time for the JVM to exit cleanly; graceful shutdown for ADK Java covers the draining logic in depth. And the in-memory session service is only a local default. VertexAiSessionService stores sessions in Google Cloud's managed session service, keyed to an Agent Engine (reasoning engine) resource: the app name must be that resource's numeric ID or full resource name, not a friendly name. Letting the backend generate session IDs, as the server does on turn one, sidesteps its stricter ID format rules.

Building an image that starts fast

Build the application at image-build time, not at start. A multi-stage Dockerfile compiles once with Maven, copies the runtime dependencies next to the application JAR, and ships only a JRE:

FROM maven:3.9-eclipse-temurin-21 AS build
WORKDIR /src
COPY pom.xml .
RUN mvn -B -q dependency:go-offline
COPY src ./src
RUN mvn -B -q package -DskipTests dependency:copy-dependencies -DincludeScope=runtime

FROM eclipse-temurin:21-jre
WORKDIR /app
COPY --from=build /src/target/*.jar /app/app.jar
COPY --from=build /src/target/dependency /app/lib
USER 1000
ENV JAVA_TOOL_OPTIONS="-XX:MaxRAMPercentage=75 -XX:+UseSerialGC"
CMD ["java", "-cp", "/app/app.jar:/app/lib/*", "com.example.AgentServer"]

MaxRAMPercentage lets the heap follow the memory limit you set on the service instead of a guess baked into the image. The serial collector is a reasonable default for one or two vCPUs; measure before switching. If startup is still slow, Cloud Run's startup CPU boost (the --cpu-boost flag) gives the instance extra CPU while the JVM loads classes, and a minimum instance count keeps one warm instance for latency-sensitive endpoints, at the cost of paying for it while idle.

Identity and the deploy command

Give the service its own runtime service account and grant it only what the agent needs: permission to call Vertex AI (the Vertex AI User role), access to the specific secrets it reads, and whatever the tools touch. Callers need the Cloud Run Invoker role on the service. Then deploy with explicit settings rather than defaults:

gcloud iam service-accounts create agent-runtime
gcloud projects add-iam-policy-binding $GOOGLE_CLOUD_PROJECT \
  --member="serviceAccount:agent-runtime@$GOOGLE_CLOUD_PROJECT.iam.gserviceaccount.com" \
  --role="roles/aiplatform.user"

gcloud run deploy capital-agent \
  --source . --region us-central1 \
  --service-account agent-runtime@$GOOGLE_CLOUD_PROJECT.iam.gserviceaccount.com \
  --no-allow-unauthenticated \
  --cpu 1 --memory 1Gi --concurrency 20 --timeout 300 \
  --min-instances 1 --max-instances 20 --cpu-boost \
  --set-env-vars="GOOGLE_CLOUD_PROJECT=$GOOGLE_CLOUD_PROJECT,GOOGLE_CLOUD_LOCATION=us-central1,GOOGLE_GENAI_USE_ENTERPRISE=True,SESSION_BACKEND=vertex,ADK_APP_NAME=$AGENT_ENGINE_ID"

URL=$(gcloud run services describe capital-agent --region us-central1 --format='value(status.url)')
curl -s -i -X POST "$URL/chat" \
  -H "Authorization: Bearer $(gcloud auth print-identity-token)" \
  -H "X-User-Id: u-42" \
  -d "What is the capital of Canada?"     # read X-Session-Id from the response headers

If you use a Gemini API key instead of Vertex AI, store it in Secret Manager and map it into the environment with --set-secrets, never with --set-env-vars, so it does not appear in the revision's configuration. The concurrency of 20 is deliberately below the default: each in-flight agent turn holds conversation history, tool results and response buffers in memory, and the model quota is shared by every instance. Start low, load test, and raise it while watching memory and model error rates. Set --max-instances from your model quota, not from your hopes: twenty instances at concurrency twenty can open four hundred simultaneous model calls.

Worked example: one conversation, two instances

Request path for an ADK Java agent on Cloud RunCallerapp, backend, or curlCloud Run front endIAM check: run.invokerroutes to any instanceID tokenContainer instancelistens on 0.0.0.0:$PORTHTTP server (virtual threads)Runner: agent + session serviceLlmAgent and FunctionToolsSIGTERM: 10 s to finishGeminiVertex AI or API keySession storeVertexAiSessionServiceSecret Managerkeys as env varsmodel callsload / appendat startRuntime service accountroles: Vertex AI user, Secret Manager accessor, nothing else; instance memory is the only local disk
Any request can land on any instance, so session state must live outside the container; the instance holds only the agent definition and in-flight work.

Here is how one conversation flows through that deployment. A support backend calls the service with an ID token minted for its own service account. Cloud Run checks the Invoker role and routes the request to instance A. No session header is present, so the server creates a session in the store, runs the agent (model call, getCapitalCity, model call again), appends the events, and returns the answer with the new session ID in a header. Nearly all of the latency is model time.

Turn two arrives a minute later with that session ID. Traffic has grown, so Cloud Run started instance B and routes the request there. Instance B has never seen the session, but it loads it from the store, so the agent sees turn one's history and answers in context. Ten minutes later traffic drops, Cloud Run stops instance A, the shutdown hook lets one in-flight request finish, and nothing is lost because no conversation lived in A's memory. Swap the session backend to in-memory and replay the same sequence: turn two on instance B finds no such session, and a client that silently starts over produces exactly the agent amnesia users report.

Failure modes

SymptomCauseFix
Agent forgets earlier turns, intermittentlyIn-memory sessions across several instancesExternal session service; test turn two on a fresh instance
First request after idle takes very longMaven compiling at container start, or JVM cold startBuild at image time; --cpu-boost; a minimum instance
Model calls fail after deploy, work locallyVariable name not understood by your ADK version, or missing Vertex AI roleTry GOOGLE_GENAI_USE_VERTEXAI on older versions; check the runtime account roles
429 or resource-exhausted errors under loadInstances times concurrency exceeds model quotaCap max instances; retry with backoff and jitter
504 on long research-style runsRequest timeout shorter than the agent runRaise --timeout or move long runs to an async job
Requests cut off during deploysWork longer than the 10-second SIGTERM windowBound turn length; drain in the shutdown hook
Unexpected bill or abuse--allow-unauthenticated on a dev serverPrivate service, Invoker role, a gateway for public traffic

Trade-offs

Cloud Run wins when traffic is uneven, the team does not want to run a cluster, and conversations fit comfortably inside a request. It loses some of its appeal when every request needs minutes of work, when you need long-lived streaming connections for many users at once, or when you must co-locate with other services in a cluster for networking reasons; Kubernetes gives you more control there at the cost of operating it. Vertex AI Agent Engine is the other alternative: a managed runtime built specifically for agents, with sessions included, in exchange for less control over the server and its API.

Inside Cloud Run, the main trade-offs are cost versus latency (minimum instances and instance-based billing buy warm starts), throughput versus quota (higher concurrency and more instances only help until the model quota is the bottleneck), and simplicity versus control (the bundled ADK server versus your own). Virtual threads make the own-server route cheap to write; virtual threads for Java agents explains why blocking on the model is fine.

What to do next

  1. Expose the root agent as a public static final BaseAgent and get the documented gcloud route working once, with the endpoint private.
  2. Write a small server around Runner with one endpoint, health checks and a shutdown hook that finishes within 10 seconds.
  3. Move the build into a multi-stage Dockerfile so containers start the JVM, not Maven.
  4. Create a dedicated runtime service account with the Vertex AI User role and only the secrets it needs; give callers the Invoker role.
  5. Switch sessions to an external store and prove turn two works on a different instance.
  6. Set concurrency, timeout, max instances and memory explicitly, then load test against your model quota.
  7. Alert on 5xx rate, model 429s, p95 latency and instance count, and log session IDs (not prompts) for every failed turn.
Key takeaway: Cloud Run is a good home for ADK Java agents if you respect its contract: listen on PORT, keep no conversation state in the instance, finish within 10 seconds of SIGTERM, and size concurrency and instance counts against model quota. Use the documented AdkWebServer route to get started, then ship a small server around Runner with an external session service, a prebuilt image and a private, least-privilege service.