Vertex AI Agent Engine is Google Cloud's managed runtime for agents. It is now documented as Agent Runtime, part of the Gemini Enterprise Agent Platform, and the underlying API resource is still called a reasoning engine. It gives you autoscaled serving, IAM-authenticated endpoints, a console playground and managed sessions without running a server fleet yourself. The catch for Java teams is that the one-command ADK deploy path is documented for ADK Python and Go only. There is no official command that takes an ADK Java JAR and deploys it.

There is still a supported route: Agent Runtime accepts a custom container in any language, provided the container honours a small runtime contract. This article builds that route end to end for ADK Java. You will write a Java server that speaks the contract, size the JVM for the runtime's resource limits, deploy the image, call it, wire up managed sessions, and learn the failure modes, so you can decide whether Agent Runtime or Cloud Run is the better home for your agent.

What is supported for Java

Agent Runtime documents several deployment methods. Two of them, deploying from source files and from a Git repository, are Python only. The in-memory SDK path serialises a Python agent object. The two that work for Java are:

  • Dockerfile: you upload source packages including a Dockerfile, and the service builds the image for you.
  • Container image: you build and push an image to Artifact Registry yourself and pass its URI. This gives you full control over the build and the fastest deployments, and it is the method used below.

Either way, the container must follow the runtime contract. Anything you read elsewhere about uploading a JAR with a manifest, or about a gcloud command that creates Java agents directly, does not match the current documentation.

A Java agent on Agent Runtime: managed front door, your container behind itCallerIAM-authenticatedAgent Runtime APIreasoningEngines/ID:queryYour container (Java 21)HTTP on 0.0.0.0:8080/api/reasoning_engineRunnerLlmAgent + toolsVertexAiSessionServicesession clientclass_method, inputGemini on Vertex AImodel callsSessions resourceseparate Agent Engine IDgenerateeventsArtifact Registryimage_uripulled at deploy
Request path for a containerised ADK Java agent. Agent Runtime owns authentication, scaling and the public API; the container owns the agent loop; sessions live in a separately created resource.

The runtime contract

RequirementWhat it means for a Java server
Listen on 0.0.0.0, port 8080Bind all interfaces; read PORT if set and default to 8080
POST /api/reasoning_engineUnary call. Body {"class_method": "query", "input": {...}}; reply {"output": ...}
POST /api/stream_reasoning_engineStreaming call. Same body; reply is newline-delimited JSON, one chunk per line
class_methods declared at deployEach method has a name and an api_mode: empty or async routes to the unary endpoint, stream or async_stream to the streaming one
SDK and playground supportThe Python SDK expects query and stream_query; the console playground needs stream_query in stream mode

The two endpoints are optional in the strict sense: a container without them still runs, but you lose the SDK and the playground, which are most of the reason to use the managed runtime. The shape of input is yours to define. This article uses {"user_id", "session_id", "message"}, returns the session ID on every response, and treats a missing session ID as the start of a new conversation.

A Java server that honours the contract

The agent itself is an ordinary ADK Java LlmAgent; nothing about it changes for Agent Runtime. The server wraps a Runner exactly as a Cloud Run server would, but routes on the contract's paths and class_method values. Jackson is used for JSON; ADK Java already depends on it, but declare it explicitly so a dependency change cannot break the build.

// Imports as in the Cloud Run server, plus com.fasterxml.jackson.databind.{ObjectMapper, JsonNode}
// and java.io.{IOException, OutputStream}. CapitalAgent is the LlmAgent with one FunctionTool.
public final class AgentRuntimeServer {
  static final ObjectMapper JSON = new ObjectMapper();
  static final String APP = System.getenv("SESSION_ENGINE_ID");  // sessions resource ID
  static BaseSessionService sessions;
  static Runner runner;

  public static void main(String[] args) throws Exception {
    sessions = new VertexAiSessionService();
    runner = Runner.builder().agent(CapitalAgent.ROOT_AGENT)
        .appName(APP).sessionService(sessions).build();
    int port = Integer.parseInt(System.getenv().getOrDefault("PORT", "8080"));
    HttpServer server = HttpServer.create(new InetSocketAddress("0.0.0.0", port), 0);
    server.setExecutor(Executors.newVirtualThreadPerTaskExecutor());
    server.createContext("/api/reasoning_engine", ex -> handle(ex, false));
    server.createContext("/api/stream_reasoning_engine", ex -> handle(ex, true));
    server.start();
  }

  static void handle(HttpExchange ex, boolean stream) throws IOException {
    JsonNode req = JSON.readTree(ex.getRequestBody());
    JsonNode in = req.path("input");
    String expected = stream ? "stream_query" : "query";
    String user = in.path("user_id").asText("");
    String msg = in.path("message").asText("");
    if (!expected.equals(req.path("class_method").asText()) || user.isEmpty() || msg.isEmpty()) {
      send(ex, 400, Map.of("error", "expected " + expected + " with user_id and message"));
      return;
    }
    String sid = in.path("session_id").asText("");
    Session s = sid.isEmpty()
        ? sessions.createSession(APP, user, Map.of(), null).blockingGet()
        : sessions.getSession(APP, user, sid, Optional.empty()).blockingGet();
    if (s == null) { send(ex, 404, Map.of("error", "unknown session")); return; }
    Content content = Content.fromParts(Part.fromText(msg));

    if (!stream) {
      StringBuilder out = new StringBuilder();
      runner.runAsync(user, s.id(), content)
          .blockingForEach(ev -> { if (ev.finalResponse()) out.append(ev.stringifyContent()); });
      send(ex, 200, Map.of("output", Map.of("text", out.toString(), "session_id", s.id())));
      return;
    }
    ex.getResponseHeaders().set("Content-Type", "application/x-ndjson");
    ex.sendResponseHeaders(200, 0);                       // chunked: length unknown
    try (OutputStream os = ex.getResponseBody()) {
      try {
        runner.runAsync(user, s.id(), content).blockingForEach(ev -> line(os, Map.of(
            "text", ev.stringifyContent(), "final", ev.finalResponse(), "session_id", s.id())));
      } catch (RuntimeException e) {                      // headers already sent:
        line(os, Map.of("error", "agent failed", "session_id", s.id()));  // report in-band
      }
    }
  }

  static void line(OutputStream os, Object o) throws IOException {
    os.write((JSON.writeValueAsString(o) + "\n").getBytes(StandardCharsets.UTF_8));
    os.flush();
  }

  static void send(HttpExchange ex, int code, Object body) throws IOException {
    byte[] b = JSON.writeValueAsBytes(body);
    ex.getResponseHeaders().set("Content-Type", "application/json");
    ex.sendResponseHeaders(code, b.length);
    try (OutputStream os = ex.getResponseBody()) { os.write(b); }
  }
}

Two details deserve attention. A streaming response has already sent status 200 before the agent runs, so a failure halfway through cannot become an HTTP error; it must be an in-band error line that clients check for. And letting the session backend mint the session ID on turn one sidesteps its constraints on what an ID may look like, as the Cloud Run deployment guide explains.

The image and JVM sizing

The default resource limits are 4 CPUs and 4 GiB of memory per instance, with container concurrency 9, which the documentation recommends as 2 x CPU + 1. Memory can be configured from 1 GiB to 32 GiB and CPU to 1, 2, 4, 6 or 8. Size the JVM to the container, not the host: cap the heap at a percentage of the container limit and leave room for thread stacks, metaspace and direct buffers. Agent requests spend most of their time waiting on the model, which is why virtual threads and a modest heap go a long way.

FROM maven:3.9-eclipse-temurin-21 AS build
WORKDIR /src
COPY pom.xml .
RUN mvn -B -q dependency:go-offline
COPY src ./src
RUN mvn -B -q package -DskipTests dependency:copy-dependencies -DincludeScope=runtime

FROM eclipse-temurin:21-jre
WORKDIR /app
COPY --from=build /src/target/*.jar /app/app.jar
COPY --from=build /src/target/dependency /app/lib
USER 1000
ENV JAVA_TOOL_OPTIONS="-XX:MaxRAMPercentage=70 -XX:+ExitOnOutOfMemoryError"
EXPOSE 8080
CMD ["java", "-cp", "/app/app.jar:/app/lib/*", "com.example.AgentRuntimeServer"]

Build and push it to Artifact Registry with your usual pipeline, for example gcloud builds submit --tag REGION-docker.pkg.dev/PROJECT/agents/capital:1.0.0. Tag immutably; redeploying a moving tag makes rollbacks guesswork.

Deploying and calling the agent

Deployment uses the Vertex AI Python SDK even though the agent is Java; the SDK is only the control-plane client. Current documentation shows the call as client.runtimes.create; earlier releases and some codelabs show the same operation as client.agent_engines.create. Use the form your installed SDK documents, and take the config keys from that same page.

import vertexai

client = vertexai.Client(project="my-project", location="us-central1")

remote = client.runtimes.create(config={
    "display_name": "capital-agent-java",
    "container_spec": {
        "image_uri": "us-central1-docker.pkg.dev/my-project/agents/capital:1.0.0",
    },
    "class_methods": [
        {"name": "query", "api_mode": ""},               # unary endpoint
        {"name": "stream_query", "api_mode": "stream"},  # ndjson endpoint
    ],
    "min_instances": 1,
    "max_instances": 20,
    "resource_limits": {"cpu": "4", "memory": "4Gi"},
    "container_concurrency": 9,
    "service_account": "agent-runtime@my-project.iam.gserviceaccount.com",
    "env_vars": {"SESSION_ENGINE_ID": "1234567890", "GOOGLE_CLOUD_LOCATION": "us-central1"},
})

Grant the service account permission to call Vertex AI models and the sessions resource, and grant the platform's service agent read access to the Artifact Registry repository so it can pull the image. Callers then use the resource's :query method with an OAuth token; the streaming method is :streamQuery.

curl -s -X POST \
  -H "Authorization: Bearer $(gcloud auth print-access-token)" \
  -H "Content-Type: application/json" \
  "https://us-central1-aiplatform.googleapis.com/v1/projects/my-project/locations/us-central1/reasoningEngines/${ENGINE_ID}:query" \
  -d '{"class_method": "query", "input": {"user_id": "u-42", "message": "Capital of Canada?"}}'

Sessions

Python ADK apps deployed to Agent Runtime get managed sessions automatically. A custom container does not; it is your code that calls the session service. ADK Java's VertexAiSessionService stores sessions in an Agent Engine resource, so the Runner's app name has to identify that resource, by its number or its full projects/.../reasoningEngines path; a display name will not resolve. The documentation describes creating an instance with no agent attached purely to hold sessions and memory. Create one per region and environment ahead of time, and pass its ID to the container as configuration, as SESSION_ENGINE_ID does above. That separation also means redeploying or deleting the serving resource never touches conversation history. For longer-lived recall across conversations, see cross-session memory; for per-region session resources and home-region pinning, see multi-region deployment.

Worked example: two turns

Follow one two-turn conversation through the system.

  1. The client calls :query with user_id u-42 and no session ID. Agent Runtime authenticates the caller with IAM, picks a warm instance (min_instances is 1, so one exists) and forwards the body to /api/reasoning_engine.
  2. The server sees no session ID and asks the sessions resource to create one. The Runner appends the user message, calls Gemini, which requests the tool, runs getCapitalCity in-process, calls Gemini again and emits a final response event.
  3. The server returns {"output": {"text": "Ottawa is the capital of Canada.", "session_id": "..."}}.
  4. The second turn carries the session ID. It may land on a different instance; that is fine, because history is loaded from the sessions resource, not from memory.
  5. Under load, each instance accepts 9 concurrent requests. Ten users at once on one instance trigger a scale-out, and new JVMs pay their start-up cost before serving.

Observability and releases

Log one structured JSON line per turn to standard output with the session ID, user ID, class method, latency, number of model calls and outcome, and confirm where the runtime delivers container output in your project before you need it during an incident. Count in-band stream errors as failures in your metrics, since the HTTP status will say 200.

For releases, prefer deploying a new resource from a new immutable image tag and moving callers to it over mutating the live one: callers address a resource ID, so switching back is a configuration change, and because sessions live in their own resource both versions see the same history. Keep the old resource until the new one has served real traffic, and plan its JVM shutdown with the graceful shutdown patterns so in-flight turns finish when instances are scaled in.

Failure modes

  • Wrong port or loopback bind. Binding 127.0.0.1 or another port makes every call fail although the container runs. Bind 0.0.0.0 and honour PORT.
  • Undeclared class methods. If stream_query is not declared with api_mode stream, the playground and SDK streaming fail even though the endpoint exists.
  • Errors after the stream starts. Clients that only check the HTTP status treat a broken stream as success. Emit and check an error line.
  • Friendly app name. Passing a name instead of the sessions resource ID makes every session call fail.
  • JVM sized for the host. Without a heap cap the JVM can exceed the container limit and be killed. Use MaxRAMPercentage and exit on out-of-memory.
  • Cold starts. min_instances 0 saves money, but the first request then waits for a JVM to start. Keep at least one instance for interactive agents.
  • Image pull permission. A missing Artifact Registry reader role fails the deployment, not the request; check the operation result.
  • Non-idempotent tools. Retries by clients or the platform can re-run a tool that charged a card. Pass idempotency keys to side-effecting tools.

Trade-offs

Agent Runtime gives you the managed API, IAM on every call, the console playground and scaling without writing infrastructure, at the cost of a fixed request shape, resource limits chosen from a short list, and a deployment path that treats Java as a generic container. Kubernetes gives the most control and the most work. Cloud Run sits between them: your own HTTP API and routing, the same sessions service, and no runtime contract. Choose Agent Runtime when other teams will call the agent through Google Cloud tooling or when the playground and managed identity matter; choose Cloud Run when you need a custom API or WebSockets.

What to do next

  1. Create a sessions-only Agent Engine resource per environment and record its ID.
  2. Wrap your Runner in a server that implements both contract endpoints and validates class_method and input.
  3. Build a Java 21 image with a capped heap, push it to Artifact Registry with an immutable tag.
  4. Deploy with container_spec, both class methods declared, a dedicated service account and min_instances of at least 1.
  5. Call :query and :streamQuery with curl, then try the console playground.
  6. Load test to find the right container_concurrency, and add an in-band error check to every streaming client.
Key takeaway: ADK Java has no one-command deploy to Agent Engine, but the custom-container route works: implement the runtime contract's query and stream_query endpoints around a Runner, size the JVM to the container, deploy the image with container_spec and declared class methods, and keep sessions in a separately created resource.