Teams that already run their web front end on Cloudflare often ask whether an ADK Java agent can live there too, close to users. The honest answer has two parts. A Cloudflare Worker runs code in V8 isolates: JavaScript, TypeScript, WebAssembly and Python. There is no JVM, and ADK Java, with its reflection, RxJava, HTTP clients and Google client libraries, cannot be squeezed into a Worker. What does work is Cloudflare Containers, which run an ordinary Linux container image next to a Worker, addressed through a Durable Object. The Worker becomes the edge layer, the container runs the agent.
This article builds that shape end to end: the server inside the container, the Container class and Worker that route to it, the wrangler configuration, session state that survives sleeping instances, streaming, cold starts, shutdown and the cases where you should not do this at all. It assumes you can already run ADK Java behind HTTP; if not, the packaging and Runner server are covered step by step in deploying ADK Java to Cloud Run.
The shape that works
Each layer has one job. The Worker terminates the request at the nearest Cloudflare location, authenticates it, applies per-user limits and decides which container instance should serve it. The Container class is a Durable Object subclass: each named instance owns at most one running container, starts it on demand, forwards requests to its port and stops it after an idle period. The JVM inside runs a normal ADK Runner behind an HTTP server. Sessions and memory live in an external store, because the container's disk is reset whenever it sleeps.
The model call dominates latency, often seconds per turn, and it goes to wherever your model endpoint is. So the gain from the edge is not faster reasoning; it is fast TLS and auth near the user, rejection of bad requests before they reach a JVM, and one platform for front end and agent. If those do not matter to you, a regional service is simpler.
Architecture
The agent server inside the container
Inside the container, run the same kind of server as on any platform, listening on the port the Container class will call. Use RunConfig.StreamingMode.SSE so events stream out as the model produces them; the event types and partial versus final events are explained in ADK Java streaming events. The Worker passes the authenticated user id in a header that only it can set.
public final class EdgeServer {
public static void main(String[] args) throws Exception {
BaseSessionService sessions = new VertexAiSessionService(); // never in-memory here
Runner runner = Runner.builder().agent(AgentDef.ROOT_AGENT)
.appName(System.getenv("ADK_APP")).sessionService(sessions).build();
RunConfig cfg = RunConfig.builder()
.setStreamingMode(RunConfig.StreamingMode.SSE).setMaxLlmCalls(20).build();
HttpServer http = HttpServer.create(new InetSocketAddress("0.0.0.0", 8080), 0);
http.setExecutor(Executors.newVirtualThreadPerTaskExecutor());
http.createContext("/healthz", ex -> reply(ex, 200, "ok"));
http.createContext("/run", ex -> {
String user = ex.getRequestHeaders().getFirst("X-Edge-User"); // set by the Worker only
String sid = ex.getRequestHeaders().getFirst("X-Session-Id");
String msg = new String(ex.getRequestBody().readAllBytes(), UTF_8);
ex.getResponseHeaders().set("Content-Type", "text/event-stream");
ex.sendResponseHeaders(200, 0);
try (OutputStream out = ex.getResponseBody()) {
runner.runAsync(user, sid, Content.fromParts(Part.fromText(msg)), cfg)
.blockingForEach(ev -> {
out.write(("data: " + ev.toJson() + "\n\n").getBytes(UTF_8));
out.flush();
});
}
});
Runtime.getRuntime().addShutdownHook(new Thread(() -> http.stop(30)));
http.start();
}
}Imports are omitted for brevity; the types come from the ADK packages com.google.adk.runner, com.google.adk.agents and com.google.adk.sessions, the GenAI Content and Part types, and the JDK's com.sun.net.httpserver. Build the image for linux/amd64; Cloudflare requires that architecture, so an image built on an ARM laptop without a platform flag will not run. Use a JRE base image, not a JDK, and set -XX:MaxRAMPercentage=75 so the heap fits the instance's memory with room for threads and native buffers. Virtual threads need Java 21 or later.
FROM --platform=linux/amd64 eclipse-temurin:21-jre
WORKDIR /app
COPY target/edge-agent-all.jar /app/agent.jar
EXPOSE 8080
USER 10001
ENTRYPOINT ["sh", "-c", "printf '%s' \"$GCP_SA_KEY_JSON\" > /tmp/sa.json && export GOOGLE_APPLICATION_CREDENTIALS=/tmp/sa.json && exec java -jar /app/agent.jar"]Run the image locally first with the same environment variables the Worker will pass, and send it a request with the two headers by hand. If it cannot stream a turn on your laptop, nothing at the edge will fix that, and the local loop is far faster to debug.
The Worker, Container class and configuration
The Worker side is a Container subclass plus a fetch handler. Secrets set on the Worker are visible in its env, so copy the ones the JVM needs into envVars in an ordinary subclass constructor; check the package typings for the exact constructor signature in your version. Google client libraries read a key file path from GOOGLE_APPLICATION_CREDENTIALS, not key JSON, so the container passes the key under a custom name and the image entrypoint writes it to a file before starting Java. ADK_APP must be the Agent Engine resource ID that VertexAiSessionService stores sessions under, not a friendly name.
import { Container } from "@cloudflare/containers";
export class AdkAgent extends Container {
defaultPort = 8080;
sleepAfter = "15m";
constructor(ctx: DurableObjectState, env: Env) {
super(ctx, env);
this.envVars = {
ADK_APP: env.ADK_APP,
GOOGLE_CLOUD_PROJECT: env.GOOGLE_CLOUD_PROJECT,
GOOGLE_CLOUD_LOCATION: env.GOOGLE_CLOUD_LOCATION,
GCP_SA_KEY_JSON: env.GCP_SA_KEY, // Worker secret; see ENTRYPOINT
JAVA_TOOL_OPTIONS: "-XX:MaxRAMPercentage=75 -XX:+UseSerialGC",
};
}
}
const SHARDS = 8;
export default {
async fetch(req: Request, env: Env): Promise<Response> {
const user = await verifyJwt(req, env); // your auth; 401 on failure
if (!user) return new Response("unauthorized", { status: 401 });
const sid = req.headers.get("X-Session-Id");
if (!sid) return new Response("missing session", { status: 400 });
const h = new Uint32Array(await crypto.subtle.digest("SHA-256",
new TextEncoder().encode(user)))[0];
const stub = env.ADK_AGENT.getByName(`shard-${h % SHARDS}`);
const fwd = new Request(new URL("/run", req.url), req);
fwd.headers.delete("X-Edge-User"); // never trust the client's copy
fwd.headers.set("X-Edge-User", user);
return stub.fetch(fwd); // SSE body streams through
}
};{
"name": "adk-edge",
"main": "src/index.ts",
"compatibility_date": "2026-10-01",
"containers": [
{ "class_name": "AdkAgent", "image": "./Dockerfile",
"instance_type": "standard-1", "max_instances": 8 }
],
"durable_objects": { "bindings": [ { "name": "ADK_AGENT", "class_name": "AdkAgent" } ] },
"exports": { "AdkAgent": { "type": "durable-object", "storage": "sqlite" } }
}Deploy with npx wrangler deploy, which builds and pushes the image. Older projects declare the class with a migrations array instead of exports; the two cannot be combined. Containers require the Workers Paid plan, and the first requests after a first deploy can fail for a few minutes while images are distributed.
Sizing and routing
| Instance type | vCPU | Memory | Disk | Fit for an ADK Java agent |
|---|---|---|---|---|
| lite | 1/16 | 256 MiB | 2 GB | Too small for a JVM with Google client libraries |
| basic | 1/4 | 1 GiB | 4 GB | Toy agents only; slow startup |
| standard-1 | 1/2 | 4 GiB | 8 GB | Reasonable default for one agent per instance |
| standard-2 | 1 | 6 GiB | 12 GB | Many concurrent sessions, heavier tools |
| standard-3 / standard-4 | 2 / 4 | 8 / 12 GiB | 16 / 20 GB | CPU-heavy tools; rarely needed since the model does the work |
These are the predefined types on Cloudflare's limits page as of October 2026; custom sizes up to standard-4 are also available. An agent mostly waits on model I/O, so one JVM with virtual threads serves many concurrent sessions, which is why the Worker shards users over a small fixed set of named instances rather than starting one container per session. A container per session gives perfect isolation but pays a JVM's memory and cold start for every conversation. The documented getRandom(binding, n) helper spreads requests over n instances without affinity; hashing the user, as above, keeps a user on one instance so connection pools and caches stay warm, while the session store keeps correctness independent of which instance serves a turn.
Session state on ephemeral disk
Cloudflare's documentation is explicit that container disk is ephemeral: when an instance sleeps, the next start uses a fresh disk from the image. The InMemorySessionService default therefore loses every conversation after sleepAfter of idleness, after a deploy, or whenever a host restarts. Use an external session service: VertexAiSessionService if you are on Google Cloud, or a database-backed store such as the Postgres schema described in the adk-postgres session table. The Durable Object's own SQLite storage survives container restarts, but it is per shard and the JVM cannot reach it directly, so do not try to make it your session store.
Worked example: one turn
Walk one turn through the system. A user in Frankfurt sends a message with a session id. The Worker in a nearby location validates the token in a few milliseconds, hashes the user to shard 5 and calls getByName("shard-5"). If shard 5 is running, the request reaches the JVM; Cloudflare notes a container may run in a different location from its Durable Object, so the hop is not always local. If shard 5 slept, the container starts: Cloudflare cites cold starts often in the 1-3 second range for the container, and the JVM's own startup and class loading come on top, so expect several seconds for the first turn. The runner loads the session from the store, calls the model, and streams events back through the Worker to the browser. After 15 idle minutes the shard sleeps and its disk is discarded; the next turn starts it again and reloads the session.
Operations
Cold starts. Keep the image small, use class data sharing (AppCDS) to cut JVM startup, set sleepAfter long enough to cover typical gaps between turns, and have the Worker show a 'starting' event rather than a blank screen. GraalVM native images would cut startup further, but native-image support for ADK Java and its dependencies is not documented, so treat it as an experiment you must test.
Shutdown. On stop, Cloudflare sends SIGTERM and waits up to 15 minutes before SIGKILL. Stop accepting new turns, let running ones finish or checkpoint, and exit; the drain pattern is in ADK Java graceful shutdown. Because sessions are external, a turn cut off mid-stream can be retried by the client with the same session id.
Trust boundary. End users cannot reach a container directly; every request passes through the Worker, which is what makes the X-Edge-User header trustworthy. Keep model credentials scoped to model access only, and rotate the Worker secret that carries them.
Observability. Log the shard, session id and a request id in both the Worker and the JVM so a slow turn can be split into edge, cold start, model and tool time.
Cost control. The Worker is the cheapest place to say no. Reject unauthenticated calls, oversized bodies and users over their turn budget there, before a container wakes up, and cap max_instances so a traffic spike cannot start more JVMs than your model quota can feed. A turn that waits on a saturated model endpoint still holds container memory, so the real ceiling is model throughput, not container count.
Trade-offs
Compared with a regional service such as Cloud Run, this design adds a platform-specific routing layer and Durable Object semantics you must learn, and the model endpoint is still regional, so per-turn latency improves little. It wins when your front end and auth already live in Workers, when you want to reject abusive traffic before it costs JVM time, or when you need many small, isolated agent instances that sleep to zero cost when idle. If you need guaranteed warm capacity, long-lived WebSocket sessions with bidirectional audio, or in-process access to Google Cloud networking, a regional deployment is the simpler choice.
Failure modes
- Trying to compile ADK Java to WebAssembly to fit a Worker; the dependency tree does not support it.
- Leaving the in-memory session service on, so conversations vanish at every sleep or deploy.
- One container per session, which multiplies JVM memory and cold starts.
- Building an ARM image on a laptop, which fails to start on the platform.
- Trusting a user id header the client can set, instead of overwriting it in the Worker.
- A sleepAfter shorter than the gap between turns, so every turn pays a cold start.
What to do next
- Run your agent behind the HTTP server above locally with an external session store.
- Build a linux/amd64 JRE image and measure its startup time and resident memory.
- Create the Worker, Container class and wrangler config, and deploy on the Workers Paid plan.
- Load-test one shard with concurrent sessions to choose instance type and shard count.
- Measure first-turn latency after sleep and tune sleepAfter and AppCDS until it is acceptable.
- Kill a shard mid-turn and confirm the client can resume the session on retry.
- Write down why the edge is worth it for you; if you cannot, deploy regionally instead.