A Docker image is the unit an ADK Java agent ships as, whether it ends up on Cloud Run, Kubernetes or a laptop. The difference between a casual image and an engineered one is large: build time on every commit, startup latency during a traffic spike, the size of the attack surface a scanner reports, and whether a leaked image also leaks your model credentials.
This article is about the image itself. It explains how layers and caching work for a Java agent, builds a small runtime with jdeps and jlink on a distroless base, compares that with Jib, sizes the JVM for a container memory limit, cuts startup with the JDK 25 AOT cache, and covers secrets, signals and supply-chain checks. Platform specifics such as Cloud Run flags or Kubernetes probes live in their own articles, linked at the end.
What belongs in an agent image
An ADK Java agent image needs four things: a Java runtime, your application classes, the dependency JARs (google-adk, the Google GenAI client, Jackson, RxJava and whatever your tools pull in), and a start command. It must not contain build tools, source code, test fixtures, credentials or a shell you do not need. Each of those adds size, and each extra binary is something a scanner flags and an attacker can use.
Note what the image should not run: ADK's development web server. The google-adk-dev artifact and com.google.adk.web.AdkWebServer are how the documentation demonstrates agents, and the documented Cloud Run route even compiles at container start. For production, build at image time and start your own server class around Runner, which the Cloud Run article walks through.
Layers and the build cache
An image is a stack of read-only layers, one per filesystem-changing instruction, each identified by a hash of its content. When you rebuild, Docker reuses a cached layer only if the instruction and everything below it are unchanged; the first changed layer invalidates every layer above it. Registries work the same way: a push uploads only layers the registry lacks, and a node pulls only layers it lacks.
So order layers from least to most frequently changed. The base image changes monthly, the runtime with your JDK version, dependencies when the POM changes, and the application on every commit. A fat JAR breaks this: it bundles dependencies and code into one file, so every commit produces a new multi-megabyte layer. Keep dependencies in their own directory and copy them in a layer below the application JAR.
A multi-stage Dockerfile with a jlink runtime
The Dockerfile below uses three stages. The build stage resolves dependencies in a layer keyed on pom.xml alone, so source edits do not re-download the internet. The runtime stage uses jdeps to list the JDK modules the application needs and jlink to assemble a runtime containing only those. The final stage is a distroless base with glibc and CA certificates but no shell or package manager, running as a non-root user.
# syntax=docker/dockerfile:1
FROM maven:3.9-eclipse-temurin-21 AS build
WORKDIR /src
COPY pom.xml .
RUN --mount=type=cache,target=/root/.m2 mvn -B -q dependency:go-offline
COPY src ./src
RUN --mount=type=cache,target=/root/.m2 \
mvn -B -q package -DskipTests dependency:copy-dependencies -DincludeScope=runtime
FROM eclipse-temurin:21-jdk AS runtime
COPY --from=build /src/target/*.jar /tmp/app.jar
COPY --from=build /src/target/dependency /tmp/lib
RUN jdeps --ignore-missing-deps --print-module-deps --multi-release 21 \
--class-path '/tmp/lib/*' /tmp/app.jar > /tmp/modules \
&& jlink --add-modules "$(cat /tmp/modules),jdk.crypto.ec" \
--strip-debug --no-man-pages --no-header-files --compress=zip-6 \
--output /opt/java
FROM gcr.io/distroless/base-debian12:nonroot
COPY --from=runtime /opt/java /opt/java
COPY --from=build /src/target/dependency /app/lib
COPY --from=build /src/target/*.jar /app/app.jar
ENV JAVA_TOOL_OPTIONS="-XX:MaxRAMPercentage=70 -XX:+ExitOnOutOfMemoryError"
USER nonroot
EXPOSE 8080
ENTRYPOINT ["/opt/java/bin/java", "-cp", "/app/app.jar:/app/lib/*", "com.example.AgentServer"]Three details are easy to get wrong. First, jdeps sees static references only. Modules loaded reflectively or through ServiceLoader, such as security providers, jdk.zipfs or java.naming, can be missing, and the symptom is a TLS handshake or class-loading error at the first model call, not at startup. The jdk.crypto.ec addition above is one such guess for JDK 21 TLS, and the right extra modules differ between JDK versions; treat the module list as something a smoke test confirms, not something you trust. Second, the build-stage JDK and the jlink JDK must be the same major version as your bytecode target. Third, ENTRYPOINT is in exec form, so the JVM is PID 1 and receives SIGTERM directly; the shell form wraps it in /bin/sh, which distroless does not have and which would swallow the signal anyway.
Pin the bases by digest in CI (FROM gcr.io/distroless/base-debian12:nonroot@sha256:...) so a rebuild of the same commit produces the same runtime, and let a bot such as Renovate or Dependabot propose digest updates as reviewable changes.
Jib: images without a Dockerfile
Jib, Google's Maven and Gradle plugin, builds images without a Dockerfile or a Docker daemon. It splits dependencies, resources and classes into separate layers automatically, sets file timestamps to a fixed value so the same input gives the same image digest, and pushes directly to a registry with mvn compile jib:build (or loads into a local daemon with jib:dockerBuild).
<plugin>
<groupId>com.google.cloud.tools</groupId>
<artifactId>jib-maven-plugin</artifactId>
<version>${jib.version}</version>
<configuration>
<from><image>eclipse-temurin:21-jre@sha256:...</image></from>
<to><image>us-central1-docker.pkg.dev/my-proj/agents/billing-agent:${git.sha}</image></to>
<container>
<mainClass>com.example.AgentServer</mainClass>
<jvmFlags><jvmFlag>-XX:MaxRAMPercentage=70</jvmFlag></jvmFlags>
<user>1000</user>
</container>
</configuration>
</plugin>The trade is control. Jib gives you good layering and reproducibility for free, but a custom jlink runtime or an AOT cache needs extra work: you build a base image containing the runtime separately and point from at it. Choose Jib when a standard JRE base is acceptable and you want the least build plumbing; choose the Dockerfile when image size or startup time justifies owning the runtime.
Sizing the JVM for a container
Modern JVMs read the container's cgroup limits, so heap defaults follow the memory limit rather than the host's RAM. The default maximum heap is a quarter of the limit, which wastes most of a small container, hence -XX:MaxRAMPercentage. Do not set it to 100: the heap is only part of the process.
| Consumer, 1 GiB limit | Typical budget | Notes |
|---|---|---|
| Java heap | 700 MiB (70%) | conversation history, tool results, JSON trees |
| Metaspace and code cache | 100-150 MiB | ADK, gRPC, Jackson and your classes; measure with NMT |
| Platform thread stacks | 1 MiB each by default | virtual threads park on the heap instead |
| Direct buffers, native libs | tens of MiB | Netty and gRPC transports allocate off-heap |
| Headroom | remainder | or the kernel OOM-kills the container with no Java stack trace |
Verify, do not guess: start the container with -XX:NativeMemoryTracking=summary, drive a load test, and run jcmd 1 VM.native_memory summary from a debug image (distroless has a :debug variant with a busybox shell for this). -XX:+ExitOnOutOfMemoryError makes a heap OOM kill the process so the platform restarts it, instead of leaving a zombie that accepts requests and fails them. If the container is limited to fractional CPUs, the JVM may size its collector and fork-join pools for one CPU; -XX:ActiveProcessorCount overrides that when measurement shows it helps.
Faster startup with the AOT cache
A JVM agent spends its first seconds loading and linking thousands of classes: ADK, the GenAI client, Jackson, RxJava, your HTTP server. Class Data Sharing archives that work. JDK 24 added ahead-of-time class loading and linking (JEP 483), and JDK 25's JEP 514 made it one step: run the application once with -XX:AOTCacheOutput=app.aot as a training run, and later runs use -XX:AOTCache=app.aot.
# In a build stage on JDK 25, with the same JDK and classpath as the final image:
RUN java -XX:AOTCacheOutput=/app/app.aot -cp "/app/app.jar:/app/lib/*" \
com.example.AgentServer --training-run
# Final image:
ENTRYPOINT ["/opt/java/bin/java", "-XX:AOTCache=/app/app.aot", \
"-cp", "/app/app.jar:/app/lib/*", "com.example.AgentServer"]The catch is the training run. A server never exits on its own, so add a --training-run mode that builds the agent and runner, sends a few requests through a stub model so the hot classes load, and exits; it must not call the real model during a build. The cache is only valid for the same JDK build and a compatible classpath, so it belongs in the same image build as the JARs it describes, and the JVM ignores an unusable cache rather than failing. Measure startup before and after on your agent; do not ship it on faith.
Secrets and the build context
Credentials never enter an image. Every layer is readable by anyone who can pull the image, and deleting a file in a later layer does not remove it from the earlier one. ARG values and ENV values are recorded in the image history. On Google Cloud, let the platform supply Application Default Credentials through the runtime service account or workload identity; elsewhere, mount a credentials file at run time and point GOOGLE_APPLICATION_CREDENTIALS at it, or inject a Gemini API key from a secret manager as an environment variable at deploy time. If the build itself needs a secret, for example a private Maven repository, use a BuildKit secret mount (RUN --mount=type=secret,id=maven), which is not written to a layer.
A .dockerignore listing target/, .git/, *.json key files and local .env files keeps them out of the build context entirely.
Supply chain: SBOMs, scanning and architectures
Generate a software bill of materials and scan it in CI: docker buildx build --sbom=true attaches one to the image, and scanners such as Trivy or Grype read images or SBOMs. A distroless jlink image typically produces far fewer findings than a full OS base, because most findings come from OS packages you never call. Build for every architecture you deploy to with docker buildx build --platform linux/amd64,linux/arm64; an ARM node pulling an amd64-only image fails to start. Tag images with the git commit, deploy by digest, and treat latest as a convenience for laptops only.
Failure modes
The failures below all reach production regularly.
- Missing module at first request.
jlinkdropped a reflectively loaded module; startup is fine and the first TLS call to the model fails. Fix: a CI smoke test that runs the image and makes one real HTTPS request. - OOM-kill with no stack trace. Heap plus native memory exceeded the limit. Fix: lower
MaxRAMPercentageand measure native memory. - Slow shutdown, dropped turns. Shell-form entrypoint, so
SIGTERMnever reached the JVM and in-flight agent turns were killed at the grace deadline. - Every build re-downloads dependencies.
COPY . .before the dependency step invalidates the cache on any file change. - Credentials in history. A key passed with
--build-argis visible indocker history. Rotate it; deleting the image does not undo a pull. - Container HEALTHCHECK ignored. Kubernetes and Cloud Run do not use Dockerfile
HEALTHCHECK; configure the platform's probes instead.
Trade-offs
| Approach | Strength | Weakness |
|---|---|---|
| Official JRE base + Dockerfile | simple, familiar, has a shell | larger, more scanner findings |
| jlink runtime on distroless | small, minimal attack surface | module list to maintain; no shell |
| Jib | no daemon, reproducible, auto-layered | custom runtime or AOT needs a custom base |
| AOT cache (JDK 25) | faster startup, same code | training-run mode; tied to JDK build and classpath |
| GraalVM native image | fastest startup, small memory | reflection configuration for ADK, Jackson and gRPC; not verified here |
What to do next
- Split your fat JAR into an application JAR plus a dependency directory and reorder the Dockerfile so dependencies sit below code.
- Build the jlink runtime on a distroless base and add a CI smoke test that makes one real model call from the container.
- Set
MaxRAMPercentagefrom a native-memory measurement under load; see virtual threads in ADK Java for why thread count matters less. - Make the JVM PID 1 and confirm in-flight turns finish on
SIGTERM, following graceful shutdown. - Move every credential to run-time injection as described in agent secrets management, and add a
.dockerignore. - Add SBOM generation, scanning and digest pinning, then deploy with Cloud Run or Kubernetes.