A container is not a virtual machine. It is an ordinary Linux process whose view of the system is narrowed by namespaces and whose resources are capped by control groups (cgroups). The JVM was designed when it could assume it owned the machine, so for years it read the host's CPU count and memory, sized its heap and thread pools for a 64-core server, and was then killed by a 2 GiB limit it never knew about.
Modern JDKs read cgroup limits, but that only solves half the problem. You still have to choose how much of the limit goes to the heap, understand what else the JVM allocates, decide how CPU limits interact with garbage collection, build an image that is small and reproducible, start fast enough for autoscaling, and shut down cleanly. This article covers each of those from first principles, with a worked memory budget, real Dockerfiles and a checklist. For general container concepts see cloud containers.
What the JVM detects, and how to check
On Linux, container support (-XX:+UseContainerSupport) is on by default in current JDKs. At startup the JVM finds its cgroup, reads the memory limit (memory.max on cgroup v2, memory.limit_in_bytes on v1) and the CPU quota (cpu.max on v2), and uses them instead of host values. cgroup v2 support arrived in JDK 15 and was backported to later updates of 8 and 11; if you run an old update on a cgroup v2 host, which is the default on current distributions and Kubernetes nodes, the JVM may silently fall back to host values. Check your vendor's release notes rather than assuming.
Never guess what the JVM saw. Ask it, inside the running container:
# What the JVM believes about its container
java -XshowSettings:system -version
# Detailed detection log at startup
java -Xlog:os+container=info -version
# The resulting ergonomic choices
java -XX:+PrintFlagsFinal -version | grep -E 'MaxHeapSize|UseSerialGC|UseG1GC|ActiveProcessorCount'
# In a running JVM (PID 1 in the container)
jcmd 1 VM.flags
jcmd 1 VM.info | grep -A5 'container'Make the first command part of your image smoke test. If it reports the host's memory instead of the limit, nothing else in this article will behave as described.
How limits become heap, GC and thread counts
Memory. Without -Xmx, the maximum heap is MaxRAMPercentage of the limit, and the default is 25 percent. That is safe and wasteful: in a 2 GiB container you get a 512 MiB heap and 1.5 GiB of mostly idle headroom. The usual choice is to set -XX:MaxRAMPercentage between 60 and 75 and leave the rest for non-heap memory, or to set -Xmx explicitly when the limit is fixed.
CPU. Runtime.availableProcessors() is derived from the CPU quota (limit divided by period, rounded up) and any cpuset. Since JDK 19, CPU shares (Kubernetes CPU requests) are no longer used to compute it, so a pod with a request but no limit sees the node's CPUs. -XX:ActiveProcessorCount=N overrides the value when you need to.
GC choice. When the JVM sees fewer than 2 processors or less than about 1792 MiB of memory, it does not treat the machine as server class and picks Serial GC instead of G1. A container with one CPU and 1 GiB therefore runs single-threaded stop-the-world collections unless you choose otherwise. Choose explicitly: G1 for most services (see G1 GC in depth), Serial for tiny single-CPU utilities, ZGC for large heaps with strict pause targets.
Everything else that is sized from availableProcessors() follows: parallel and concurrent GC threads, JIT compiler threads, the common ForkJoin pool (processors minus one), the virtual thread scheduler's carrier threads (processors), and many frameworks' default pools. A wrong CPU count distorts all of them at once.
Worked example: a memory budget for a 2 GiB container
The heap is only one region. Metaspace, the code cache, thread stacks, GC bookkeeping, direct buffers and native allocations all count against the same limit, which is why resident memory always exceeds -Xmx; the regions are explained in the JVM memory model. For a typical Spring or Micronaut service with 200 platform threads, a realistic budget looks like this:
| Region | Budget | Controlled by |
|---|---|---|
| Java heap | 1,280 MiB | -XX:MaxRAMPercentage=62.5 |
| Metaspace | 150 MiB | -XX:MaxMetaspaceSize |
| Code cache | 100 MiB | -XX:ReservedCodeCacheSize |
| Thread stacks (200 x up to 1 MiB, touched pages only) | about 100 MiB | -Xss, thread count |
| GC structures (G1) | about 60 MiB | heap size, region count |
| Direct buffers | 128 MiB | -XX:MaxDirectMemorySize |
| Native: malloc, libraries, TLS | about 100 MiB | measure with NMT |
| Headroom | about 130 MiB | - |
The numbers are estimates to replace with measurements. Run the service under realistic load with -XX:NativeMemoryTracking=summary and read jcmd 1 VM.native_memory summary, then compare the committed total with the container's memory.current. Cap direct memory explicitly; by default its limit equals the maximum heap size, which alone can push the total past the container limit. Add -XX:+ExitOnOutOfMemoryError so a Java heap OOM ends the process cleanly instead of leaving a half-working JVM behind a passing health check.
CPU limits and throttling
A CPU limit in Kubernetes becomes a CFS quota: for example, 200 ms of CPU time per 100 ms period for a limit of 2. A JVM that runs a burst of parallel GC threads, JIT compilations and request threads can consume the whole quota early in the period and then be throttled, meaning paused, for the rest of it. That appears as latency spikes that do not show in the GC logs. Watch nr_throttled and throttled_usec in cpu.stat or the equivalent container metrics.
Three practical rules. Give latency-sensitive Java services enough CPU limit for startup and GC bursts, or remove the limit and rely on requests where your platform policy allows it. Do not set the limit below 2 for a G1 service unless you have measured it. And if you remove the limit, set -XX:ActiveProcessorCount so thread pools are sized for the CPUs you actually expect to get, not for the whole node.
Building the image
A good Java image is small, reproducible and layered so that a code change does not re-upload dependencies. A multi-stage build compiles with a full JDK, uses jdeps and jlink to build a runtime containing only the modules the application needs, and copies the result into a minimal base:
# syntax=docker/dockerfile:1
FROM eclipse-temurin:25-jdk AS build
WORKDIR /src
COPY . .
RUN ./mvnw -q -DskipTests package dependency:copy-dependencies \
-DincludeScope=runtime -DoutputDirectory=target/lib
RUN jdeps --ignore-missing-deps --print-module-deps --multi-release 25 \
--class-path "$(ls target/lib/*.jar | paste -sd:)" target/app.jar > modules.txt
RUN jlink --add-modules $(cat modules.txt),jdk.jcmd,jdk.jfr --strip-debug --no-man-pages \
--no-header-files --output /opt/jre
FROM debian:bookworm-slim
RUN useradd --uid 10001 app
COPY --from=build /opt/jre /opt/jre
COPY --from=build /src/target/lib /app/lib
COPY --from=build /src/target/app.jar /app/app.jar
USER 10001
ENV JAVA_TOOL_OPTIONS="-XX:MaxRAMPercentage=62.5 -XX:+ExitOnOutOfMemoryError"
ENTRYPOINT ["/opt/jre/bin/java", "-cp", "/app/app.jar:/app/lib/*", "com.example.Main"]Note the order: dependencies are copied in a separate layer from the application jar, so a typical rebuild changes only the last layer. Spring Boot 3.3 and later can do this for a fat jar with java -Djarmode=tools -jar app.jar extract --layers --destination extracted, which splits it into dependency, loader, snapshot and application directories that you copy as separate layers. Jib builds equivalent layered images from Maven or Gradle without a Dockerfile or a Docker daemon.
Keep the image honest: run as a non-root user, pin base images by digest in CI, and include jdk.jcmd and jdk.jfr (as above) or plan to attach diagnostic tools from a debug container, because minimal images have no shell and no jcmd by default.
Startup: CDS, the AOT cache, CRaC and native images
Autoscaling and short-lived jobs make startup time matter. The options, from least to most invasive:
- Class Data Sharing. The JDK ships a default CDS archive for its own classes. An application CDS archive created at build time adds your classes and typically cuts class-loading time noticeably.
- AOT cache (JDK 24 and 25). JEP 483 in JDK 24 extends CDS to store classes already loaded and linked from a training run. JDK 25's JEP 514 simplifies creation to one step,
java -XX:AOTCacheOutput=app.aot -jar app.jarduring a representative training run, thenjava -XX:AOTCache=app.aot -jar app.jarin production; JEP 515 adds method profiles to the cache so the JIT warms up sooner. The cache must be created with the same JDK and class path it is used with, so build it inside the image build. - CRaC. Checkpoint a warmed-up JVM and restore it in milliseconds; it needs application cooperation for open files and sockets, as described in Java CRaC.
- Native image. GraalVM compiles ahead of time to a native executable with very fast startup and small memory, at the cost of build time, reflection configuration and peak throughput.
Signals, shutdown and probes
Kubernetes stops a pod by sending SIGTERM to PID 1, waiting terminationGracePeriodSeconds (30 by default), then sending SIGKILL. The JVM handles SIGTERM by running shutdown hooks, which is where frameworks drain requests and close pools. This only works if the JVM is PID 1 or receives the signal. The shell form ENTRYPOINT java -jar app.jar runs under /bin/sh -c, which may not forward the signal, so the JVM is killed without shutting down. Use the exec (JSON array) form, or exec java ... in a wrapper script.
Readiness and liveness probes should check different things. Readiness says whether to send traffic and may fail during warm-up or dependency outages. Liveness says whether to restart and should fail only when the process is truly stuck; a liveness probe that calls the database restarts every pod during a database incident. Give slow-starting JVMs a startup probe so liveness does not kill them before they are ready.
Diagnosing the common failures
| Symptom | Likely cause | What to check or change |
|---|---|---|
| Exit code 137, no Java stack trace | Kernel OOM kill: total memory over the limit | NMT summary against memory.current; lower heap percentage; cap direct memory |
java.lang.OutOfMemoryError in logs | Heap too small or a leak | -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/dumps on a mounted volume |
| Long pauses on a small pod | Serial GC chosen by ergonomics | Check PrintFlagsFinal; select G1 explicitly; add a CPU |
| Latency spikes, clean GC logs | CFS throttling | cpu.stat throttling counters; raise or remove the limit |
| Thread pools sized for 64 cores | No CPU limit, or an old JDK on cgroup v2 | -XshowSettings:system; -XX:ActiveProcessorCount |
| Dropped requests on every deploy | SIGTERM not reaching the JVM | Exec-form entrypoint; graceful shutdown enabled; preStop delay |
For live diagnosis in a minimal image, attach an ephemeral container that shares the process namespace, for example with kubectl debug -it pod/app --image=eclipse-temurin:25-jdk --target=app, and run jcmd from there; the JDK versions should match. Java Flight Recorder (jcmd 1 JFR.start duration=60s filename=/tmp/rec.jfr) is cheap enough to use in production.
What to do next
- Run
java -XshowSettings:system -versioninside each production image and confirm it reports the container's limits. - Set
MaxRAMPercentage(or-Xmx),MaxDirectMemorySizeand-XX:+ExitOnOutOfMemoryErrorexplicitly, then verify the budget with Native Memory Tracking under load. - Choose the garbage collector explicitly and check that pods have at least the CPUs it needs; measure throttling before and after.
- Move to a multi-stage, layered, non-root image, using jlink or Spring Boot layer extraction, with base images pinned by digest.
- On JDK 25, add an AOT cache step to the image build and measure startup time and time-to-peak against the current image.
- Switch entrypoints to exec form, enable graceful shutdown, and separate readiness, liveness and startup probes.