Project Leyden is the OpenJDK effort to improve the startup time, time to peak performance and footprint of Java programs by shifting work out of each production run into an earlier phase. Its shipped result so far is the AOT cache: a training run records what the JVM loads, links and profiles, an assembly step writes that into a file, and later runs map the file and skip the recorded work. How the cache works internally, what each JEP added and the consistency rules the JVM enforces are explained in detail in Project Leyden architecture in depth; read that first if the mechanism is new to you.

This article is about the step after understanding: putting the cache into production across many services without surprises. It covers how to measure whether Leyden helps you, how to design the training run, what Spring Boot and Quarkus do for you, how to manage the cache as a build artifact, how to roll it out and observe it, the failure modes that only appear at fleet scale, and what is on the roadmap. The goal is that you can leave with a plan for one pilot service this sprint.

Advertisement

Where Leyden stands in 2026

Leyden ships in pieces through ordinary JEPs, and each piece is usable on its own. According to the project page, four are delivered: JEP 483, ahead-of-time class loading and linking, in JDK 24; JEP 514, command-line ergonomics that train and assemble in one step, and JEP 515, cached method profiles, both in JDK 25; and JEP 516, cached heap objects usable with any garbage collector including ZGC, in JDK 26. Ahead-of-time code compilation, which would put JIT-compiled native code into the cache, is an unnumbered draft JEP, JDK-8335368, with no target release.

Two practical consequences follow. JDK 25, the current long-term-support release line, is the natural baseline: it has the one-step workflow and profiles, and it is what Spring Boot requires for its AOT cache support. And the cache's contents will keep growing with each release, so your investment is in the pipeline, the training run and the telemetry, not in any one release's feature set. Experimental work happens in the project's premain branch, with early-access builds published separately; use those to evaluate what is coming, never in production. Note also that JEP 483 now records that the strict mode -XX:AOTMode=on is spelled -XX:AOTMode=required as of JDK 27, which matters for pipelines that span releases.

Measure before you adopt

Leyden accelerates class loading, linking and warmup, so its benefit is proportional to how much of your startup is spent on those. A service that spends four seconds loading ten thousand framework classes gains a lot; a service that spends four seconds waiting for a connection pool and a remote configuration server gains little. Before building anything, measure three numbers for a candidate service: time from process start to the first successful health check, time until throughput or p99 latency reaches its steady state under load, and resident memory at steady state.

Measure them the same way every time. Run on the same machine type as production, with the same container CPU limit, because startup is CPU-bound and a two-CPU limit changes everything. Take at least a dozen runs and report the median, because startup times have a long tail from disk cache and scheduling noise. Keep external dependencies local or stubbed so you are measuring the JVM, not the network. A small harness is enough:

#!/usr/bin/env bash
# startup.sh MODE  -> prints milliseconds from exec to first HTTP 200, and RSS in KB
set -euo pipefail
MODE=$1                                        # "base" or "aot"
OPTS=""; [ "$MODE" = aot ] && OPTS="-XX:AOTCache=app.aot"
for i in $(seq 1 15); do
  t0=$(date +%s%N)
  java $OPTS -jar app.jar --server.port=18080 > /dev/null 2>&1 &
  pid=$!
  until curl -sf http://localhost:18080/health > /dev/null; do sleep 0.02; done
  t1=$(date +%s%N)
  rss=$(ps -o rss= -p $pid)
  echo "$MODE $(( (t1 - t0) / 1000000 )) $rss"
  kill $pid; wait $pid 2>/dev/null || true
done | sort -k2 -n | awk '{a[NR]=$0} END {print "median:", a[int((NR+1)/2)]}'

Run it with base and with aot once you have a cache. For warmup, drive a fixed load and record per-second throughput for the first few minutes; the interesting number is how many seconds it takes to reach, say, 90 percent of peak. Record memory too: the cache is mapped from a file, so it can change resident memory in either direction, and you want to know which before you change container limits.

Advertisement

Designing the training run

The cache contains what the training run did and nothing else, so the training run is the real design decision. Think of it in three tiers. The minimal tier starts the application and exits once it is initialised; it captures most framework class loading, which is usually the bulk of the startup gain, and it is cheap and deterministic. The second tier adds a smoke pass over the main endpoints, which loads the classes on the request paths so that the first real requests do not pay for them. The third tier drives a short scripted load so that method profiles reflect realistic hot paths, which is what JEP 515 needs to improve warmup.

Whatever the tier, the training run executes your code at build time, so treat it like a test with side effects. Point it at in-memory or containerised fixtures, disable outbound calls, schedulers and message consumers, and make the build fail if it touches the real network. Keep it deterministic: a run that sometimes exercises a code path and sometimes does not produces caches of varying quality, which shows up as unexplained variance in canary startup times. And keep the training profile close to production configuration, because a different profile can activate different beans and load different classes.

The cached profiles are hints, not constraints. HotSpot keeps profiling in production and can deoptimise and recompile when behaviour differs, so an unrepresentative training load costs some warmup rather than correctness. How the JIT uses profiles, and what deoptimisation does, is covered in JVM JIT architecture.

What the frameworks do for you

Both major frameworks have made the training run a packaging step. Spring Boot supports the AOT cache on Java 25 and later, and its documentation points earlier releases to classic class data sharing instead. Its documented flow is to extract the executable jar into an exploded layout, run a training pass with -Dspring.context.exit=onRefresh, which refreshes the context, creating the eager singletons, then exits before any lifecycle beans start; production then runs with the cache. The cache only works with the extracted form, because the fat jar's nested class path is not something the JVM can validate.

Quarkus takes the third-tier approach by default. With quarkus.package.jar.aot.enabled=true, its Maven build runs the existing @QuarkusIntegrationTest suite against the packaged application as the training workload, and writes app.aot next to quarkus-run.jar. If your integration tests already cover the real request paths, you get representative profiles for free.

# Spring Boot on Java 25+: the cache only works with the extracted layout
java -Djarmode=tools -jar target/orders.jar extract --destination build/app
cd build/app
java -XX:AOTCacheOutput=app.aot -Dspring.context.exit=onRefresh \
     -Dspring.profiles.active=training -jar orders.jar
java -XX:AOTCache=app.aot -jar orders.jar

# Quarkus (from the project root): training is driven by @QuarkusIntegrationTest
./mvnw verify -Dquarkus.package.jar.aot.enabled=true -DskipITs=false
cd target/quarkus-app && java -XX:AOTCache=app.aot -jar quarkus-run.jar

For applications on neither framework, the JDK's own tools suffice: -XX:AOTCacheOutput for a one-step training run on JDK 25 and later, and the JDK_AOT_VM_OPTIONS environment variable for options that should apply only to the assembly phase.

The cache as a build artifact

The AOT cache as a build artifact: lifecycle in a delivery pipelineBuildjar + dependenciesTraining runfixtures, scripted loadapp.aotkeyed to JDK + classpathStrict smoke testAOTMode=on/requiredImagesame JDK buildCanarystartup, warmup, RSSFleetAOTMode=autoTelemetrycache used? how fast?Invalidation triggersJDK update, classpath change, new agentassemblepasspromoterebuildTreat the cache like a compiled binary: produced by the pipeline, verified strictly, never copied between builds.Production uses the forgiving mode so a stale cache costs startup time, not availability.
The cache is produced and strictly verified in the same pipeline run as the image it ships in, promoted through a canary, and rebuilt whenever the JDK build, class path or agents change.

The most important operational rule is that the cache belongs to exactly one build. It is valid only for the JDK build, operating system, architecture and class path it was trained with, so it must be produced in the same pipeline run as the image that will use it, from the same artifacts, and never cached between runs, copied from another branch or downloaded from a shared store. When it is invalid, the default mode silently ignores it, and the failure is invisible.

The pipeline therefore needs one strict check. Start the packaged application with the cache in the strict mode, which exits with an error if the cache cannot be used, and fail the build if it does. Record a fingerprint of the JDK version and the class path next to the cache so a later investigation can tell what it was built for:

#!/usr/bin/env bash
# ci/aot-cache.sh: build, verify and fingerprint the cache in the same job as the image
set -euo pipefail
JDK_VERSION=$(java -XshowSettings:properties -version 2>&1 | awk -F'= ' '/java.runtime.version/ {print $2}')
CP_HASH=$(cd build/app && find . -name '*.jar' | sort | xargs sha256sum | sha256sum | cut -c1-16)
KEY="${JDK_VERSION}-${CP_HASH}"

./ci/train.sh                                   # produces build/app/app.aot
echo "$KEY" > build/app/app.aot.key

# strict smoke test, same directory and command line as production:
# must start WITH the cache or fail the build (-XX:AOTMode=required from JDK 27)
(cd build/app && java -XX:AOTMode=on -XX:AOTCache=app.aot \
     -Dspring.context.exit=onRefresh -jar orders.jar)
echo "aot cache ok for $KEY"

Production then runs in the default mode, in which an unusable cache logs a warning and the JVM starts normally. That split, strict in CI and forgiving in production, means a mismatch is caught before release but can never stop a service from starting.

Rolling out and observing

Roll out like any performance change. Ship the cached image to a canary, compare the three baseline numbers on real traffic, and promote only if startup improved and neither warmup nor memory got worse. Revisit what depends on startup time: readiness probe delays, autoscaler cool-downs and deployment surge settings were tuned for the old numbers, and leaving them unchanged hides the benefit.

Then make cache use visible, because the silent fallback is the fleet-scale risk. At minimum, log at startup whether the cache was mapped, and alert when a service's startup time returns to its pre-Leyden baseline. The JVM's unified logging can report cache mapping and validation decisions; the available tags differ between releases, so list them with java -Xlog:help on your JDK. For deeper checks, compare class loading with -verbose:class or the JFR jdk.ClassLoad event, which is how JEP 483 suggests examining what a run loads; classes served from the cache report the shared archive as their source rather than a jar. The shared-archive mechanics behind this are covered in Java class data sharing.

Fleet-scale failure modes

  • Base image drift. The base image picks up a JDK patch release without a rebuild of the cache step, and every service loses its gain at once. Pin base images by digest and make the cache step part of every image build.
  • Agent mismatch. An APM or security agent that rewrites classes is attached in production but not in training, or is upgraded independently. Such agents invalidate the cache; train with the same agent configuration or measure without it.
  • Class path drift. Production starts from a different layout than training, such as the fat jar instead of the extracted directory, or an entrypoint script that reorders the class path. Appending entries is tolerated; reordering or replacing them is not.
  • Collector choice. Before JDK 26, cached heap objects could not be used with ZGC, so a ZGC service on JDK 25 gets a smaller gain than a G1 service. Compare like with like.
  • Nondeterministic training. Variable canary startup times across builds of the same code usually trace to a training run that depends on timing or data.
  • Image bloat. A big application produces a big cache file; keep it in a separate, late image layer so a code-only change does not force nodes to re-pull the dependencies.

When Leyden is not enough

Leyden keeps full Java semantics, which makes it the safest place to start: the application is unchanged, and the worst case is an ordinary startup. It does not make a JVM start in milliseconds. If a service must scale from zero within a request's latency budget, compare CRaC, which restores a whole warmed process from a checkpoint at the cost of managing live state across the checkpoint, and native images, which compile ahead of time under a closed-world assumption. Adopt Leyden first, measure what remains, and move to one of those only for the services where the remaining startup still matters.

What to do next

  1. Pick one service with a slow, class-loading-heavy startup and measure its startup, warmup and memory with a repeatable harness.
  2. Move it to JDK 25 or later if it is not there already.
  3. Write a deterministic training run against fixtures, starting with boot-only and adding smoke requests.
  4. Use the framework path: extracted layout and onRefresh for Spring Boot, the aot property and integration tests for Quarkus.
  5. Build the cache in the same pipeline run as the image, fingerprint it, and fail the build on a strict-mode start.
  6. Canary, compare against the baseline, and retune probes and autoscaler timings to the new startup.
  7. Alert when startup regresses to the old baseline, and rebuild caches on every JDK or agent change.
  8. Track the Leyden project page for the AOT code compilation JEP, and re-measure when you upgrade.
Key takeaway: Leyden's AOT cache is a compatible, low-risk startup and warmup improvement, but its value depends on operating it well. Measure first to confirm class loading and warmup dominate your startup, design a deterministic training run that exercises real paths without side effects, let Spring Boot or Quarkus drive it on JDK 25 or later, produce and strictly verify the cache in the same pipeline as its image, run production in the forgiving mode, and watch for the silent fallback that base-image, agent or class path drift causes across a fleet.