HotSpot's just-in-time compilers decide, method by method, what your Java code becomes on the CPU. They choose what to inline, which branches to treat as never taken and which allocations to remove. All of that is recorded only if you ask for it, as an XML log that is hard to read by hand. JITWatch, an open-source tool from the AdoptOpenJDK project, parses that log and shows its decisions next to your source, bytecode and generated assembly.
This article is about using the tool well. It covers how to produce a log JITWatch can trust, how to read its main views on a worked example, what the common inlining failure messages mean and what to change for each, and how to run JITWatch headless in a build so that a regression in compiled code fails CI rather than surprising you in production. The compiler mechanics themselves (tiers, profiling, speculation and deoptimisation) are covered in the JIT compiler in depth; here we assume them and focus on evidence.
What JITWatch reads, and what it does not
JITWatch is not a profiler and never attaches to a live process. It reads one file, the log written by -XX:+LogCompilation, which records each compilation task: the method, the tier, the bytecode size, every inlining decision with its reason, uncommon traps planted in the code, the resulting nmethod and, if hsdis is installed, the disassembly. JITWatch then loads your source and class files so it can line each decision up with a line of Java and a bytecode index.
Two consequences follow. First, JITWatch tells you what the compiler did, not whether it mattered, so pair it with a profiler that shows where time goes. Second, the log covers HotSpot's C1 and C2 compilers. If you run a different JIT, such as Graal through JVMCI or OpenJ9, check what your runtime logs before assuming JITWatch can read it.
Producing a log you can trust
The project documents three flags as the minimum: -XX:+UnlockDiagnosticVMOptions, -Xlog:class+load=info (so JITWatch can see which classes were loaded) and -XX:+LogCompilation. By default the file is named after the process id, for example hotspot_pid860.log; -XX:LogFile=name.log gives it a stable name. Assembly needs -XX:+PrintAssembly plus a disassembler: either a debug JVM build or the hsdis library built from OpenJDK source and placed where the JVM can load it. The project recommends adding -XX:+DebugNonSafepoints so the assembly can be mapped back to more source lines.
javac Shapes.java
java -XX:+UnlockDiagnosticVMOptions \
-Xlog:class+load=info \
-XX:+LogCompilation -XX:LogFile=shapes_jit.log \
-XX:+PrintAssembly -XX:+DebugNonSafepoints \
Shapes > shapes_stdout.txt
# Java 8 and older: replace -Xlog:class+load=info with -XX:+TraceClassLoading.
# Without an hsdis library on the JVM's library path, PrintAssembly prints a warning and no code.Three operational details decide whether the log is usable. The JVM writes each compiler thread's records to a temporary file and merges them into the final log when it exits normally, so a SIGKILL or out-of-memory kill leaves a log missing most of the compilation detail. Second, long runs with PrintAssembly can produce gigabytes, so use this on a benchmark, canary or reproduction, not every production host. Third, the log reflects the workload that ran. A two-second smoke test never reaches C2 for most methods, so drive the code with realistic data until the compile activity settles.
Worked example: a call site that will not inline
The program below sums areas over an array that holds three implementations of one interface. It is a small version of a common production shape: a hot loop over a collection of an abstract type.
// Shapes.java: a hot loop over an interface with three receivers, plus one oversized method.
interface Shape { double area(); }
record Circle(double r) implements Shape { public double area() { return Math.PI * r * r; } }
record Square(double s) implements Shape { public double area() { return s * s; } }
record Tri(double b, double h) implements Shape { public double area() { return 0.5 * b * h; } }
public class Shapes {
static double total(Shape[] shapes) {
double sum = 0;
for (Shape s : shapes) sum += s.area(); // three receiver types reach this call site
return sum;
}
public static void main(String[] args) {
Shape[] shapes = new Shape[10_000];
for (int i = 0; i < shapes.length; i++) {
shapes[i] = switch (i % 3) { case 0 -> new Circle(i); case 1 -> new Square(i); default -> new Tri(i, 2); };
}
double acc = 0;
for (int round = 0; round < 20_000; round++) acc += total(shapes);
System.out.println(acc); // keep the result alive
}
}Run it with the flags above, start JITWatch (mvn clean package && java -jar ui/target/jitwatch-ui-shaded.jar from a checkout; on JDK 11 and later Maven pulls JavaFX for you), open shapes_jit.log, add the directory holding Shapes.java as a source path and the directory with the class files as a class path, then start the analysis.
Select Shapes.total. Expect a C1 compilation followed by a C2 compilation at tier 4, and probably an on-stack replacement (OSR) compile of main, whose loop runs long enough to be replaced mid-execution. In the bytecode pane, expect the invokeinterface for area() to be annotated as not inlined, with a reason about its receiver profile (exact wording varies by JDK): three classes were seen, more than C2's type-profile-based inlining will handle, so the call stays virtual. The compiled loop therefore performs an interface dispatch per element and cannot vectorise or hoist anything across the call.
Fixes follow from the reason. A sealed interface with a switch over record patterns gives the compiler a static set of targets; one array per shape type removes the polymorphism altogether. Rerun, confirm in JITWatch that the annotation changed, then confirm with a benchmark that the time did too.
Reading TriView
TriView is the screen you will spend most time in: Java source on the left, bytecode in the middle and assembly on the right, linked so that selecting a line highlights its counterparts. The member list shows each method's compilations; choose the one you care about, usually the last tier 4 compile, because earlier ones may have been thrown away after deoptimisation.
Read in this order. Start with the source line you expected to be fast and find its bytecode. The bytecode annotations show call sites inlined or not, with the compiler's reason, plus branch statistics and any allocation that escape analysis removed. Only then open the assembly, and use it to answer narrow questions such as whether a bounds check is still in the loop, whether the loop was unrolled and whether a call instruction remains where you expected inlined code. Reading assembly first, without a question, wastes an afternoon.
For triage, the top lists rank methods by size, compile time or number of compilations; a method compiled many times is usually being deoptimised and recompiled. The compile chain shows the inlining tree for one compilation.
Inlining failure reasons and what each one asks of you
Inlining is the optimisation that enables the others, so its failure reasons are the most valuable lines in the log. The wording differs slightly between JDK versions; these are the common families.
| Reason (paraphrased) | What it means | What to do |
|---|---|---|
| callee is too large | A callee that is not hot exceeds MaxInlineSize (35 bytes of bytecode by default) | Usually nothing. If the call is on a hot path but the profile is not yet hot, extract the common case into a small method |
| hot method too big | A hot callee exceeds FreqInlineSize (325 bytes by default on x86-64) | Split the method so the hot path is small and the rare branches move to a separate method |
| already compiled into a big method | The callee already has compiled code larger than InlineSmallCode | Shrink the callee, or accept it; raising the threshold globally grows code everywhere |
| inlining too deep | The inline tree reached MaxInlineLevel | Flatten deep chains of tiny delegating wrappers on the hot path |
| not inlined because of the receiver profile (megamorphic) | More receiver types than the compiler will speculate on | Reduce types at that site, use a sealed hierarchy and a switch, or split the data by type |
| native method / no static binding | There is no bytecode to inline, or the target cannot be resolved statically | Move the work out of the hot loop or batch the native calls |
Resist tuning flags such as FreqInlineSize first. They change every method's decisions and grow the code cache; a source change that shrinks one method is local and survives JDK upgrades.
Uncommon traps, deoptimisation and recompilation
C2 compiles optimistically. It leaves out branches the profile says never run and plants an uncommon trap in their place, so if the branch is ever taken, execution falls back to the interpreter and the method is later recompiled. The log records each trap with a reason and an action. An abridged entry looks like this:
<uncommon_trap bci='14' reason='unstable_if' action='reinterpret' .../>A few traps after warm-up are normal. A method whose compilations keep alternating with traps is not: the workload's shape keeps changing, for example a rare type that appears in bursts. Find such methods by compilation count, read the trap reasons, and consider moving the rare case into a separate cold method.
Escape analysis, eliminated allocations and locks
When C2 proves an object never leaves a compiled method, it can replace the object with its fields in registers and remove the allocation; it can also remove locking on objects no other thread can see. JITWatch reports the allocations and locks it finds eliminated in the log, per method. This is the evidence you need before arguing about whether a short-lived iterator or builder costs anything: if the allocation is eliminated in the hot compile, it costs nothing there.
The reports also explain surprises. An allocation that is eliminated in a microbenchmark but survives in production usually escapes through a call that inlined in one context and not the other. See escape analysis in depth for the rules the compiler applies.
JarScan and headless mode: JIT checks in CI
Two tools in the JITWatch tree work without the GUI. JarScan, started with jarScan.sh, statically scans a jar and reports, as CSV, every method whose bytecode exceeds the hot-method inlining threshold (see the script for its other modes). Running it over your own code is a cheap way to find large methods before they cause inlining failures.
Headless mode runs the analyser on a log and writes text output: -e shows parse errors, -m the model, -c only compiled methods, -s the code suggestions, -t the compilation timeline and -f writes the output to headless.csv with a vertical bar as separator. That is enough to guard a performance-critical package across builds.
# ci_jit_check.sh: fail the build if a known-hot method gains a new inlining failure.
set -euo pipefail
java -XX:+UnlockDiagnosticVMOptions -Xlog:class+load=info \
-XX:+LogCompilation -XX:LogFile=build/jit.log \
-jar build/bench.jar --duration 60s
java -cp "$JITWATCH_CP" org.adoptopenjdk.jitwatch.launch.LaunchHeadless -s -f build/jit.log
mv headless.csv build/jit_suggestions.csv # vertical-bar separated
python3 tools/jit_diff.py baseline/jit_suggestions.csv build/jit_suggestions.csv \
--watch 'com.acme.pricing.'# tools/jit_diff.py: compare two headless suggestion exports and flag new findings in watched packages.
import sys
def load(path):
with open(path, encoding="utf-8") as f:
return {line.strip() for line in f if line.strip() and not line.startswith("sep=")}
baseline, current, watch = sys.argv[1], sys.argv[2], sys.argv[4] # argv[3] is --watch
new = sorted(r for r in load(current) - load(baseline) if watch in r)
for r in new:
print("NEW:", r)
sys.exit(1 if new else 0)Keep the guard narrow. Thread timing changes the profiles, so a whole-application diff is noisy. Watch only packages you have measured, and treat a new finding as a prompt to rerun a benchmark, not as proof of a regression.
Failure modes and traps
| Symptom | Cause | Fix |
|---|---|---|
| Log loads but most methods have no detail | JVM killed before it merged the per-thread logs | Stop the process normally; give containers time to exit |
| Assembly pane empty | No hsdis library, or PrintAssembly not set | Build hsdis for your JDK and platform; check the startup warning |
| Source pane blank or misaligned | Wrong source or class path, or sources from another build | Point JITWatch at the exact build that produced the log |
| Nothing reached C2 | Run too short or too little data | Drive realistic load until compile activity flattens |
| Conclusions do not survive production | Benchmark data differs from production | Capture logs on a canary with production traffic |
Where JITWatch fits among the other tools
-XX:+PrintCompilation and -XX:+PrintInlining are cheaper and good enough for a quick question about one method, but they print text with no source mapping. In practice the order is: profile first with async-profiler or JDK Flight Recorder to find the hot method, use JITWatch to explain why that method compiles as it does, change the code, and measure again. JITWatch is the explanation step, not the discovery step.
What to do next
- Pick one method your profiler ranks in the top five and capture a LogCompilation log from a realistic run that exits cleanly.
- Install hsdis for your JDK so TriView shows assembly, and add
-XX:+DebugNonSafepoints. - In TriView, list every call site in that method that failed to inline and note the reason for each.
- Fix the one with the clearest reason, such as a method above the hot inlining threshold, rerun, and confirm both the annotation and the benchmark changed.
- Check the compilation counts for methods that are compiled repeatedly and read their uncommon trap reasons.
- Run JarScan over your jars and review the largest methods on hot paths.
- Add a headless suggestions export for one critical package to CI, comparing it against a stored baseline.