YourKit Java Profiler and JProfiler are the two long-standing commercial profilers for the JVM. Both attach an agent to your application, both offer CPU, memory, thread and lock views in a desktop UI, and both can save snapshots for later analysis. Teams usually buy one when the free tools have answered "which method is hot" but not "why is this request slow", "what is holding on to this memory" or "which call path allocates these objects".
Owning a licence is not the hard part; using it without misleading yourself is. This article explains how a profiler agent works, what sampling and instrumentation each measure and distort, how heap snapshots answer retention questions, and how to load each agent in development, in a container and unattended in CI. Option names below were checked against each vendor's current documentation on 2026-10-03. Where something was not confirmed, the article says so rather than guessing.
How a profiler agent works
A Java profiler of this kind is a native library loaded into the JVM with -agentpath at start-up, or attached later through the JVM's attach mechanism. Inside the process it uses JVMTI, the JVM Tool Interface, to do three things: ask for thread stack traces, intercept class loading so it can rewrite bytecode, and walk the heap. The agent keeps its data in the target process and serves it over a socket to the desktop UI, or writes it to a snapshot file. JProfiler's agent listens on port 8849 by default; YourKit's tries 10001 first and then the following ports.
This design has a consequence people forget: the agent runs inside your application. It shares your CPU, your heap headroom and your safepoints, so every measurement includes some of the profiler's own cost, and the amount depends on the mode you choose.
Sampling versus instrumentation
Sampling wakes up periodically, records the stack of each thread and counts how often each method appears. Its overhead is low and roughly independent of how many calls your code makes, so it is safe on realistic load. It gives proportions, not call counts: a method seen in 30% of samples used about 30% of the sampled time. Classic JVMTI stack sampling can only observe threads at safepoints, so time spent in tight loops between safepoints tends to be attributed to the next safepoint location. That is safepoint bias, and it is the main reason to cross-check with a tool like async-profiler that samples outside safepoints.
Instrumentation (YourKit calls it tracing) inserts timing code at method entry and exit. It gives exact invocation counts and call trees, which is what you need for questions such as "how many database calls does one request make". The cost scales with the number of calls, and it distorts what it measures: a three-line getter called ten million times gets the same probe as a slow method, the extra bytecode can stop the JIT from inlining, and small hot methods look far more expensive than they are. Use instrumentation on a narrowed set of packages for short periods.
YourKit exposes these as the cpu start-up option with the values sampling, tracing, counting and off. JProfiler calls the setting the call tree recording mode, and its offline agent accepts callTreeMode=sampling.
Memory: who allocates and who retains
Memory work splits into two questions. "Who allocates?" is answered by allocation recording, which samples or records allocation sites with their stacks. It is the tool for GC pressure: if young collections are frequent, allocation recording shows which code path creates the garbage. YourKit's alloc option controls this mode at start-up; JProfiler records allocations when you start allocation recording in the UI or with recording=cpu:allocation in offline mode.
"Who retains?" is answered by a heap snapshot. Both tools walk the object graph and compute retained size, the memory that would be freed if an object became unreachable, and both show the dominator tree and the shortest path from a GC root to any object. A leak investigation is almost always: take a snapshot, sort classes by retained size, open the biggest suspicious instance and follow its path to a GC root until you reach the static map, cache or listener list that should not be holding it. Comparing two snapshots taken minutes apart shows which classes grew. A heap snapshot pauses the application while the heap is walked and needs disk space of the same order as the live heap, which matters on a production node; heap tuning covers how to size the heap itself.
Loading the agent
For local development both products integrate with IDEs and launch the application with the agent already loaded. Everywhere else you add the agent path to the JVM options. Install paths vary by version and platform, so the paths below are placeholders.
# YourKit: listen on localhost only, start CPU sampling at start-up,
# keep snapshots in a known directory and write one when the JVM exits
java -agentpath:/opt/yourkit/bin/linux-x86-64/libyjpagent.so=port=10001,listen=localhost,cpu=sampling,dir=/var/tmp/yk,on_exit=snapshot \
-jar orders-service.jar
# JProfiler: listen on the default port, do not block start-up waiting for the UI
java -agentpath:/opt/jprofiler/bin/linux-x64/libjprofilerti.so=port=8849,nowait \
-jar orders-service.jarWithout nowait, JProfiler's agent holds the JVM at start-up until a UI connects, which is useful for profiling start-up and a nasty surprise in a service. For an already running JVM, JProfiler ships jpenable, which lists local JVMs, loads the agent into the one you pick and prints the port; run it as the same user as the target process. The YourKit UI can also attach to a running local JVM. Loading an agent late means classes that are already loaded must be retransformed before instrumentation sees them, and recent JDKs print a warning when an agent is loaded dynamically, so prefer start-up loading for anything you plan to profile regularly.
In containers, keep the agent listening on localhost (YourKit listen=localhost; JProfiler binds with address=) and reach it with kubectl port-forward or an SSH tunnel. Never expose a profiler port to a network: it gives whoever connects deep access to the process.
Unattended capture and CI
The most valuable mode for teams is unattended capture, because problems rarely wait for someone to open a UI. YourKit's start-up options include on_exit to save a snapshot when the JVM exits, used_mem to capture a memory snapshot when heap use crosses a percentage, and periodic_perf to save performance snapshots every given number of seconds, all into dir. JProfiler's offline mode runs without any UI connection, using either a saved session configuration or inline recording options:
# JProfiler offline mode with a session exported from the UI
-agentpath:/opt/jprofiler/bin/linux-x64/libjprofilerti.so=offline,id=1234,config=/etc/jprofiler/config.xml
# Or record CPU and allocations for ten minutes, then write a snapshot
-agentpath:/opt/jprofiler/bin/linux-x64/libjprofilerti.so=offline,snapshot=/var/tmp/run.jps,recording=cpu:allocation,duration=10mJProfiler also offers jpcontroller to start recordings and save snapshots in a running agent, jpdump for heap snapshots, and jpexport and jpcompare to turn snapshots into reports and compare two of them, which is how you build a performance regression check in CI. For programmatic control there is a Controller API (com.jprofiler.api.controller.Controller, shipped in bin/agent.jar and on Maven as com.jprofiler:jprofiler-probe-injected), and a Gradle plugin with a TestProfile task that runs tests under offline profiling. Check the exact flags of each tool against your installed version's help; they are not repeated here.
Worked example: a p99 regression
An orders service's p99 latency rose from 180 ms to 420 ms after a release, while average CPU only went up a little. The investigation took four steps.
- Run the old and new builds under the same load test with sampling enabled from start-up. Average CPU barely differs, so a whole-application hot spot is unlikely.
- Switch the views to wall time for the request threads. The new build spends a large share of request time waiting on a lock in a shared date formatter cache introduced in the release. Sampling showed proportions, which was enough to name the suspect.
- Enable allocation recording for one minute. The same code path allocates a new formatter on a cache miss and the cache is evicting constantly, so garbage and lock contention rise together.
- Narrow instrumentation to the cache package for a short run to get exact invocation counts: about 40 lookups per request instead of the expected one, from a loop that formats each line item. Hoist the formatter, make it immutable and thread-safe, and the p99 returns to its old value. Save both snapshots and compare them to confirm the change and document it.
Notice the order: low-overhead sampling to find where, allocation recording to find why, then short, narrow instrumentation to count. Starting with full instrumentation would have inflated every small method and hidden the lock wait under probe overhead.
Where they fit next to JFR and async-profiler
| Need | Good first tool | Why |
|---|---|---|
| Always-on production recording | JDK Flight Recorder | Built into the JDK, designed for continuous low overhead |
| CPU flame graph without safepoint bias | async-profiler | Samples outside safepoints, includes native frames |
| Leak hunting with retained sizes and GC-root paths | YourKit or JProfiler | Mature heap snapshot analysis and comparison |
| Exact call counts per request | YourKit or JProfiler instrumentation | Counts every call on a narrowed scope |
| Interactive exploration of threads, locks and probes | YourKit or JProfiler | Live UI views of many subsystems at once |
Failure modes
- Profiling the profiler. Full instrumentation on all packages makes getters the "hot spots". Sample first; instrument narrowly.
- Wrong clock. CPU time hides waiting; wall time includes it. A latency problem is a wall-time question.
- Unrepresentative load. Profiling an idle service measures JIT warm-up and start-up. Drive realistic traffic and discard the first minutes.
- Heap snapshot outage. Walking a large heap pauses the service and can fill the disk. Take it from a node removed from the load balancer.
- Exposed agent port. Bind to localhost and tunnel.
- Version mismatch. A snapshot or agent from one release may not open in a much older UI. Keep agent and UI versions aligned across the team.
- Blocking start-up. JProfiler without
nowaitwaits for a UI connection; a service that never becomes ready fails its health checks.
Trade-offs
Choosing between the two is mostly about workflow and licensing rather than capability: both cover CPU, memory, threads, locks and snapshot comparison. Trial both on a real problem from your own service, check how each fits your IDE, containers and CI, and compare licence terms for your team size directly with the vendors. Against the free tools, the trade is money and a heavier agent for deeper heap analysis, instrumentation and a richer UI. Many teams run JFR continuously, use async-profiler for flame graphs, and reach for a commercial profiler for leaks and call-count questions. Whatever you choose, read the GC logs alongside allocation data, because GC pressure is the most common reason memory profiling starts.
What to do next
- Install the trial of one product, profile a service you know under a load test, and find its top three methods by sampled CPU time and by wall time.
- Take two heap snapshots ten minutes apart under load, compare them, and follow the largest growing class to a GC root.
- Add agent options for unattended capture to a staging deployment: snapshot on exit and on a heap threshold (YourKit) or an offline recording with a duration (JProfiler).
- Bind every agent port to localhost and document the port-forward command your team uses.
- Build a CI job that profiles one benchmark offline and exports or compares snapshots between releases.
- Write down which questions you answer with JFR, async-profiler and the commercial profiler, so nobody starts a production investigation with full instrumentation.