Every garbage collector in HotSpot does two jobs: it hands out memory, and it takes back memory nobody can reach any more. Epsilon, added by JEP 318 in JDK 11, does only the first. It allocates objects as fast as the JVM can and never reclaims a single byte. When the heap is used up, the JVM throws OutOfMemoryError and, unless something catches it, the process ends.

For most services that would be a disaster, but a collector that never runs is a precise instrument. It gives a performance baseline with zero GC cost, turns a hidden allocation regression into a test failure, and suits processes that finish before filling their heap. This article covers how Epsilon allocates, its flags, JMH and test setups for each job, heap sizing, and what quietly stops working when nothing collects.

Advertisement

What Epsilon is, and what it refuses to do

Epsilon is enabled with two flags, because it is still marked experimental: -XX:+UnlockExperimentalVMOptions -XX:+UseEpsilonGC. It ships in standard OpenJDK builds from JDK 11 onward, so you do not need a special distribution. The JEP is explicit about scope: it implements a bounded allocation limit and the lowest possible latency overhead, at the expense of memory footprint and memory throughput. In other words, it has the cheapest allocation and the most expensive memory bill.

Three consequences follow directly. First, System.gc() does nothing, because there is no reclamation code to run. Second, the barrier set is empty: the small fragments of code that other collectors inject around reference reads and writes are simply absent. Third, the heap only grows. Committed memory is never uncommitted, so resident memory climbs monotonically toward -Xmx for the life of the process.

Epsilon is not a fast production collector or a cure for pauses in a long-running service: one allocating 100 MB per second fills a 32 GB heap in about five and a half minutes. It is a measuring tool, and a runtime for processes whose total allocation is known in advance.

How allocation works with no collector

Allocation in HotSpot is already mostly lock-free, and Epsilon keeps that machinery. Each thread owns a thread-local allocation buffer, or TLAB: a private slice of heap where an allocation is just a pointer bump and a bounds check. Only when a TLAB is exhausted does the thread go to the shared heap, where it carves a new TLAB with an atomic compare-and-swap on the heap top. Objects too big for a TLAB are allocated directly from the shared space the same way.

Epsilon's heap is one contiguous reserved range of -Xmx bytes. It starts with -Xms committed and commits more as the top crosses the committed edge, in steps of at least EpsilonMinHeapExpand. Once the top reaches the reserved limit, the allocation fails and the JVM raises the usual java.lang.OutOfMemoryError: Java heap space, so the familiar options -XX:+HeapDumpOnOutOfMemoryError, -XX:+ExitOnOutOfMemoryError and -XX:OnOutOfMemoryError=... all apply.

Thread Anew Order(...)Thread Bnew byte[4096]TLAB Abump pointer, no lockTLAB Bbump pointer, no lockShared heap topCAS to carve next TLABCommit moreEpsilonMinHeapExpand stepReserved limit (-Xmx)nothing left to carveOutOfMemoryErrorJava heap spaceOOM actionsheap dump, exit, hookallocateallocateTLAB fullTLAB fullcommitted fullreserved fullNo arrow ever points back: memory is handed out once and never reclaimed
The whole of Epsilon's allocation path. Threads bump pointers inside their own TLABs, refill TLABs from a shared top with a compare-and-swap, the heap commits more memory in steps up to -Xmx, and then the JVM throws OutOfMemoryError. There is no box for marking, sweeping or compacting because there is none.

Because nothing is reused, a thread that grabs a 4 MB TLAB and goes idle strands that memory forever. So Epsilon sizes TLABs elastically: they grow while a thread keeps allocating and decay back after it goes quiet.

Advertisement

The flags, and the ones that actually matter

All Epsilon tuning flags are experimental, so they also need -XX:+UnlockExperimentalVMOptions. Names and defaults below are taken from epsilon_globals.hpp in the current OpenJDK source; check java -XX:+UnlockExperimentalVMOptions -XX:+PrintFlagsFinal -version on your exact JDK before relying on any of them.

FlagDefaultWhat it controls
EpsilonPrintHeapSteps20How many heap occupancy reports to print over the heap's capacity; 0 turns them off. Visible with -Xlog:gc.
EpsilonUpdateCountersStep1MAllocation volume between updates of heap counters seen by monitoring tools. Higher is faster, lower resolution.
EpsilonMaxTLABSize4MLargest TLAB a thread can get. Larger is faster per allocation but wastes more per thread.
EpsilonElasticTLABtrueGrow TLABs for actively allocating threads instead of handing large ones to everyone.
EpsilonTLABElasticity1.1Multiplier for the next TLAB size while a thread keeps allocating.
EpsilonElasticTLABDecaytrueShrink a thread's TLAB size back after a quiet period.
EpsilonTLABDecayTime1000 msHow long a thread must be idle before its TLAB size decays to the initial value.
EpsilonMinHeapExpand128MSmallest step by which committed heap grows.

In practice you rarely touch these. The flags that matter are the standard ones: -Xmx sets the allocation budget, -Xms equal to -Xmx plus -XX:+AlwaysPreTouch moves page-fault cost to startup so it does not pollute measurements, and the OOM flags decide what happens at the end.

Barriers: the cost you pay even when no GC runs

To see why a GC-free baseline is useful, you need to know what a collector costs even when it is not collecting. Concurrent and generational collectors need to know when the application changes the object graph. They learn that through barriers: G1 adds a pre-write barrier for concurrent marking and a post-write barrier that records cross-region references in remembered sets; Parallel and Serial mark cards on reference stores; ZGC and Shenandoah add load barriers that check or fix references as they are read. Each is a few instructions, but on hot reference-storing loops they cost throughput and code size.

Epsilon's barrier set is empty. A run under it measures your code plus allocation, with no barriers, no concurrent GC threads competing for cores and no collection safepoints. The difference from the same run under G1 or ZGC is the total cost of memory management for that workload, which tells you whether GC tuning is worth an afternoon.

Use 1: a GC-free baseline in JMH

The cleanest place to use this is JMH. Add a fork configuration that runs the same benchmark under Epsilon with a heap large enough to hold every allocation the run will make:

@BenchmarkMode(Mode.AverageTime)
@OutputTimeUnit(TimeUnit.NANOSECONDS)
@Warmup(iterations = 5, time = 1)
@Measurement(iterations = 5, time = 1)
@State(Scope.Thread)
public class ParseBench {
    private byte[] payload;

    @Setup
    public void setup() throws IOException {
        payload = Files.readAllBytes(Path.of("order-sample.json"));
    }

    @Benchmark
    @Fork(value = 2, jvmArgsAppend = {"-XX:+UseG1GC", "-Xms8g", "-Xmx8g"})
    public Order g1() { return OrderParser.parse(payload); }

    @Benchmark
    @Fork(value = 2, jvmArgsAppend = {
        "-XX:+UnlockExperimentalVMOptions", "-XX:+UseEpsilonGC",
        "-Xms8g", "-Xmx8g", "-XX:+AlwaysPreTouch"})
    public Order epsilon() { return OrderParser.parse(payload); }
}

Run it with JMH's GC profiler, java -jar benchmarks.jar ParseBench -prof gc, and read gc.alloc.rate.norm, the bytes allocated per operation. Multiply that by the total number of operations across warmup and measurement to check the run fits in 8 GB. If it does not, the Epsilon fork dies with an OOM part way through, which is loud and easy to fix: shorten iterations or raise the heap.

If Epsilon is 3 percent faster, memory management is a small slice of this path; optimise the parser. If it is 30 percent faster, allocation or barrier cost dominates; allocate less or change collector. Epsilon's number is a ceiling, not a production target. Epsilon also never compacts, so a small gap can be object-layout noise rather than GC cost.

Use 2: allocation budgets as failing tests

The second honest use is an allocation budget. Many latency-sensitive paths are supposed to allocate little or nothing per request, and that property erodes silently: someone adds a lambda that captures, a boxed Long in a map key, or a log statement that builds a string even when the level is off. With a normal collector the regression shows up weeks later as slightly worse tail latency. With Epsilon it can be a test failure the same day.

The pattern is a forked test JVM with a deliberately small heap. Warm up, then run a known number of operations. If the code allocates more than budgeted, the heap runs out and the fork fails:

// build.gradle.kts: a separate test task that forks with Epsilon and a 256 MB heap
tasks.register<Test>("allocationBudgetTest") {
    testClassesDirs = sourceSets.test.get().output.classesDirs
    classpath = sourceSets.test.get().runtimeClasspath
    useJUnitPlatform { includeTags("alloc-budget") }
    maxHeapSize = "256m"
    minHeapSize = "256m"
    jvmArgs("-XX:+UnlockExperimentalVMOptions", "-XX:+UseEpsilonGC",
            "-XX:+ExitOnOutOfMemoryError")
    forkEvery = 1
}

// The test: 1,000,000 ticks must fit in what is left after startup.
@Tag("alloc-budget")
class TickPathBudgetTest {
    @Test
    void tickPathStaysWithinBudget() {
        var book = new OrderBook();
        var tick = new Tick();
        for (int i = 0; i < 50_000; i++) book.apply(tick.reset(i));   // warm up, let JIT settle
        for (int i = 0; i < 1_000_000; i++) book.apply(tick.reset(i)); // the budgeted run
    }
}

The arithmetic sets the budget. Suppose startup, JUnit and class loading use about 60 MB, measured once by logging heap usage with -Xlog:gc at the end of a passing run. That leaves roughly 196 MB. Spread over a million ticks it permits about 200 bytes per tick, which is a single small object. If the target is zero allocation, the run should use almost nothing beyond startup, and you can shrink the heap until the margin is tight. A thousand extra bytes per tick would need about 1 GB and fails immediately.

A complementary check works on any collector: com.sun.management.ThreadMXBean.getThreadAllocatedBytes read before and after the loop gives an exact figure for the failure message. Epsilon's check is cruder but also catches allocation on other threads.

Use 3: processes that finish before the heap does

The third use is a process whose total allocation is smaller than the memory you are willing to give it. Command-line tools, build steps, short batch transforms and serverless-style handlers that start, do one job and exit can run under Epsilon with no GC threads, no GC pauses and no barrier overhead. The JEP also mentions latency-critical applications that can be restarted before they exhaust the heap, but that is an operational commitment, not a flag.

Work the numbers before trying it. A report generator that allocates 1.5 GB in total over a 40 second run fits easily in a 4 GB heap. The same tool fed a file ten times larger allocates 15 GB and dies. So measure total allocation with -Xlog:gc under Epsilon across your largest realistic inputs, add a safety factor of two, and set -Xmx from that. If inputs are unbounded, Epsilon is the wrong choice; Serial or Parallel GC with a modest heap costs little for short runs and never dies of success.

For restart-before-exhaustion, replace input size with time: a gateway allocating 20 MB per second with a 64 GB heap lasts about 53 minutes, so it must be drained and cycled well inside that, automatically, with the allocation rate on a dashboard.

Failure modes: what depends on a collector running

The obvious failure is running out of heap. The less obvious failures come from everything in the JVM that quietly depends on a collector running:

  • Direct buffers and Cleaners. A DirectByteBuffer's native memory is freed by a Cleaner once the buffer becomes unreachable, and only a GC discovers unreachability. Under Epsilon that never happens. When -XX:MaxDirectMemorySize is hit, the JDK calls System.gc() and retries, which is a no-op here, so the process fails with an out-of-memory error for direct buffer memory. Netty-style pooled buffers and explicit release avoid it.
  • Weak and soft references. References are cleared only by a collector, so WeakHashMap entries and soft-reference caches are never cleared, and stale ThreadLocal entries are never expunged. A cache that relies on GC pressure to shrink simply grows.
  • Class unloading. Unloading classes requires a GC to prove a class loader unreachable. Applications that generate classes or redeploy plugins keep every class in metaspace forever, which is separate from -Xmx.
  • Container memory limits. Epsilon touches every heap page it ever uses and never returns any. If -Xmx plus metaspace, thread stacks and native memory exceeds the container limit, the kernel OOM-kills the process before Java can throw, and you get no heap dump. Keep -Xmx well under the limit.

Trade-offs against the collectors that reclaim

CollectorReclaims memoryBarriersTypical role
EpsilonNeverNoneBaselines, allocation tests, bounded short-lived runs
SerialYes, stop-the-world, one threadCard markingSmall heaps, single-core containers, short jobs
ParallelYes, stop-the-world, many threadsCard markingBatch throughput where pauses do not matter
G1Yes, mostly concurrent marking, pause-time goalPre- and post-writeDefault general-purpose server collector
ZGC / ShenandoahYes, concurrent compactionLoad and store barriersLarge heaps with very low pause targets

Epsilon sits at the far end: the application pays nothing for memory management and therefore must not need any. For a deeper comparison of the reclaiming collectors, see the G1 deep dive, the ZGC deep dive and Shenandoah. Heap regions, metaspace and native memory, which matter for the container sizing above, are covered in the JVM memory model article, and sizing a reclaiming heap is covered in heap tuning.

What to do next

  1. Check your JDK prints the Epsilon flags: java -XX:+UnlockExperimentalVMOptions -XX:+PrintFlagsFinal -version and search for Epsilon.
  2. Pick one hot path and add an Epsilon fork next to your production collector in its JMH benchmark; run with -prof gc and record gc.alloc.rate.norm.
  3. Compute the gap between the two forks. Under 5 percent, stop tuning GC for that path; over 20 percent, reduce allocation first, then revisit collector choice.
  4. Add an allocation-budget test task with Epsilon, a small fixed heap and -XX:+ExitOnOutOfMemoryError; calibrate the heap from one passing run and tighten it.
  5. Back it with a getThreadAllocatedBytes assertion so failures report a number, not just a crash.
  6. Before using Epsilon in any real process, measure total allocation on the largest realistic input, double it, and confirm it fits under -Xmx and the container limit.
  7. Audit that process for direct buffers, weak-reference caches and generated classes, and remove or bound each one.
  8. Never enable Epsilon on a long-running service unless restarts are automated, drained and alerted, and allocation rate is on a dashboard.
Key takeaway: Epsilon allocates and never reclaims, so a run under it shows your code's cost with zero memory-management overhead and turns any allocation beyond a budget into an OutOfMemoryError. Use it as a JMH baseline to decide whether GC tuning is worth doing, as a small-heap test that fails on allocation regressions, and only for processes whose total allocation is measured and bounded. Remember that System.gc, Cleaners, weak references and class unloading all stop working.