The LMAX Disruptor is a Java library for passing events between threads. It was built at LMAX, a financial exchange, for an order-matching engine that had to process very high message rates with low and predictable latency. Martin Fowler's 2011 article on the LMAX architecture reports the business logic thread handling around six million orders per second; that is LMAX's claim about their system at the time, not a general benchmark, but it explains why the design attracted attention. Today the Disruptor sits inside logging frameworks, trading systems and other places where a queue between threads has become the bottleneck.
This article explains the design from first principles: why ordinary queues are slow under contention, what the ring buffer and sequences replace them with, exactly how a producer claims and publishes a slot, how consumers wait and batch, and how to build a dependency graph of handlers. A worked order pipeline in Java ties it together, followed by failure modes and the trade-offs that decide whether you should use it at all. The API notes are for the 4.x line, which requires Java 11.
Why a queue becomes the bottleneck
A conventional bounded queue such as ArrayBlockingQueue protects its head, tail and size with a lock. Every producer and every consumer contends for that lock, and each acquisition that is contended can cost a trip through the operating system. Lock-free linked queues avoid the lock but allocate a node per message, which creates garbage and scatters messages across memory, so the consumer misses the cache on each one.
There is a subtler cost too. Head and tail are written by different threads, and if they sit on the same cache line every write by one thread invalidates the other thread's copy. This is false sharing, and the false sharing article shows how it turns independent counters into a shared bottleneck. A queue has the worst possible shape for it: two hot variables, always written, by different cores.
The Disruptor addresses all three costs: no locks on the hot path, no allocation per message, and every counter on its own cache line with a single writer.
The pieces
| Component | Role |
|---|---|
| RingBuffer | Fixed array of pre-allocated event objects. Slots are reused forever; nothing is allocated per message. |
| Sequence | A padded long counter. Each consumer owns one and is its only writer. |
| Sequencer | Hands out sequence numbers to producers, tracks the published cursor, and checks gating sequences so producers never overwrite unread slots. |
| SequenceBarrier | What a consumer waits on: the cursor, plus the sequences of any consumers it depends on. |
| EventProcessor / EventHandler | The consumer thread loop, and your callback onEvent(event, sequence, endOfBatch). |
| WaitStrategy | How a consumer waits when no new events are available: block, sleep, yield or spin. |
The ring buffer and sequence numbers
The buffer size must be a power of two, so the slot for a sequence is computed with a mask instead of a division: index = sequence & (bufferSize - 1). Sequences are 64-bit and only ever increase; they never wrap in practice, because at a billion events per second a signed long lasts about 292 years. That removes the classic ambiguity of circular buffers about whether head equal to tail means full or empty. Full means the producer's next sequence minus the buffer size is greater than the slowest consumer's sequence; empty means a consumer's sequence equals the cursor.
All event objects are created at startup by an EventFactory. Producers do not hand over objects; they copy data into an existing slot. This is why events in Disruptor code are mutable classes, and why a handler must never keep a reference to an event after onEvent returns: the slot will be overwritten one lap later.
Claim and publish: the producer protocol
Publishing is two steps. The producer first claims a sequence, writes its data into that slot, then publishes the sequence so consumers may read it. With a single producer the claim needs no atomic operation at all, because only one thread advances the counter.
// Single-producer claim, simplified pseudocode of the idea (not the library source).
long nextValue = -1; // last sequence this producer claimed
long cachedGate = -1; // cached min(consumer sequences)
long next() {
long next = nextValue + 1;
long wrapPoint = next - bufferSize; // the slot we would overwrite
if (wrapPoint > cachedGate) { // only look at consumers when we might lap them
long min;
while (wrapPoint > (min = minimumSequence(gatingSequences, nextValue))) {
LockSupport.parkNanos(1); // ring full: wait for the slowest consumer
}
cachedGate = min;
}
nextValue = next;
return next;
}
void publish(long sequence) {
cursor.setRelease(sequence); // makes the slot's writes visible to readers
waitStrategy.signalAllWhenBlocking(); // wakes consumers only if they are blocked
}Two details matter. The producer caches the minimum consumer sequence and only re-reads the consumers' counters when it might be about to lap them, so in the common case a claim touches no shared memory. And the publish is a release store: every write to the slot's fields made before it becomes visible to any thread that reads the cursor with an acquire load. This is the happens-before edge that makes the protocol correct without locks; the Java memory model article explains release and acquire semantics in detail.
With multiple producers, claiming requires an atomic claim on the shared cursor, so producers contend there. They also complete out of order: producer A may claim 10 and producer B claim 11, and B may finish first. The cursor therefore cannot simply be the highest published sequence. The multi-producer sequencer keeps an extra array, one entry per slot, recording which lap each slot was last published for. Consumers scan it to find the highest contiguous published sequence. This is correct but more expensive, which is why choosing ProducerType.SINGLE when you genuinely have one publishing thread is one of the largest optimisations available.
Consumers, barriers and the batching effect
A consumer asks its barrier for the next sequence it wants. The barrier returns the highest sequence that is safe to read, which may be many slots ahead. The consumer then processes everything up to that point without touching shared state again.
// Event processor loop, simplified.
long nextSequence = mySequence.get() + 1;
while (running) {
long available = barrier.waitFor(nextSequence); // may return much more than nextSequence
while (nextSequence <= available) {
T event = ringBuffer.get(nextSequence);
handler.onEvent(event, nextSequence, nextSequence == available); // endOfBatch
nextSequence++;
}
mySequence.setRelease(available); // one store per batch, not per event
}This is the batching effect, and it is what makes the Disruptor degrade gracefully. When the consumer keeps up, batches are size one and latency is minimal. When it falls behind, batches grow, the per-event cost of synchronisation falls, and the consumer catches up faster. The endOfBatch flag lets a handler exploit this directly: a journaling handler can buffer writes and flush to disk once per batch rather than once per event.
Dependencies between consumers are expressed through barriers. A consumer that depends on others waits until all of their sequences have passed a slot, not only the cursor. This lets several handlers read the same event in parallel and a later handler run only when all of them are done, with no queue between stages and no copying.
Wait strategies
| Strategy | How it waits | Use when |
|---|---|---|
| BlockingWaitStrategy | Lock and condition variable | Default. CPU is shared and lock wake-up latency is acceptable. |
| TimeoutBlockingWaitStrategy | Blocking with a timeout | Handlers need periodic timeout callbacks, for example to flush. |
| SleepingWaitStrategy | Spin, then yield, then short sleeps | Background work such as asynchronous logging; low CPU, variable latency. |
| YieldingWaitStrategy | Spin, then Thread.yield() | Low latency with fewer consumer threads than cores. |
| BusySpinWaitStrategy | Pure spin | Lowest latency, only when each consumer has a dedicated core. |
| PhasedBackoffWaitStrategy | Spin, then yield, then a fallback strategy | Bursty load where you want spin latency without spinning forever. |
The wait strategy is the main dial between latency and CPU cost. A busy-spinning consumer burns a full core even when idle. On an oversubscribed machine, or in a container with a CPU quota, spinning threads compete with the threads that would produce work and latency gets worse, not better. Start with blocking, measure, and move to yielding or spinning only with pinned threads and spare cores.
Padding and false sharing in the implementation
Each Sequence is surrounded by unused padding fields so that the counter occupies a cache line by itself. Without this, two consumers' counters could share a line and every progress update by one would invalidate the other. The ring buffer itself is also padded at both ends of its array so neighbouring objects do not share lines with the first and last slots. The lock-free data structures article covers the broader family of techniques the Disruptor draws on.
Worked example: an order pipeline
The pipeline mirrors LMAX's original shape. Incoming orders must be journaled to disk and replicated to a standby before the matching engine acts on them, so that a crash loses nothing that was matched. Journaling and replication are independent and can run in parallel; matching depends on both.
import com.lmax.disruptor.BusySpinWaitStrategy;
import com.lmax.disruptor.EventHandler;
import com.lmax.disruptor.RingBuffer;
import com.lmax.disruptor.dsl.Disruptor;
import com.lmax.disruptor.dsl.ProducerType;
import com.lmax.disruptor.util.DaemonThreadFactory;
public final class OrderPipeline {
public static final class OrderEvent { // mutable, pre-allocated, reused
long orderId; long price; int qty;
void clear() { orderId = 0; price = 0; qty = 0; }
}
public static void main(String[] args) throws Exception {
Disruptor<OrderEvent> disruptor = new Disruptor<>(
OrderEvent::new, // EventFactory fills every slot up front
1 << 16, // 65,536 slots, must be a power of two
DaemonThreadFactory.INSTANCE, // 4.x takes a ThreadFactory, not an Executor
ProducerType.SINGLE, // one publishing thread: no contended claim
new BusySpinWaitStrategy()); // only with a dedicated core per consumer
EventHandler<OrderEvent> journal = (e, seq, endOfBatch) -> Journal.append(e, endOfBatch);
EventHandler<OrderEvent> replicate = (e, seq, endOfBatch) -> Replica.send(e, endOfBatch);
EventHandler<OrderEvent> matcher = (e, seq, endOfBatch) -> { Book.match(e); e.clear(); };
disruptor.handleEventsWith(journal, replicate) // run in parallel on the same events
.then(matcher); // runs only after both have seen seq
RingBuffer<OrderEvent> ring = disruptor.start();
// Publishing: the translator writes into the claimed slot, then publishes.
ring.publishEvent((e, seq, id, px, q) -> { e.orderId = id; e.price = px; e.qty = q; },
42L, 10_050L, 3);
disruptor.shutdown(); // waits until all published events are processed, then halts
}
}A few points are worth noting. The event is a plain mutable class and the matcher clears it, so no stale data leaks into the next lap. The journal handler uses endOfBatch to group disk writes. handleEventsWith(a, b).then(c) builds the diamond in the diagram without any queue between stages. The 4.x line removed the worker-pool API and handleEventsWithWorkerPool, so each handler sees every event; if you want work shared across threads, partition explicitly, for example by having handler i process only sequences where sequence % n == i.
Failure modes
- A slow consumer stalls everyone. The ring is bounded and the producer gates on the slowest last-stage consumer. A handler that blocks on a network call freezes ingestion. That is backpressure working as designed, but it must be expected: either keep handlers non-blocking or decide what the producer does when the ring is full, using
tryPublishEventto fail fast rather than wait. - A claimed sequence that is never published. If a producer using the low-level
next()andpublish()calls throws between them, that sequence is never published and every consumer waits forever. Always publish in afinallyblock, or use the translator methods, which do this for you. - Exceptions in handlers. The default exception handler logs and halts the processor. Set an explicit handler with
setDefaultExceptionHandlerand decide whether a bad event is skipped or fatal. - Retained references. Storing the event object in a map leads to data corruption one lap later. Copy the fields you need.
- Spinning in containers. Busy-spin strategies under a CPU quota cause throttling and multi-millisecond stalls.
- Wrong producer type. Declaring
SINGLEwhile two threads publish corrupts the sequence silently. DeclareMULTIunless one thread provably owns publishing.
Operating it and choosing it
Size the ring for bursts, not averages: the buffer must absorb the largest burst the slowest consumer cannot keep up with, and each slot is a fully allocated event, so memory is the slot size times the slot count. Export remainingCapacity() as a metric; a ring that is regularly near full means a consumer is under-provisioned. Pin latency-critical threads to cores, and use shutdown() for a graceful drain rather than halt(), which stops processors immediately.
The Disruptor is not a general replacement for queues. It shines with a fixed topology, in-process communication, high rates, and a team willing to reason about mutable reused events. It is a poor fit when consumers do slow I/O per event, when topology changes at runtime, or when throughput is modest and a BlockingQueue is simply easier to read. For flow control between asynchronous stages over networks, reactive streams backpressure is usually the better model, and for goroutine-style communication see channels.
What to do next
- Profile first: confirm that the queue between your threads, not the work itself, is the bottleneck.
- Model the topology: list producers, handlers and the dependencies between them, and draw the diamond.
- Pick the producer type honestly; use SINGLE only if one thread owns publishing.
- Start with BlockingWaitStrategy, measure latency percentiles, then try Yielding or BusySpin on pinned cores.
- Use translator-based publishing so claimed sequences are always published.
- Make events mutable, clear them in the last handler, and never retain references.
- Use endOfBatch to batch I/O, export remainingCapacity, and set an explicit exception handler before going live.