The java.util.concurrent.atomic package updates single variables safely from many threads without a lock. AtomicInteger, AtomicLong, AtomicBoolean and AtomicReference wrap one value and expose operations such as increment and compare-and-set that complete as one indivisible step. Because no thread ever blocks, there is no lock to leak, no deadlock and no context switch.

This article explains the hardware primitive underneath, the full API including the memory-ordering variants added in JDK 9, the ABA problem as it actually applies to Java, LongAdder and when it beats AtomicLong, field updaters and VarHandle, a worked example that builds the metrics and admission control for a service, failure modes, and when a lock is still the right answer. For the language-neutral theory, read atomics and memory ordering first; this page is about the Java API.

Compare-and-set: the primitive underneath

Every atomic class rests on compare-and-set (CAS): "if this location still holds the value I expect, store my new value, and tell me whether it worked", done by the processor as one step. On x86 the JIT emits a lock cmpxchg instruction. On ARMv8 it uses the LSE atomic instructions where the CPU has them, or a load-exclusive/store-exclusive retry loop where it does not. Simple arithmetic such as getAndAdd usually compiles to a dedicated instruction (lock xadd on x86) instead of a CAS loop.

Anything more complex is a loop you write, or that updateAndGet writes for you:

// what a CAS loop looks like: read, compute, try, retry on interference
long prev, next;
do {
    prev = counter.get();
    next = Math.min(prev + 1, LIMIT);
} while (!counter.compareAndSet(prev, next));

If another thread changed the value between the read and the CAS, the CAS fails and the loop retries with the fresh value. Some thread always succeeds, which is what lock-free means: the system as a whole always makes progress, although one unlucky thread could in theory retry for a long time. The real cost under contention is not the instruction but the cache line holding the value, which must move exclusively to whichever core writes it next.

The API, including memory modes

The everyday operations, all on AtomicInteger and AtomicLong with equivalents on the other classes:

MethodEffectReturns
get() / set(v)read or write with volatile semanticsvalue / nothing
incrementAndGet(), getAndAdd(n)atomic arithmeticnew or old value
compareAndSet(e, v)store v only if current equals eboolean success
compareAndExchange(e, v)same, JDK 9+the value actually seen (the witness)
updateAndGet(f), getAndUpdate(f)CAS loop applying fnew or old value
accumulateAndGet(x, f)CAS loop applying f(current, x)new value

Functions passed to updateAndGet and accumulateAndGet may run several times under contention, so they must be pure: no logging, no I/O, no incrementing another counter.

JDK 9 added memory-ordering variants. getPlain/setPlain give no ordering; getOpaque/setOpaque guarantee only that the write eventually becomes visible; getAcquire/setRelease give one-way ordering suitable for publishing data; and lazySet is the older name for setRelease. The weakCompareAndSet family may fail spuriously and is meant for retry loops on weakly ordered CPUs. Use the plain get, set and compareAndSet unless a profiler tells you otherwise; the weaker modes are easy to misuse, and the rules are explained in the Java memory model.

There are array forms too: AtomicIntegerArray, AtomicLongArray and AtomicReferenceArray give each element atomic operations without one wrapper object per slot.

AtomicReference and immutable snapshots

AtomicReference is the most useful class in the package because it makes any immutable object atomically replaceable. Hold configuration or a statistics snapshot as an immutable record and swap the whole record:

record Limits(int maxInFlight, long timeoutMs, long version) {
    Limits withMax(int m) { return new Limits(m, timeoutMs, version + 1); }
}

final AtomicReference<Limits> limits = new AtomicReference<>(new Limits(64, 500, 0));

// writer: lock-free read-modify-write of several fields at once
limits.updateAndGet(l -> l.withMax(128));

// readers: one volatile read gives a consistent view of every field
Limits now = limits.get();

Readers never see half an update, because they see either the old record or the new one. The rule that keeps this correct is that the referenced object must never be mutated after publication.

One trap catches people every year. compareAndSet on a reference compares identity with ==, not equals. AtomicReference<Integer> with compareAndSet(1000, 1001) usually fails, because the autoboxed 1000 is a different object from the stored one; only values from -128 to 127 are cached by default. Use AtomicInteger for numbers, and compare against the exact object you read.

The ABA problem, as it applies to Java

CAS checks that the value equals what you expected, not that it never changed. If another thread changes A to B and back to A between your read and your CAS, your CAS succeeds. Whether that is a bug depends on what the value means.

In C and C++ the classic ABA failure is a lock-free stack whose popped node is freed and its memory reused for a new node at the same address. In Java, garbage collection prevents that particular failure for freshly allocated nodes: while your thread holds a reference to a node, that object cannot be reclaimed and reused, so a reference CAS against it is not fooled. ABA still bites Java code in two situations. The first is pooled or recycled nodes, where your own code reuses an object that another thread may still be holding. The second is values whose identity carries no history, such as an index, a version-less state code or a boxed small integer.

AtomicStampedReference pairs the reference with an int stamp that you increment on every change, and its CAS checks both, so A to B to A is detected. It allocates a small pair object on every successful update. AtomicMarkableReference carries a single boolean instead, which lock-free linked lists use to mark a node as logically deleted. Often the cheapest fix is to put a version field in your immutable record, as Limits above does.

LongAdder: scaling a hot counter

When many threads write one AtomicLong, every write needs the same cache line exclusively. The line bounces between cores and throughput stops scaling, then falls. LongAdder (JDK 8) spreads the count out.

One hot AtomicLong vs a LongAdder: same counter, different cache-line trafficAtomicLong.incrementAndGet()core 0incrementcore 1incrementcore 2incrementone cache linevalue (8 bytes)every increment needs the line exclusively:it bounces between cores and throughput flattensLongAdder.increment()core 0probe -> cellcore 1probe -> cellcore 2probe -> cellbaseused while uncontendedcell[0] (padded)cell[1] (padded)cell[2] (padded)sum() = base + all cellsnot an atomic snapshotCells are created only after a CAS on base fails, and the table grows (up to about the CPU count) when cells collide.Writes scale with cores; reads cost more. Choose by whether you write or read the value more often.
AtomicLong funnels every core through one cache line; LongAdder hashes threads onto padded cells and pays at read time instead.

Inside, LongAdder holds a base field and, lazily, an array of cells. Uncontended, it adds to base with one CAS. When a CAS fails, it creates the cell array; each thread is mapped to a cell using a per-thread probe hash, and on collision the probe is re-hashed and the array may double, up to roughly the number of CPUs. Threads are not given private cells; they share cells, just far fewer at a time. Cells are padded so neighbours do not share a cache line.

sum() adds base and every cell. It is not an atomic snapshot: increments that happen during the sum may or may not be counted. LongAccumulator generalises the idea to any function, for example new LongAccumulator(Long::max, Long.MIN_VALUE) for a running maximum, but because the order of combination is not guaranteed the function must give the same answer in any order. DoubleAdder and DoubleAccumulator exist too.

Rule of thumb: use LongAdder for counters that are written often and read rarely, such as metrics. Keep AtomicLong when every update needs the result, as with ID generators, or when you enforce a limit on the value.

Field updaters and VarHandle

Each atomic is a separate object, about 16 bytes for an AtomicInteger and 24 for an AtomicLong on a typical 64-bit JVM, plus a pointer to it. For a million cache entries each with an atomic hit count, that overhead adds up. Field updaters and VarHandle give atomic operations on a field of your own class instead.

final class Entry {
    volatile int hits;                       // must be volatile, must be int (not Integer)
    private static final VarHandle HITS;
    static {
        try {
            HITS = MethodHandles.lookup().findVarHandle(Entry.class, "hits", int.class);
        } catch (ReflectiveOperationException e) { throw new ExceptionInInitializerError(e); }
    }
    void hit() { HITS.getAndAdd(this, 1); }
}

AtomicIntegerFieldUpdater.newUpdater(Entry.class, "hits") does the same with an older, reflection-based API. VarHandle (JDK 9) is the modern choice: it covers fields and array elements and every memory-ordering mode, and the JIT optimises it well when the handle is held in a static final field. It is the same per-field CAS technique that ConcurrentHashMap and the other java.util.concurrent classes use internally. Use it in library code; in application code the wrapper objects are usually fine.

Worked example: metrics and admission control

Suppose a pricing service needs four pieces of shared state: a request counter for metrics, the slowest request seen, a sequence for request ids, and a cap of 200 requests in flight. Each needs a different tool:

final class ServiceStats {
    final LongAdder requests = new LongAdder();                               // written always, read on scrape
    final LongAccumulator maxLatencyMicros = new LongAccumulator(Long::max, 0L);
    final AtomicLong nextId = new AtomicLong();                               // caller needs each value
    final AtomicInteger inFlight = new AtomicInteger();
    static final int MAX_IN_FLIGHT = 200;

    boolean tryAdmit() {                                  // CAS loop: never exceeds the cap
        int cur;
        do {
            cur = inFlight.get();
            if (cur >= MAX_IN_FLIGHT) return false;
        } while (!inFlight.compareAndSet(cur, cur + 1));
        return true;
    }

    void finish(long micros) {
        inFlight.decrementAndGet();
        requests.increment();
        maxLatencyMicros.accumulate(micros);
    }
}

Note why tryAdmit is a loop rather than incrementAndGet followed by a check. With increment-then-check, 50 threads arriving together at 199 all increment, the counter overshoots to 249, and each must then decrement and reject; meanwhile other readers see a value above the cap. The CAS loop checks and claims in one step, so the counter never exceeds 200. The handler calls tryAdmit(), returns HTTP 503 if it fails, and calls finish in a finally.

The metrics scrape reads requests.sum() and maxLatencyMicros.getThenReset() every 15 seconds. Since finish makes three separate atomic updates, a scrape can see them half applied. That is acceptable for metrics. If you need several values to change together, put them in one immutable record behind an AtomicReference, or use a lock.

Failure modes

The common mistakes with atomics:

  • Check-then-act across two calls. if (a.get() < max) a.incrementAndGet() is a race. Put the check inside a CAS loop, as in tryAdmit.
  • Two atomics that must agree. Each is atomic on its own; together they are not. A balance in one AtomicLong and a count in another can be observed inconsistent.
  • Side effects in update functions. They re-run on contention, so a log line or a call inside appears twice.
  • Identity comparison on boxed values. AtomicReference<Long> CAS fails outside the small-value cache.
  • False sharing. Two hot atomic fields allocated next to each other share a cache line and slow each other down. @Contended is internal to the JDK and requires -XX:-RestrictContended for application classes, so for your own code prefer LongAdder or separate objects.
  • Unbounded spinning. A CAS loop around expensive work under heavy contention wastes CPU. If retries are frequent, the data wants a lock or sharding.

Choosing the right tool

SituationUse
One counter, many writers, rare readsLongAdder
Counter whose every result mattersAtomicLong
Flag, one-time initialisationAtomicBoolean with compareAndSet(false, true)
Several fields changed togetherimmutable record in AtomicReference
Invariant across several objectsa lock, see ReentrantLock
Visibility only, single writervolatile field, see volatile
Millions of per-object countersVarHandle on a field

Atomics win when the shared state is one variable or fits in one immutable object, and the update is short. Once an invariant spans several variables, a lock is simpler and usually just as fast.

What to do next

Apply this to your own code:

  1. Find synchronized blocks that protect a single counter or flag and replace them with atomics.
  2. Switch high-write metric counters from AtomicLong to LongAdder.
  3. Search for get() followed by set or an increment on the same atomic, and turn each into a CAS loop or an updateAndGet.
  4. Check every function passed to updateAndGet for side effects.
  5. Replace AtomicReference over boxed numbers with the primitive atomic class.
  6. Where several fields must change together, move them into an immutable record behind one reference.
  7. Benchmark with JMH at your real thread count before and after; contention behaviour does not show up in single-threaded tests.

Key takeaway: Atomics give lock-free, single-variable updates built on compare-and-set. Use CAS loops for check-and-claim logic, LongAdder for hot counters, AtomicReference over immutable records for multi-field state, keep update functions pure, and switch to a lock when an invariant spans several variables.