An atomic operation is an operation on shared memory that other threads can never observe half-done. That sounds like a small guarantee, but it is the foundation every mutex, reference count, lock-free queue and metrics counter is built on. Most programmers meet atomics as a fix for a racy counter and stop there, which leaves the important questions unanswered: which operations are atomic, what the processor actually does to make them so, why an atomic counter can be slower than a mutex under load, and when compare-and-swap is the right tool.
This article answers those questions in order. It starts with the lost update, defines exactly what atomicity promises and what it does not, tours the operations available in C++, Java, Rust and Go, explains the hardware mechanisms on x86 and Arm, writes correct compare-and-swap loops, measures the cost of contention with a worked counter example, and ends with failure modes and a checklist.
The lost update
The statement count++ on a plain integer compiles to three steps: load the value into a register, add one, store it back. Two threads can interleave those steps:
Thread A Thread B count in memory
load r = 41 41
load r = 41 41
add r = 42 41
add r = 42 41
store 42 42
store 42 42 <- one increment lostRun that across eight threads incrementing a million times each and the final total is usually well short of eight million, and different on every run. In C and C++ it is also undefined behaviour, which means the compiler is entitled to assume it never happens and optimise accordingly. The fix is to make the whole load-add-store one indivisible step: an atomic read-modify-write.
What atomicity guarantees, and what it does not
An atomic object promises two things. First, no torn values: a reader sees either the old value or the new one, never a mix of bytes from both (a real risk for 64-bit values on 32-bit platforms or for misaligned data). Second, read-modify-write operations such as fetch-and-add are indivisible: no other write to that object can land between the read and the write.
It does not promise three things people often assume. It does not make two atomic operations atomic together: checking if (balance.load() >= amount) and then calling balance.fetch_sub(amount) is still a race. It does not order surrounding non-atomic memory unless you ask for an ordering. And it does not make an operation fast. Ordering is a separate subject, covered in the memory model article; this article treats it only as far as choosing the right option.
The operation set across languages
Every mainstream language exposes the same small set. Learn it once and the APIs map directly:
| Operation | C++ | Java | Rust | Go |
|---|---|---|---|---|
| Load / store | load(), store() | get(), set() | load(), store() | Load(), Store() |
| Exchange | exchange() | getAndSet() | swap() | Swap() |
| Fetch-and-add | fetch_add() | getAndAdd() | fetch_add() | Add() (returns new) |
| Bitwise | fetch_and/or/xor() | getAndBitwiseOr() via VarHandle | fetch_and/or/xor() | And(), Or() (Go 1.23) |
| Compare-and-swap | compare_exchange_weak/strong() | compareAndSet() | compare_exchange(_weak)() | CompareAndSwap() |
The types are std::atomic<T> in C++ (plus std::atomic_ref<T> since C++20 for operating atomically on an ordinary object), AtomicInteger, AtomicLong, AtomicReference and VarHandle in Java, AtomicU64 and friends in Rust's std::sync::atomic, and the typed atomic.Int64 family in Go's sync/atomic since Go 1.19. Note the Go difference: Add returns the new value, while C++ fetch_add returns the old one. Off-by-one bugs in ticket counters often come from exactly this.
What the hardware does
Modern CPUs keep caches coherent with a protocol in the MESI family: a core may write a cache line only when it holds it exclusively, and acquiring exclusivity invalidates every other copy. An atomic read-modify-write works by acquiring the line exclusively and refusing to give it up until the read, modify and write are all done. No bus is locked for ordinary cacheable memory; the cache line is the unit of atomicity, which is also why the cost depends on who else wants that line. Coherence states and their performance consequences are covered in false sharing.
On x86, read-modify-write instructions take a lock prefix: lock xadd for fetch-and-add, lock cmpxchg for compare-and-swap, and xchg with memory is implicitly locked. Locked instructions on x86 are also full memory barriers, which is one reason sequentially consistent atomics are relatively cheap there.
Arm took a different route. Up to Armv8.0 the only tool is load-linked/store-conditional: ldxr loads and marks the address as monitored, and stxr stores only if nobody else wrote the line in between, reporting failure otherwise. A fetch-and-add is a small loop that retries until the store succeeds. Armv8.1 added the Large System Extensions (LSE) with single instructions such as ldadd, swp and cas, which scale much better under contention because the retry loop disappears and the operation can be performed closer to where the line lives. Whether you get LSE depends on compile flags: building for -march=armv8.1-a or later uses them directly, and GCC and Clang support -moutline-atomics, which selects LSE at run time when the CPU has it. On a large Arm server it is worth checking the disassembly of a hot atomic to see which one you got.
Compare-and-swap loops done right
Compare-and-swap (CAS) is the universal primitive: it writes a new value only if the current value equals what you expected, and tells you whether it did. Any read-modify-write that has no dedicated instruction, such as an atomic maximum, is built as a CAS loop:
#include <atomic>
void atomic_max(std::atomic<long>& target, long value) {
long seen = target.load(std::memory_order_relaxed);
while (seen < value &&
!target.compare_exchange_weak(seen, value,
std::memory_order_relaxed)) {
// on failure, compare_exchange_weak stored the current value in `seen`
}
}Three details matter. First, on failure the call writes the value it actually found into seen, so the loop does not need a separate reload. Second, the loop re-checks the condition each time: if another thread has already raised the maximum above value, the loop exits without writing. Third, the weak form may fail spuriously, reporting failure even though the value matched, because on LL/SC machines the store-conditional can fail for reasons such as an interrupt or a neighbouring write to the same line. Inside a loop that is harmless and lets the compiler emit a tighter sequence. Use compare_exchange_strong when you are not in a loop and a spurious failure would be wrong, for example a one-shot "claim this slot" attempt. Java's compareAndSet and Go's CompareAndSwap are strong; Rust offers both.
CAS compares values, not histories. If a value changes from A to B and back to A between your read and your CAS, the CAS succeeds even though the world changed. For counters that is fine; for pointers in a lock-free stack it is the ABA problem, which corrupts the structure. The standard remedies are tagged pointers and safe reclamation schemes, covered in hazard pointers.
The price of contention: a worked counter
Uncontended atomics are cheap: the line sits in one core's cache and the operation completes locally. Contended atomics are expensive in a way that surprises people, because every operation must pull the line across the chip. Throughput no longer scales with threads; past a handful of cores it often falls, since cores spend their time waiting for ownership.
Worked example: a request-handling service increments a global requests_total counter on every request across 32 worker threads. Profiling shows the increment, a single lock xadd, among the hottest instructions in the process. The fix is to stop sharing the line: give each thread (or each CPU) its own counter, padded to a cache line, and sum them only when someone reads the metric.
#include <atomic>
#include <cstdint>
#include <new>
constexpr std::size_t kLine = 64; // or std::hardware_destructive_interference_size
struct alignas(kLine) Shard { std::atomic<uint64_t> n{0}; };
class ShardedCounter {
static constexpr int kShards = 64;
Shard shards_[kShards];
public:
void add(uint64_t v, unsigned thread_index) {
shards_[thread_index % kShards].n.fetch_add(v, std::memory_order_relaxed);
}
uint64_t read() const { // approximate while writers run
uint64_t total = 0;
for (const auto& s : shards_) total += s.n.load(std::memory_order_relaxed);
return total;
}
};Writes now touch a line that, most of the time, only one core uses, so they stay local. The cost moves to the reader, which sums 64 values, a good trade for a metric read every few seconds. The read is not a snapshot: shards are summed at slightly different moments, which is acceptable for monitoring and unacceptable for, say, allocating unique IDs. Java ships this exact design as LongAdder, whose sum() is documented as not an atomic snapshot under concurrent updates. Go has no built-in equivalent, but the same struct-of-padded-shards pattern works.
The alignment is essential. Without alignas, several shards share one line and the contention returns under a different name, false sharing.
Is it actually lock-free?
An atomic type is only lock-free if the hardware can do the operation directly. For larger types the library may fall back to a hidden lock. In C++, std::atomic<T>::is_always_lock_free (C++17) answers at compile time and is_lock_free() at run time. Integers and pointers up to the machine word are lock-free on mainstream platforms. Sixteen-byte types, often wanted for a pointer plus a counter, are the grey zone: x86-64 has cmpxchg16b and Armv8 has paired exclusives, but whether your compiler and library actually use them, or route through libatomic, depends on the toolchain and flags, so check is_lock_free on yours rather than assuming. The type must also be trivially copyable, and before C++20 (which made compare-exchange ignore padding bits) comparison was bytewise, so on older toolchains padding bytes in a struct can make CAS fail forever when values look equal. C++20 added fetch_add for floating-point atomics and wait() / notify_one() so a thread can block on an atomic value instead of spinning.
Choosing a memory order
Every atomic operation takes an ordering argument, and the default (sequentially consistent) is always correct. Weaker orders are an optimisation with two common safe uses:
- Relaxed for values that are their own point, such as statistics counters and the sharded counter above. Atomicity holds; nothing else is ordered.
- Release on the write, acquire on the read for publishing data: everything the writer did before the release store is visible to a reader whose acquire load sees that store.
std::string payload; // plain data
std::atomic<bool> ready{false};
void producer() {
payload = build(); // 1
ready.store(true, std::memory_order_release); // 2 publishes 1
}
void consumer() {
while (!ready.load(std::memory_order_acquire)) {} // 3 pairs with 2
use(payload); // safe: sees 1
}With relaxed on either side, the consumer may see ready true and a partly built payload, a bug that x86 tends to hide and Arm tends to expose. If you are unsure, keep the default; the measurable gain from weaker orders is usually small next to the cost of contention. Spinning on the flag as above is acceptable only for very short waits; the spinlock article covers when to spin and when to sleep.
Failure modes
- Check-then-act across two operations. A load followed by a store is two atomics, not one. Use a single RMW or a CAS loop that re-checks.
- Mixing atomic and plain access to the same variable. One plain write reintroduces the data race. Make every access atomic, or use
atomic_refconsistently during the concurrent phase. - volatile as a substitute. In C and C++
volatilegives neither atomicity nor inter-thread ordering. (Java'svolatileis different: it gives visibility and ordering, butcount++on a volatile field is still not atomic.) - Relaxed publication. Using relaxed for a ready flag works on the developer's x86 laptop and fails intermittently on Arm servers.
- One hot line. A global counter, reference count or sequence number that every thread hammers. Shard it, batch updates, or restructure ownership.
- Return-value confusion. Old value versus new value differs across APIs, which produces off-by-one ticket and ID bugs.
- Overflow. Atomic unsigned arithmetic wraps silently; a 32-bit reference count can wrap under a leak.