Locks make individual operations safe, but they do not compose. Moving an item from one thread-safe queue to another is not atomic just because each queue is, and taking both queues' locks invites deadlock unless every caller agrees on an order. Transactional memory (TM) offers a different contract: mark a block atomic, and the runtime makes it appear to run all at once, with no visible intermediate state, retrying it if it collides with another thread.
This article explains how that promise is kept. It covers the semantics, the design choices every TM makes, the TL2 software algorithm step by step with a working Python implementation, hardware TM and the fallback path it always needs, which hardware still offers it, and the operational rules that decide whether TM helps or hurts. It assumes familiarity with memory models and basic atomics.
What a transaction must guarantee
A transaction should give three guarantees. Atomicity: either all of its writes take effect or none do. Isolation: no other transaction sees its intermediate state. Opacity: even a transaction that will later abort never observes an inconsistent snapshot. Opacity is the guarantee people forget. Without it, a doomed transaction can read x from before a concurrent commit and y from after it, then divide by (y − x) = 0 or loop on a corrupted pointer before anyone gets the chance to abort it. These "zombie" transactions are why serious STMs validate as they read, not only at commit.
One more choice concerns code outside transactions. Under strong isolation, transactions are also isolated from plain non-transactional accesses. Under weak isolation they are not, so mixing the two on the same data is a race. Most software TMs provide weak isolation, which is why the rule "shared data is accessed only inside transactions" matters.
The design space
| Decision | Option A | Option B |
|---|---|---|
| Version management | eager: write in place, keep an undo log | lazy: buffer writes, publish at commit |
| Conflict detection | eager: on each access (encounter-time locks) | lazy: at commit (validation) |
| Granularity | word or stripe (hash of address) | object or field |
| Isolation | weak (only between transactions) | strong (also against plain accesses) |
| Progress | blocking (locks at commit) | obstruction- or lock-free (complex, rarer) |
Eager versioning makes commits cheap and aborts expensive, because the undo log must be replayed. Lazy versioning does the opposite, and it adds a lookup on every read to see whether the transaction already wrote that location. Detecting conflicts late wastes work on doomed transactions but avoids aborting ones that would have finished. TL2 (Dice, Shalev and Shavit, 2006) is the reference design: lazy versioning, commit-time locking, per-read validation against a global clock, and stripe granularity. It is worth understanding because most later STMs are variations on it.
TL2 step by step
TL2 keeps a global version clock and, for each memory stripe, a versioned lock: a lock bit and the clock value of the last commit that wrote the stripe. A transaction proceeds in five steps:
- Begin. Read the global clock into rv, the read version.
- Read. If the location is in the write set, return the buffered value. Otherwise read the stripe version, then the value, then check that the stripe is unlocked, its version is unchanged, and the version is ≤ rv. Any failure aborts. Every value returned is consistent with the snapshot at time rv, which is what gives opacity.
- Write. Record the new value in a private write set. Nothing is visible yet.
- Commit. Try-lock every stripe in the write set, aborting if any is busy, so no thread can deadlock. Increment the clock to get wv. If wv ≠ rv + 1, someone else committed meanwhile, so re-check every read stripe: its version must still be ≤ rv and it must not be locked by another thread.
- Publish. Write the buffered values, set each stripe's version to wv, and release the locks.
Read-only transactions skip commit entirely, since every read was already validated against rv. That makes them cheap, a property shared with MVCC databases. The read check orders its steps deliberately. A writer stores the value and then the version while holding the lock, so a reader that saw a new value either still sees the lock held or sees a changed version.
A working STM and a concurrency test
A compact TL2 in Python, with one versioned lock per variable. The global interpreter lock means this is a teaching model, not a fast STM, but the protocol is the real one:
import threading
class Abort(Exception):
pass
class TVar:
__slots__ = ("value", "version", "lock")
def __init__(self, value):
self.value, self.version, self.lock = value, 0, threading.Lock()
class Clock:
def __init__(self):
self.v, self.m = 0, threading.Lock()
def tick(self):
with self.m:
self.v += 1
return self.v
CLOCK = Clock()
class Txn:
def __init__(self):
self.rv, self.reads, self.writes = CLOCK.v, set(), {}
def read(self, tv):
if tv in self.writes:
return self.writes[tv]
v1 = tv.version
val = tv.value
if tv.lock.locked() or tv.version != v1 or v1 > self.rv:
raise Abort # never hand out an inconsistent value
self.reads.add(tv)
return val
def write(self, tv, val):
self.writes[tv] = val
def commit(self):
if not self.writes:
return # read-only: already validated
held = []
try:
for tv in sorted(self.writes, key=id):
if not tv.lock.acquire(blocking=False):
raise Abort
held.append(tv)
wv = CLOCK.tick()
if wv != self.rv + 1:
for tv in self.reads:
if tv.version > self.rv or (tv.lock.locked() and tv not in self.writes):
raise Abort
for tv, val in self.writes.items():
tv.value, tv.version = val, wv
finally:
for tv in held:
tv.lock.release()
def atomically(fn, max_retries=10_000):
for _ in range(max_retries):
tx = Txn()
try:
result = fn(tx)
tx.commit()
return result
except Abort:
pass # add backoff here under contention
raise RuntimeError("transaction starved")The test was eight threads each running 3,000 random transfers among 20 accounts of 100, while an auditor thread ran 500 read-only transactions summing every balance. The thread switch interval was forced down to one microsecond to provoke interleavings. Across three runs the final total was always 2,000, and the auditor never saw any other total, even mid-transfer. That is opacity working. Each run committed 24,500 transactions and aborted roughly 49,000 to 53,000 times, about two aborts per commit under this deliberately extreme contention. That ratio is the number to watch in production.
Hardware transactional memory
Hardware TM uses the cache as the transaction log. Lines read inside a transaction are tagged in the read set, written lines are held in the L1 data cache and kept invisible, and the coherence protocol detects conflicts: an invalidation or a remote read of a written line aborts the transaction and rolls back registers and memory. There are no per-access instrumentation costs, which is why short critical sections run fast. The same design makes HTM best-effort, though. A transaction can abort because it exceeded cache capacity or associativity, took an interrupt or page fault, made a system call, or hit an instruction the hardware does not support in transactions. So every HTM program needs a non-transactional fallback, usually a lock.
Intel's RTM interface shows the standard pattern, lock elision with a subscribed fallback lock (compile with -mrtm):
#include <immintrin.h>
#include <stdatomic.h>
static atomic_int fallback; /* 0 = free, 1 = held */
void transfer(long *from, long *to, long amt) {
for (int attempt = 0; attempt < 3; attempt++) {
unsigned status = _xbegin();
if (status == _XBEGIN_STARTED) {
if (atomic_load_explicit(&fallback, memory_order_relaxed))
_xabort(0xff); /* lock held: must not run beside it */
*from -= amt; *to += amt;
_xend();
return;
}
if ((status & _XABORT_EXPLICIT) && _XABORT_CODE(status) == 0xff) {
while (atomic_load(&fallback)) _mm_pause();
continue;
}
if (!(status & _XABORT_RETRY)) /* capacity etc.: retrying will not help */
break;
}
while (atomic_exchange(&fallback, 1)) _mm_pause();
*from -= amt; *to += amt;
atomic_store(&fallback, 0);
}Reading the fallback lock inside the transaction is the critical line. It puts the lock in the read set, so a thread taking the lock aborts every speculating thread, and speculative and locked execution never overlap.
Platform status, checked October 2026. Intel shipped a June 2021 microcode update that disables TSX on a range of Core and Xeon parts, and some distributions (Red Hat Enterprise Linux among them) now leave it off by default. Check the rtm CPU flag on the actual machine before relying on it. IBM removed transactional memory from Power ISA v3.1, so POWER10 has none. Arm specifies the Transactional Memory Extension (TME) as an optional Armv9 feature, but no shipping core implementing it could be confirmed as of October 2026, so treat it as unavailable until your target reports it. In practice, design portable code for software TM or locks, and treat HTM as an optional accelerator behind a feature check.
Where TM is used
TM is at its most successful in languages that control side effects. Haskell's STM monad provides atomically, retry (block until a read variable changes) and orElse (try an alternative), and its type system keeps I/O out of transactions. Clojure's refs, updated inside dosync, use multi-version concurrency. ZIO STM brings the same model to Scala. In C and C++, GCC implements __transaction_atomic blocks under -fgnu-tm with the libitm runtime, following the ISO Transactional Memory Technical Specification (ISO/IEC TS 19841:2015). That specification was not merged into the C++ standard, so treat those blocks as a compiler extension.
Operational guidance
- No irrevocable actions inside transactions. I/O, system calls, and sending messages cannot be rolled back. Buffer them and perform them after commit.
- Keep transactions short and narrow. Abort probability grows with read-set size and duration, and one hot counter that every transaction writes makes everything serialise.
- Measure the abort ratio. Export commits, aborts, and abort reasons (conflict, capacity, explicit). Start with an alert on a sustained ratio around 1 and tune it from your own baseline: past that point most work is being thrown away. Find the hot variable as you would with lock contention diagnosis.
- Back off and cap retries. Use randomised exponential backoff, and after N aborts fall back to an irrevocable or serialised mode so long transactions cannot starve.
- Watch privatization. When a transaction unlinks a node and then plain code uses it, a doomed transaction may still be reading or writing it. Either keep all accesses transactional or use a runtime that guarantees quiescence.
Failure modes
- Zombie transactions in STMs that validate only at commit: crashes and infinite loops on inconsistent reads.
- Starvation: a long reader keeps losing to short writers. Fix with retry caps and an irrevocable mode.
- False conflicts: two variables hashed to one stripe, or sharing a cache line under HTM, abort each other without a real conflict.
- HTM that never commits: capacity aborts on every attempt mean every call pays for a failed transaction and then takes the lock anyway.
- Mixed access: non-transactional writes to transactional data under weak isolation are data races.
Trade-offs
| Approach | Strengths | Weaknesses |
|---|---|---|
| Locks | predictable, any side effect allowed | do not compose; deadlock; contention |
| Lock-free (CAS) | progress guarantees; fast single-word updates | hard to extend to multiple words |
| Software TM | composes; portable; opacity possible | per-access overhead; aborts; no I/O |
| Hardware TM | near-zero overhead when it commits | best-effort; needs fallback; patchy availability |
What to do next
- Run the Python TL2 with the bank test, then remove the post-read lock check and keep rerunning until the auditor reports a wrong total; that failure is what opacity prevents.
- Add commit, abort, and abort-reason counters before using any TM in a service.
- Audit the transactional code paths for I/O and move it after commit.
- On x86, check
grep rtm /proc/cpuinfoon your production hosts before planning around HTM. - Try the same transfer in Haskell STM or ZIO STM to see
retryandorElsecompose. - Compare against a striped-lock version under realistic contention, and keep whichever has the lower p99.