Locks make individual operations safe, but they do not compose. Moving an item from one thread-safe queue to another is not atomic just because each queue is, and taking both queues' locks invites deadlock unless every caller agrees on an order. Transactional memory (TM) offers a different contract: mark a block atomic, and the runtime makes it appear to run all at once, with no visible intermediate state, retrying it if it collides with another thread.

This article explains how that promise is kept. It covers the semantics, the design choices every TM makes, the TL2 software algorithm step by step with a working Python implementation, hardware TM and the fallback path it always needs, which hardware still offers it, and the operational rules that decide whether TM helps or hurts. It assumes familiarity with memory models and basic atomics.

What a transaction must guarantee

A transaction should give three guarantees. Atomicity: either all of its writes take effect or none do. Isolation: no other transaction sees its intermediate state. Opacity: even a transaction that will later abort never observes an inconsistent snapshot. Opacity is the guarantee people forget. Without it, a doomed transaction can read x from before a concurrent commit and y from after it, then divide by (y − x) = 0 or loop on a corrupted pointer before anyone gets the chance to abort it. These "zombie" transactions are why serious STMs validate as they read, not only at commit.

One more choice concerns code outside transactions. Under strong isolation, transactions are also isolated from plain non-transactional accesses. Under weak isolation they are not, so mixing the two on the same data is a race. Most software TMs provide weak isolation, which is why the rule "shared data is accessed only inside transactions" matters.

The design space

DecisionOption AOption B
Version managementeager: write in place, keep an undo loglazy: buffer writes, publish at commit
Conflict detectioneager: on each access (encounter-time locks)lazy: at commit (validation)
Granularityword or stripe (hash of address)object or field
Isolationweak (only between transactions)strong (also against plain accesses)
Progressblocking (locks at commit)obstruction- or lock-free (complex, rarer)

Eager versioning makes commits cheap and aborts expensive, because the undo log must be replayed. Lazy versioning does the opposite, and it adds a lookup on every read to see whether the transaction already wrote that location. Detecting conflicts late wastes work on doomed transactions but avoids aborting ones that would have finished. TL2 (Dice, Shalev and Shavit, 2006) is the reference design: lazy versioning, commit-time locking, per-read validation against a global clock, and stripe granularity. It is worth understanding because most later STMs are variations on it.

TL2 step by step

TL2 transaction lifecyclebeginrv = global clockread xversion(x) ≤ rv, unlockedwrite xbuffer in write setlock write settry-lock each stripewv = ++clockcommit timestampvalidate readsskip if wv == rv + 1write backvalues, version = wvrelease lockscommit visibleabortdrop buffers, release, retrystale readlock busyread changedretryWrites stay private until commit; reads are checked against rv as they happen, so a running transaction never sees an inconsistent state.
TL2: speculative reads validated against rv, buffered writes, and commit-time locking with read-set validation.

TL2 keeps a global version clock and, for each memory stripe, a versioned lock: a lock bit and the clock value of the last commit that wrote the stripe. A transaction proceeds in five steps:

  1. Begin. Read the global clock into rv, the read version.
  2. Read. If the location is in the write set, return the buffered value. Otherwise read the stripe version, then the value, then check that the stripe is unlocked, its version is unchanged, and the version is ≤ rv. Any failure aborts. Every value returned is consistent with the snapshot at time rv, which is what gives opacity.
  3. Write. Record the new value in a private write set. Nothing is visible yet.
  4. Commit. Try-lock every stripe in the write set, aborting if any is busy, so no thread can deadlock. Increment the clock to get wv. If wv ≠ rv + 1, someone else committed meanwhile, so re-check every read stripe: its version must still be ≤ rv and it must not be locked by another thread.
  5. Publish. Write the buffered values, set each stripe's version to wv, and release the locks.

Read-only transactions skip commit entirely, since every read was already validated against rv. That makes them cheap, a property shared with MVCC databases. The read check orders its steps deliberately. A writer stores the value and then the version while holding the lock, so a reader that saw a new value either still sees the lock held or sees a changed version.

A working STM and a concurrency test

A compact TL2 in Python, with one versioned lock per variable. The global interpreter lock means this is a teaching model, not a fast STM, but the protocol is the real one:

import threading

class Abort(Exception):
    pass

class TVar:
    __slots__ = ("value", "version", "lock")
    def __init__(self, value):
        self.value, self.version, self.lock = value, 0, threading.Lock()

class Clock:
    def __init__(self):
        self.v, self.m = 0, threading.Lock()
    def tick(self):
        with self.m:
            self.v += 1
            return self.v

CLOCK = Clock()

class Txn:
    def __init__(self):
        self.rv, self.reads, self.writes = CLOCK.v, set(), {}

    def read(self, tv):
        if tv in self.writes:
            return self.writes[tv]
        v1 = tv.version
        val = tv.value
        if tv.lock.locked() or tv.version != v1 or v1 > self.rv:
            raise Abort                   # never hand out an inconsistent value
        self.reads.add(tv)
        return val

    def write(self, tv, val):
        self.writes[tv] = val

    def commit(self):
        if not self.writes:
            return                        # read-only: already validated
        held = []
        try:
            for tv in sorted(self.writes, key=id):
                if not tv.lock.acquire(blocking=False):
                    raise Abort
                held.append(tv)
            wv = CLOCK.tick()
            if wv != self.rv + 1:
                for tv in self.reads:
                    if tv.version > self.rv or (tv.lock.locked() and tv not in self.writes):
                        raise Abort
            for tv, val in self.writes.items():
                tv.value, tv.version = val, wv
        finally:
            for tv in held:
                tv.lock.release()

def atomically(fn, max_retries=10_000):
    for _ in range(max_retries):
        tx = Txn()
        try:
            result = fn(tx)
            tx.commit()
            return result
        except Abort:
            pass                          # add backoff here under contention
    raise RuntimeError("transaction starved")

The test was eight threads each running 3,000 random transfers among 20 accounts of 100, while an auditor thread ran 500 read-only transactions summing every balance. The thread switch interval was forced down to one microsecond to provoke interleavings. Across three runs the final total was always 2,000, and the auditor never saw any other total, even mid-transfer. That is opacity working. Each run committed 24,500 transactions and aborted roughly 49,000 to 53,000 times, about two aborts per commit under this deliberately extreme contention. That ratio is the number to watch in production.

Hardware transactional memory

Hardware TM uses the cache as the transaction log. Lines read inside a transaction are tagged in the read set, written lines are held in the L1 data cache and kept invisible, and the coherence protocol detects conflicts: an invalidation or a remote read of a written line aborts the transaction and rolls back registers and memory. There are no per-access instrumentation costs, which is why short critical sections run fast. The same design makes HTM best-effort, though. A transaction can abort because it exceeded cache capacity or associativity, took an interrupt or page fault, made a system call, or hit an instruction the hardware does not support in transactions. So every HTM program needs a non-transactional fallback, usually a lock.

Intel's RTM interface shows the standard pattern, lock elision with a subscribed fallback lock (compile with -mrtm):

#include <immintrin.h>
#include <stdatomic.h>

static atomic_int fallback;               /* 0 = free, 1 = held */

void transfer(long *from, long *to, long amt) {
    for (int attempt = 0; attempt < 3; attempt++) {
        unsigned status = _xbegin();
        if (status == _XBEGIN_STARTED) {
            if (atomic_load_explicit(&fallback, memory_order_relaxed))
                _xabort(0xff);            /* lock held: must not run beside it */
            *from -= amt; *to += amt;
            _xend();
            return;
        }
        if ((status & _XABORT_EXPLICIT) && _XABORT_CODE(status) == 0xff) {
            while (atomic_load(&fallback)) _mm_pause();
            continue;
        }
        if (!(status & _XABORT_RETRY))    /* capacity etc.: retrying will not help */
            break;
    }
    while (atomic_exchange(&fallback, 1)) _mm_pause();
    *from -= amt; *to += amt;
    atomic_store(&fallback, 0);
}

Reading the fallback lock inside the transaction is the critical line. It puts the lock in the read set, so a thread taking the lock aborts every speculating thread, and speculative and locked execution never overlap.

Platform status, checked October 2026. Intel shipped a June 2021 microcode update that disables TSX on a range of Core and Xeon parts, and some distributions (Red Hat Enterprise Linux among them) now leave it off by default. Check the rtm CPU flag on the actual machine before relying on it. IBM removed transactional memory from Power ISA v3.1, so POWER10 has none. Arm specifies the Transactional Memory Extension (TME) as an optional Armv9 feature, but no shipping core implementing it could be confirmed as of October 2026, so treat it as unavailable until your target reports it. In practice, design portable code for software TM or locks, and treat HTM as an optional accelerator behind a feature check.

Where TM is used

TM is at its most successful in languages that control side effects. Haskell's STM monad provides atomically, retry (block until a read variable changes) and orElse (try an alternative), and its type system keeps I/O out of transactions. Clojure's refs, updated inside dosync, use multi-version concurrency. ZIO STM brings the same model to Scala. In C and C++, GCC implements __transaction_atomic blocks under -fgnu-tm with the libitm runtime, following the ISO Transactional Memory Technical Specification (ISO/IEC TS 19841:2015). That specification was not merged into the C++ standard, so treat those blocks as a compiler extension.

Operational guidance

  • No irrevocable actions inside transactions. I/O, system calls, and sending messages cannot be rolled back. Buffer them and perform them after commit.
  • Keep transactions short and narrow. Abort probability grows with read-set size and duration, and one hot counter that every transaction writes makes everything serialise.
  • Measure the abort ratio. Export commits, aborts, and abort reasons (conflict, capacity, explicit). Start with an alert on a sustained ratio around 1 and tune it from your own baseline: past that point most work is being thrown away. Find the hot variable as you would with lock contention diagnosis.
  • Back off and cap retries. Use randomised exponential backoff, and after N aborts fall back to an irrevocable or serialised mode so long transactions cannot starve.
  • Watch privatization. When a transaction unlinks a node and then plain code uses it, a doomed transaction may still be reading or writing it. Either keep all accesses transactional or use a runtime that guarantees quiescence.

Failure modes

  • Zombie transactions in STMs that validate only at commit: crashes and infinite loops on inconsistent reads.
  • Starvation: a long reader keeps losing to short writers. Fix with retry caps and an irrevocable mode.
  • False conflicts: two variables hashed to one stripe, or sharing a cache line under HTM, abort each other without a real conflict.
  • HTM that never commits: capacity aborts on every attempt mean every call pays for a failed transaction and then takes the lock anyway.
  • Mixed access: non-transactional writes to transactional data under weak isolation are data races.

Trade-offs

ApproachStrengthsWeaknesses
Lockspredictable, any side effect alloweddo not compose; deadlock; contention
Lock-free (CAS)progress guarantees; fast single-word updateshard to extend to multiple words
Software TMcomposes; portable; opacity possibleper-access overhead; aborts; no I/O
Hardware TMnear-zero overhead when it commitsbest-effort; needs fallback; patchy availability

What to do next

  1. Run the Python TL2 with the bank test, then remove the post-read lock check and keep rerunning until the auditor reports a wrong total; that failure is what opacity prevents.
  2. Add commit, abort, and abort-reason counters before using any TM in a service.
  3. Audit the transactional code paths for I/O and move it after commit.
  4. On x86, check grep rtm /proc/cpuinfo on your production hosts before planning around HTM.
  5. Try the same transfer in Haskell STM or ZIO STM to see retry and orElse compose.
  6. Compare against a striped-lock version under realistic contention, and keep whichever has the lower p99.
Key takeaway: Transactional memory lets atomic blocks compose by running them speculatively and retrying on conflict. TL2 achieves this in software with a global clock, buffered writes, per-read validation for opacity, and commit-time try-locks. Hardware TM is faster but best-effort and unevenly available, so it always needs a subscribed fallback lock. Keep I/O out, keep transactions small, and treat the abort ratio as your main health metric.