Cassandra is built for writes that never wait for each other: every write carries a timestamp, replicas accept it independently, and the newest timestamp wins when values meet. That is why it scales, and it is also why a plain write cannot express "create this username only if nobody has it". Two clients can both read that the name is free, both write, and both be told they succeeded.

Lightweight transactions, LWTs, close that gap for a single partition. A statement with an IF clause runs a Paxos consensus round among the partition's replicas, so the condition check and the write happen as one linearizable step. This article explains what those rounds do, what they cost, how to choose consistency levels for them, how contention and timeouts behave, and how to operate them without turning a cluster designed for throughput into a queue.

Advertisement

What an LWT guarantees, and what it does not

An LWT is a compare-and-set on one partition. The forms are INSERT ... IF NOT EXISTS, UPDATE ... IF col = value (and other comparisons), UPDATE ... IF EXISTS and the matching DELETE forms. The result set contains an [applied] column; when it is false, the result also returns the current values of the columns in the condition, so the client can see why it lost without another read.

The guarantee is linearizability for operations on that partition that go through Paxos: there is a single order of those operations consistent with real time, and each condition is evaluated against the result of all operations before it. It is not a multi-partition transaction, it is not isolation for arbitrary reads, and it says nothing about plain writes that bypass Paxos.

Why QUORUM writes are not enough

The instinct is to read at QUORUM, check, and write at QUORUM. Walk through it with a replication factor of 3. Client X reads username ada at QUORUM: not found. Client Y does the same a millisecond later: not found. X writes ada to user 1 at QUORUM; Y writes ada to user 2 at QUORUM. Both succeed. Whichever write has the higher timestamp wins on every replica, silently, and one user has an account that points at a name they do not own. The quorum overlap guarantees each read sees the latest completed write; it does not stop two clients from acting on the same stale observation. See consistency levels for what quorums do guarantee.

What is missing is agreement on a single winner before either write becomes visible. That is exactly the problem consensus solves, and Cassandra solves it per partition with Paxos, whose general form is explained in the Paxos article.

Advertisement

Paxos v1, round by round

One lightweight transaction under Paxos v1: four round trips from coordinator to replicasClientINSERT ... IF NOT EXISTSCoordinatorpicks a ballotReplica 1system.paxosReplica 2system.paxosReplica 3system.paxos1.Prepare / promiseno higher ballot promised? report any accepted proposal2.Readread current row at the serial level to evaluate IF3.Propose / acceptsend ballot + mutation; a quorum must accept4.Commitapply the mutation at the commit consistency level[applied]true, or false + current rowPaxos v2 (4.1+, paxos_variant: v2): expect 2 round tripsfor a write and 1 or 2 for a read
The coordinator drives four exchanges with the partition's replicas. The first three wait for a quorum at the serial consistency level and the commit waits for the statement's own consistency level, so latency is roughly four replica round trips plus writes to the Paxos state table.

1. Prepare and promise. The coordinator picks a ballot, a time-based unique id, and asks the replicas to promise not to accept any proposal with a lower ballot. Each replica records the promise in its local system.paxos table and replies with any proposal it has accepted but not seen committed. If one exists, the coordinator must finish that earlier proposal first; this is how a transaction interrupted by a coordinator crash is completed by the next one.

2. Read. With the promise held by a quorum, the coordinator reads the current row to evaluate the IF condition. If the condition is false, the operation returns [applied] = false with the current values.

3. Propose and accept. If the condition holds, the coordinator sends the mutation with the ballot. A replica accepts unless it has since promised a higher ballot. With a quorum of accepts, the value is chosen.

4. Commit. The coordinator tells replicas to apply the mutation to the normal table, waiting for the commit consistency level. Only now is the write visible to ordinary reads.

The cassandra.yaml for 5.0 summarises the cost plainly: the default v1 variant should be expected to take four round trips for a write and three for a read at SERIAL. In one data centre at 1 ms per round trip that is a few milliseconds; across regions with SERIAL it is four WAN round trips, easily hundreds of milliseconds.

Paxos v2 in Cassandra 4.1 and later

Cassandra 4.1 added an optimised protocol, selected with paxos_variant: v2, which the configuration file describes as expecting two round trips for a write and one or two for a read, and marks as recommended. The default remains v1. The documented upgrade path is: make sure every node runs 4.1 or later, run nodetool repair --full -pr on each node, then set paxos_variant: v2 and do a rolling restart. The same file states that rolling back to v1 is safe and needs no data migration.

A further saving concerns the commit step. With any v2 variant and paxos_state_purging: repaired, the configuration notes that it is safe to use a commit consistency of ANY, which removes another wait from the write path. That setting depends on regular repairs that include Paxos state, so treat it as part of your repair plan, not a latency knob on its own; see repair architecture.

Variants whose names end in without_linearizable_reads or without_linearizable_reads_or_rejected_writes buy fewer round trips by weakening guarantees. Use them only if you have written down which guarantee you are giving up and why the application does not need it.

Two consistency levels on one statement

Every LWT carries two levels. The serial consistency level governs the Paxos phases and is either SERIAL, a quorum of all replicas in all data centres, or LOCAL_SERIAL, a quorum of replicas in the coordinator's data centre. The ordinary consistency level on the statement governs the commit phase, typically QUORUM or LOCAL_QUORUM.

LOCAL_SERIAL is much cheaper across regions, but it is only linearizable among clients that use the same data centre for that partition. If two regions both run LOCAL_SERIAL conditions on the same key, each region runs its own consensus and both can win. Use it when each partition has a home region that serves all of its conditional writes; otherwise pay for SERIAL.

Reads matter too. A read at QUORUM after an LWT may not see a proposal that was accepted but not yet committed. A read at SERIAL or LOCAL_SERIAL runs through Paxos and first finishes any in-progress proposal, so it returns the linearizable value. Use serial reads when the next decision depends on the answer.

Worked example: usernames, versions and leases

-- Unique username: first writer wins
INSERT INTO users_by_name (username, user_id, created_at)
VALUES ('ada', 5b6962dd-3f90-4c93-8f61-eabfa4a803e2, toTimestamp(now()))
IF NOT EXISTS;

-- Optimistic update with a version column
UPDATE accounts SET balance = 70, version = 8
WHERE account_id = 'acc-1'
IF version = 7;

-- Lease with fencing: take it only if no owner is currently set
UPDATE leases USING TTL 30 SET owner = 'worker-3', token = 42
WHERE resource = 'shard-17'
IF owner = null;

The username claim is the canonical case: rare writes, a strict uniqueness rule, one partition per name. The version-column update is optimistic concurrency: read the row, compute the new state, write it only if nobody changed the version in between, and on [applied] = false re-read and retry. The lease pattern gives a worker exclusive ownership of a resource for thirty seconds; the token is a fencing token, taken from a source that only increases (never reset when the lease expires), which downstream systems check, so a worker paused past its lease cannot overwrite a newer owner's work.

from cassandra import ConsistencyLevel, WriteTimeout
from cassandra.query import SimpleStatement

claim = SimpleStatement(
    "INSERT INTO users_by_name (username, user_id) VALUES (%s, %s) IF NOT EXISTS",
    consistency_level=ConsistencyLevel.LOCAL_QUORUM,          # commit phase
    serial_consistency_level=ConsistencyLevel.LOCAL_SERIAL,   # Paxos phase
)
check = SimpleStatement(
    "SELECT user_id FROM users_by_name WHERE username = %s",
    consistency_level=ConsistencyLevel.LOCAL_SERIAL,          # linearizable read
)

def register(session, username, user_id):
    try:
        rs = session.execute(claim, (username, user_id))
        return rs.was_applied            # False: someone else owns the name
    except WriteTimeout:
        # Outcome unknown: the proposal may or may not have been accepted.
        row = session.execute(check, (username,)).one()
        return row is not None and row.user_id == user_id

The Python driver exposes the outcome as was_applied on the result set. The except branch is the part most code omits, and the next section explains why it is required.

A timeout means unknown, not failed

When an LWT times out, the driver reports a write timeout whose write type is CAS. That tells you the coordinator did not get the responses it needed in time. It does not tell you whether a quorum accepted the proposal. If it did, the next Paxos operation on that partition, by any client, will find the accepted proposal during its prepare phase and commit it. So a timed-out claim may become a successful claim a moment later.

Never treat that timeout as a clean failure and never retry blindly. Resolve it with a serial read, as the code does, or make the write idempotent by including a request id in the value and checking for your own id. Retrying an IF NOT EXISTS that actually succeeded will return [applied] = false with your own row, which looks like a loss unless you compare ids.

Contention and hot partitions

Paxos serialises operations on a partition. When clients compete, a coordinator whose ballot is superseded backs off and retries with a higher ballot, for up to cas_contention_timeout, 1,000 ms by default. Under sustained contention, latency climbs and throughput on that partition can fall below what one client alone would achieve, because rivals keep invalidating each other's rounds.

The design rule follows: LWT partitions must be naturally low-contention. A per-user or per-order key is fine. A global counter, a single "next id" row or one lock row for a whole service is not. If you need many conditional updates to one logical thing, shard it, move the coordination into one process that owns the key, or use a different store.

Do not mix LWT and non-LWT writes on the same data

Paxos only orders the operations that go through Paxos. A plain UPDATE or DELETE on the same partition uses the ordinary timestamp path, can land between an LWT's read and its commit, and can win or lose against the LWT according to timestamps rather than the agreed order. The resulting state can violate the condition the LWT checked. Once a table uses conditions for correctness, route every write to those rows through conditions, including deletes, which have IF EXISTS and IF col = value forms. Apply the same caution to batches that touch those rows; the BATCH article explains what batches do and do not guarantee.

What to monitor

Metric (ClientRequest scope CASWrite / CASRead)What it tells you
LatencyEnd-to-end LWT latency; compare with plain writes to see the Paxos premium
ContentionHistogramHow often rounds collided; rising values point to hot partitions
ConditionNotMetConditions that were false; high values may be normal (uniqueness) or a retry storm
UnfinishedCommitIn-progress proposals completed by a later operation, often after timeouts
Timeouts, Unavailables, FailuresUnknown outcomes and quorum loss; each timeout needs client-side resolution

When to use something else

LWTs fit rare, correctness-critical decisions on one partition: uniqueness, leases, state-machine transitions, idempotency keys. They are the wrong tool for high-rate counters, for updates that must span partitions atomically, or for every write in a table "just to be safe". Cassandra 6.0 introduces Accord (CEP-15), a general-purpose transaction protocol for multi-partition transactions; as of mid-2026 it is in pre-release, so check its status and guarantees for your version before designing around it.

What to do next

  1. List every conditional statement in your code and the invariant each one protects; delete the ones that protect nothing.
  2. Check that every write to those tables, including deletes, uses a condition.
  3. Handle CAS write timeouts with a serial read or a request-id check, never a blind retry.
  4. Choose SERIAL or LOCAL_SERIAL per table based on whether each partition has a single home region.
  5. On 4.1 or later, plan the move to paxos_variant: v2 with the documented repair and rolling restart, and test it on staging.
  6. Dashboard CASWrite latency, ContentionHistogram, UnfinishedCommit and timeouts, and alert on contention growth.
Key takeaway: A Cassandra LWT is a per-partition compare-and-set implemented with Paxos: prepare, read, propose and commit in v1, about half the round trips with v2 in 4.1 and later. It gives linearizability only to operations that go through it, so every write to protected rows must be conditional. Pick SERIAL or LOCAL_SERIAL deliberately, treat a timeout as an unknown outcome resolved by a serial read, keep contended keys out of Paxos, and watch the CASWrite and CASRead metrics.