A lightweight transaction (LWT) in Cassandra is a compare-and-set: a write that applies only if a condition on the current data holds, evaluated and applied atomically with respect to every other LWT on the same partition. You write one by adding an IF clause to an INSERT, UPDATE or DELETE. Underneath, the coordinator runs a Paxos round among the partition's replicas, which is why LWTs cost several times a normal write.

The protocol, its round trips and the meaning of a timeout are covered in Cassandra LWT architecture, and whether to use an LWT at all is the subject of Cassandra LWT: when and how. This page is the practical companion: the full condition language, what comes back and how to act on it in code, and the one feature that turns LWTs from single-row checks into partition-wide invariants, conditional batches with static columns.

Advertisement

The unit of atomicity is a partition

Every LWT is scoped to one partition. Paxos runs per partition key, so two LWTs on different partitions do not see or serialise against each other, and no LWT can check a condition on one partition and write another. This is the single most important design constraint: the data your invariant covers must live in one partition. Inside that partition, though, the scope is wide. Conditions may name any row or column of the partition, including static columns, and a batch may combine several conditional statements on different rows, which are then checked and applied together.

A conditional batch: one partition, one Paxos round, all or nothingClientBATCH with IF clausesLOCAL_SERIALCoordinatorone partition keyPaxos roundPartition account a1Static rowbalance, versionentry 2026-10-01amount, memoentry 2026-10-03new, IF NOT EXISTSEvaluate every condition on the current partitionIF version = 42 on the static row AND IF NOT EXISTS on the new entryAll conditions trueevery statement appliedAny condition falsenothing appliedResult set[applied] plus current values
The coordinator evaluates every condition in the batch against the current state of one partition inside one Paxos round, and applies all statements or none. The client learns which by reading the [applied] column.

The condition language

The IF clause takes one of three shapes. INSERT accepts only IF NOT EXISTS. UPDATE and DELETE accept IF EXISTS, or one or more column conditions joined by AND. The column condition forms, checked against the CQL grammar in the Cassandra source, are these:

FormExampleNotes
ComparisonIF status = 'open', IF version < 9Operators =, !=, <, <=, >, >=
MembershipIF status IN ('open', 'held')Any listed value matches
Collection elementIF tags['owner'] = 'ana'Map key or list index
UDT fieldIF addr.city = 'Pune'A field of a UDT column
Collection contentsIF roles CONTAINS 'admin', IF prefs CONTAINS KEY 'tz'Conditional UPDATE and DELETE, Cassandra 4.1 and later (CASSANDRA-10537)
Null testIF owner = nullTrue when the column has no value

Several conditions joined with AND must all hold. There is no OR, and conditions compare against literals or bind markers, never against another column, so IF balance >= amount where both are columns is not expressible. You read, compute on the client and condition on what you read, which is the CAS loop below.

Two cases are easy to get wrong. IF EXISTS on an UPDATE means the row exists, which turns the normally upserting UPDATE into a true update. And a column condition on a row that does not exist is evaluated as if every column were null: IF version = 42 fails, while IF owner = null holds and, because UPDATE upserts, creates the row. That is how a lease is acquired on a key nobody has used yet; confirm it on your version with one cqlsh test before relying on it.

Advertisement

Reading the result

A conditional statement returns a result set, which an unconditional write does not. The first column is the boolean [applied]. When it is true the write happened, and usually nothing else is returned. When it is false, Cassandra also returns the current values it compared, so you can decide what to do without another read:

cqlsh> INSERT INTO users (username, email) VALUES ('ana', 'ana@example.com') IF NOT EXISTS;

 [applied] | username | email
-----------+----------+------------------
     False |      ana | ana@old.example

cqlsh> UPDATE docs SET body = 'v2', version = 8 WHERE id = 42 IF version = 7;

 [applied] | version
-----------+---------
     False |       9

For IF NOT EXISTS a false result carries the existing row; for column conditions it carries the current values of the columns in the conditions. Those values come from the serial read inside the Paxos round, so they are the linearizable truth at that instant, which is exactly what a retry needs. Drivers expose the flag directly: ResultSet.was_applied in the Python driver and wasApplied() in the Java driver. Do not parse column zero by name in application code.

A compare-and-set loop in code

The standard pattern is optimistic concurrency: keep a version column, read the row, compute the new value, write it on condition that the version has not changed, and retry on a conflict using the values the failed write returned. In Python:

from cassandra import ConsistencyLevel, WriteTimeout
from cassandra.query import SimpleStatement

UPD = SimpleStatement(
    "UPDATE docs SET body = %s, version = %s WHERE id = %s IF version = %s",
    consistency_level=ConsistencyLevel.LOCAL_QUORUM,
    serial_consistency_level=ConsistencyLevel.LOCAL_SERIAL)
READ = SimpleStatement("SELECT body, version FROM docs WHERE id = %s",
                       consistency_level=ConsistencyLevel.LOCAL_SERIAL)

def edit(session, doc_id, change, attempts=5):
    row = session.execute(READ, (doc_id,)).one()
    for _ in range(attempts):
        expected, new_body = row.version, change(row.body)
        try:
            rs = session.execute(UPD, (new_body, expected + 1, doc_id, expected))
        except WriteTimeout:
            # Outcome unknown: our write may or may not have committed. Re-read serially.
            row = session.execute(READ, (doc_id,)).one()
            if row.version == expected + 1 and row.body == new_body:
                return row.version          # it landed
            continue                        # it did not, or another writer won: recompute
        if rs.was_applied:
            return expected + 1
        # [applied]=False: rs.one().version is the winner's version. The change must be
        # recomputed from the winner's body, so fetch it with a serial read.
        row = session.execute(READ, (doc_id,)).one()
    raise RuntimeError(f"doc {doc_id}: gave up after {attempts} conflicts")

Three choices here matter. The read uses LOCAL_SERIAL so it sees any LWT that is in flight, rather than a possibly stale quorum read. The timeout branch treats a WriteTimeout as unknown, re-reads and checks whether its own version and body landed, instead of blindly retrying; and because the condition names the version it read, a blind retry of a write that did commit would simply fail its condition, never apply twice. When the returned values are all you need, a counter-like version or a single status, skip the re-read and retry from them directly. And the attempt limit is low: if five writers in a row beat you, the partition is too hot for compare-and-set, and adding retries only adds load.

Conditional batches: several rows, one decision

A BATCH may contain conditional statements, and then the whole batch becomes one LWT. Cassandra evaluates every condition in the batch against the current partition, and applies every statement only if all conditions hold. The rules are strict:

  • Every statement must target the same table and the same partition key. A conditional batch that spans partitions is rejected with Batch with conditions cannot span multiple partitions.
  • Conditions and plain statements may be mixed; the plain ones apply only if all conditions hold.
  • If any condition fails, nothing is applied, and the result returns [applied] = False with current values for the conditioned rows.
  • USING TIMESTAMP is not allowed on conditional statements, because Paxos assigns the timestamp.

This is what makes multi-row invariants possible: an insert of a child row together with a check of a header row, or a move of an item between two clustering rows of one partition, decided once. How unconditional batches behave, and why multi-partition logged batches are a different tool entirely, is in Cassandra batches.

Worked example: an account ledger in one partition

Suppose each account has a balance that must never go negative, plus an append-only list of entries. Put both in one partition: entries as clustering rows, balance and version as static columns, which belong to the partition rather than to any row.

CREATE TABLE ledger (
    account_id text,
    entry_id   uuid,
    balance    bigint STATIC,
    version    bigint STATIC,
    amount     bigint,
    memo       text,
    PRIMARY KEY (account_id, entry_id)
);

-- open the account: a static-only insert, once
INSERT INTO ledger (account_id, balance, version) VALUES ('a1', 10000, 0) IF NOT EXISTS;

-- debit 3000: read balance=10000, version=42 serially, check 10000 - 3000 >= 0 on the client
BEGIN BATCH
  UPDATE ledger SET balance = 7000, version = 43 WHERE account_id = 'a1' IF version = 42;
  INSERT INTO ledger (account_id, entry_id, amount, memo)
       VALUES ('a1', 5f0c1b9e-6a1d-4c39-9f53-2d8f1e4b7a10, -3000, 'invoice 881') IF NOT EXISTS;
APPLY BATCH;

Walk through what can happen. If nobody else wrote, both conditions hold: the balance and version change and the entry appears, together. If another debit committed first, version = 42 fails, nothing is written, and the result returns the new version, so the client re-reads and recomputes. If the client timed out on an earlier attempt that actually committed, the retry finds the entry id already present, IF NOT EXISTS fails, and the client knows the debit is done rather than applying it twice. The entry id is generated once per business operation, before the first attempt, which makes it the idempotency key.

Every account is its own partition, so different accounts never contend, and one hot account serialises only with itself. A transfer between two accounts cannot be one LWT, because it spans two partitions; that needs a saga with entries in both ledgers, or a different system.

Restrictions and consistency levels

  • Counters cannot be used with conditions, and counter tables cannot take LWTs.
  • Two consistency levels apply to every LWT: the serial consistency, SERIAL or LOCAL_SERIAL, for the Paxos phase, and the normal consistency for the commit. In one datacenter they behave the same; across datacenters LOCAL_SERIAL is cheaper but only linearizable within the local datacenter, so pick one per table and never mix them on the same data.
  • Mixing LWT and plain writes on the same columns breaks the guarantee, because the plain write is not ordered by Paxos and its client timestamp can win or lose arbitrarily.
  • Reads that must see the latest LWT result use SERIAL or LOCAL_SERIAL as their consistency level.
  • TTL is allowed on conditional writes, which is how leases are usually built; the expired value then reads as null, so IF owner = null sees an expired lease as free.

The interplay of the two levels with replication factor is in Cassandra consistency.

Failure modes and trade-offs

  • Contention: many writers on one partition fail each other in turn. Watch the CAS write contention and unfinished-commit metrics, and treat a high [applied] = False rate as a schema problem, not a retry problem.
  • Unknown outcome: a timeout on an LWT does not mean it failed. Without an idempotency key or version check, the retry can double-apply.
  • Hot static rows: the ledger pattern serialises every write to an account, so one very active account caps out at the LWT throughput of a single partition.
  • Growing partitions: append-only entries make the partition grow forever; bucket entries by month or archive them, while keeping the balance where the batch can condition on it.
  • Silent scope errors: a design that needs a condition across two partitions cannot be fixed with retries; redesign the key so the invariant lives in one partition.

The trade-off is simple to state. LWTs buy linearizable compare-and-set per partition at several times the latency of a plain write. They are worth it for invariants that cannot tolerate a lost update, uniqueness, versions, balances and leases, at modest write rates, and not for bulk or high-frequency writes.

What to do next

  1. List every invariant your application relies on and confirm each can be checked within one partition.
  2. Add a version column to rows edited by more than one writer and switch their updates to IF version = ? with a bounded retry loop.
  3. Generate an idempotency key per business operation and include it as an IF NOT EXISTS row in conditional batches.
  4. Use was_applied or wasApplied() and the returned current values instead of a second read after a conflict.
  5. Pick SERIAL or LOCAL_SERIAL per table, set it on every LWT and serial read, and never write the same columns without a condition.
  6. Track the conflict rate and LWT latency per table, and redesign any partition where conflicts exceed a few percent.
Key takeaway: A Cassandra LWT is compare-and-set on one partition: add IF NOT EXISTS, IF EXISTS or column conditions (comparisons, IN, collection elements, UDT fields, and CONTAINS from 4.1) and Paxos applies the write only if they hold. Read [applied] through the driver, and use the current values returned on failure to retry. A conditional batch checks and applies several rows of one partition together, which with static columns and an idempotency row gives safe balances and versions. Never use USING TIMESTAMP or counters with conditions, never mix plain writes on the same data, and treat a timeout as unknown.