A lightweight transaction (LWT) is Cassandra's compare-and-set: a write that applies only if a condition on the current row holds, decided by a Paxos round among the partition's replicas. It is the one tool in Cassandra that can say "only one of these concurrent writers wins". It is also several times more expensive than a normal write, serialised per partition and easy to misuse.

How the protocol works, why a timeout means the outcome is unknown and why LWT and plain writes must not be mixed on the same data are covered in lightweight transactions in depth. This article is the practical companion. It asks when an LWT is the right answer, how much it will cost before you build it, how to shape the schema, how to write client code that handles all three outcomes, and how to test the result. The worked example is a ticketing system that must never sell the same seat twice.

Advertisement

The question an LWT answers

Normal Cassandra writes are blind: the replica stores the cell with a timestamp, and the newest timestamp wins on read. That is ideal when the new value does not depend on the old one, such as recording a sensor reading or overwriting a profile field. It fails when correctness depends on the current state: "create this account only if the name is free", "sell this seat only if it is still held by this customer", "process this payment only if we have not processed this request id".

An LWT turns such a check-then-write into one linearizable step on a single partition. Every concurrent attempt is ordered, the condition is evaluated against the latest committed state, and the result tells the client whether its write applied. Two limits follow from that design. The condition and the write must touch one partition, and every LWT on that partition, whatever row it targets, competes for the same Paxos state.

A decision matrix

Use caseLWT?Why, or what instead
Unique claim (username, email, external id)YesRare writes, strict invariant, one partition per key
Idempotency key for a non-idempotent actionYesINSERT IF NOT EXISTS on the request id before acting
Reservation with a hold (seats, slots, stock units)Yes, if contention per item is modestOne partition per item; see the worked example
State machine transitions (order placed, paid, shipped)UsuallyCondition on the current state; reject illegal transitions
High-rate counters, likes, page viewsNoCounter columns or aggregation; LWT serialises every increment
Last-writer-wins updates (profile, settings)NoPlain writes; timestamps already resolve them
Append-only eventsNoUnique clustering keys such as a timeuuid make writes non-conflicting
Transfer between two partitionsNoOne LWT cannot span partitions; use a ledger with idempotent steps, or wait for Accord
Leader election for a servicePossibleA dedicated coordinator such as etcd or ZooKeeper is usually simpler to operate

A useful test: if two concurrent writes with different values would both be acceptable, you do not need an LWT. If accepting both would break an invariant someone will be paged about, you probably do.

Advertisement

Budgeting the cost before you build

An LWT costs several round trips between the coordinator and the replicas instead of one: four in the original Paxos implementation and roughly half that with Paxos v2, introduced in Cassandra 4.1. Because every LWT on a partition is serialised, the latency of one operation also caps the throughput of that partition. A back-of-envelope calculation before design review avoids unpleasant surprises:

# lwt_budget.py: an upper bound on successful LWTs per second against ONE partition.
# All LWTs on a partition are serialised through its Paxos state, so contention is per partition.
def max_lwt_per_partition(round_trips, rtt_ms, replica_ms):
    per_op_ms = round_trips * (rtt_ms + replica_ms)
    return 1000 / per_op_ms

# Illustrative numbers: replace with your measured coordinator-to-replica RTT and replica time.
for label, rt, rtt in [("Paxos v1, one DC", 4, 0.5), ("Paxos v2, one DC", 2, 0.5),
                       ("Paxos v1, SERIAL across DCs", 4, 40.0)]:
    print(f"{label:30} ~{max_lwt_per_partition(rt, rtt, 0.3):7.0f} ops/s ceiling per partition")
# Real throughput under contention is lower: competing proposers abort and retry each other.

The numbers are illustrative, but the shape is not. Inside one datacenter, a single partition can sustain hundreds to low thousands of uncontended LWTs per second. Using SERIAL across datacenters puts WAN round trips into every phase, and the ceiling falls by an order of magnitude or more. Contention reduces it further, because competing proposers invalidate each other's ballots and retry. If your design concentrates thousands of conditional writes per second on one key, such as a single global stock counter, the answer is a different design, not a bigger cluster.

Schema shapes that keep LWTs cheap

Three rules carry most of the weight. Give each contended item its own partition, because LWTs on different rows of the same partition still contend. Keep conditional tables narrow and separate from read models, so that wide displays are served by plain writes to another table. Write to a conditional table only through LWTs, deletes included, so ordering is never decided by timestamps behind Paxos's back.

-- One seat per partition: Paxos state is per partition, so seats must not share one.
CREATE TABLE seats (
    show_id    text,
    seat       text,
    state      text,          -- 'free' | 'held' | 'sold'
    holder     uuid,
    hold_until timestamp,
    PRIMARY KEY ((show_id, seat))
);

-- Read model for the seat map: plain writes, eventually consistent, separate table.
CREATE TABLE seat_map_by_show (
    show_id text, section text, seat text, state text,
    PRIMARY KEY ((show_id, section), seat)
);

-- Hold a free seat for ten minutes.
UPDATE seats SET state = 'held', holder = ?, hold_until = ?
 WHERE show_id = ? AND seat = ? IF state = 'free';

-- Take over a hold that has expired (conditions support comparisons).
UPDATE seats SET holder = ?, hold_until = ?
 WHERE show_id = ? AND seat = ? IF state = 'held' AND hold_until < ?;

-- Confirm the sale only while this customer still holds an unexpired hold.
UPDATE seats SET state = 'sold'
 WHERE show_id = ? AND seat = ? IF state = 'held' AND holder = ? AND hold_until >= ?;

Putting seat into the partition key means a show with 2,000 seats has 2,000 independent Paxos states, so two customers fighting over A7 never slow down a customer buying B12. The price is that listing a show's seats needs the second table, updated after each successful transition and allowed to lag by a moment. Conditions can compare as well as match, which lets one statement express "the hold has expired" without a sweeper job racing customers.

Worked example: holding a seat with the Java driver

Client handling of one conditional write: three outcomes, not twoApplicationUPDATE ... IF state = 'free'CoordinatorPaxos on one partitionReplica 1Replica 2Replica 3bound statement[applied] = truecondition held; write committed[applied] = falserow returns current valuesWriteTimeoutExceptionwrite type CAS: outcome unknownProceede.g. take paymentDecide from returned rowretry, offer another seatSERIAL read, then decidenever a blind retry
Figure 1. A conditional write has three outcomes. A timeout with write type CAS means the proposal may or may not commit, so the client must read at SERIAL before deciding.
import com.datastax.oss.driver.api.core.CqlSession;
import com.datastax.oss.driver.api.core.DefaultConsistencyLevel;
import com.datastax.oss.driver.api.core.cql.*;
import com.datastax.oss.driver.api.core.servererrors.DefaultWriteType;
import com.datastax.oss.driver.api.core.servererrors.WriteTimeoutException;
import java.time.*;
import java.util.UUID;

public final class SeatHolds {
    enum Outcome { HELD, TAKEN, UNKNOWN_RESOLVED_HELD, UNKNOWN_RESOLVED_LOST }

    private final CqlSession session;
    private final PreparedStatement holdFree, readSerial;

    SeatHolds(CqlSession session) {
        this.session = session;
        this.holdFree = session.prepare(
            "UPDATE seats SET state = 'held', holder = ?, hold_until = ? "
          + "WHERE show_id = ? AND seat = ? IF state = 'free'");
        this.readSerial = session.prepare(
            "SELECT state, holder FROM seats WHERE show_id = ? AND seat = ?");
    }

    Outcome hold(String show, String seat, UUID customer) {
        Instant until = Instant.now().plus(Duration.ofMinutes(10));
        BoundStatement stmt = holdFree.bind(customer, until, show, seat)
            .setConsistencyLevel(DefaultConsistencyLevel.LOCAL_QUORUM)        // applies once accepted
            .setSerialConsistencyLevel(DefaultConsistencyLevel.LOCAL_SERIAL)  // ballot agreement, this DC
            .setIdempotent(false);               // no speculative execution, no automatic retry
        try {
            ResultSet rs = session.execute(stmt);
            if (rs.wasApplied()) return Outcome.HELD;
            Row current = rs.one();              // [applied]=false row carries the current values
            return Outcome.TAKEN;                // current.getString("state") says why
        } catch (WriteTimeoutException e) {
            if (e.getWriteType() != DefaultWriteType.CAS) throw e;
            // Outcome unknown: the proposal may still commit. Read linearizably and decide.
            Row now = session.execute(readSerial.bind(show, seat)
                .setConsistencyLevel(DefaultConsistencyLevel.LOCAL_SERIAL)).one();
            boolean mine = now != null && "held".equals(now.getString("state"))
                        && customer.equals(now.getUuid("holder"));
            return mine ? Outcome.UNKNOWN_RESOLVED_HELD : Outcome.UNKNOWN_RESOLVED_LOST;
        }
    }
}

Several details matter. The statement sets two consistency levels: LOCAL_SERIAL for the Paxos phase and LOCAL_QUORUM for the commit. Use SERIAL instead only if a partition can be written from more than one datacenter and must be linearizable across them. The statement is marked non-idempotent so the driver never retries it or runs it speculatively: running it twice could turn a success into an apparent failure, since the second attempt would see the seat held, by us.

On wasApplied() == false the driver returns the current row, so the client can tell "someone else holds it" from "it is already sold" without another query, and can try the expired-hold statement if hold_until is in the past. On a timeout whose write type is CAS, the client reads the row at serial consistency, which completes any in-progress Paxos round, and decides from the result. In the ticketing flow, an unresolvable result shows the customer "checking availability" and retries the read, never the write.

Idempotency keys: the second workhorse

Many systems need LWTs for one thing only: making a non-idempotent action safe to retry. A payment, an email or an external API call should run once even if the client times out and resends. Claim the request id first:

CREATE TABLE request_claims (
    request_id uuid PRIMARY KEY,
    owner      text,
    status     text,          -- 'in_progress' | 'done'
    result     text
) WITH default_time_to_live = 604800;     -- keep claims for 7 days

INSERT INTO request_claims (request_id, owner, status)
VALUES (?, ?, 'in_progress') IF NOT EXISTS;

UPDATE request_claims SET status = 'done', result = ?
 WHERE request_id = ? IF owner = ? AND status = 'in_progress';

If the insert applies, this worker performs the action and records the result conditionally. If it does not, the returned row says whether another worker is still in progress or has already finished, in which case the stored result is returned to the caller. The time to live must exceed the longest window in which a client might retry; otherwise a late retry finds no claim and runs the action again.

Testing that the invariant holds

LWT bugs rarely show up in single-threaded tests, because the code is correct when nobody competes. Test the invariant directly with real concurrency against a real multi-node cluster, such as a three-node cluster in containers, not a mock:

// Concurrency test: 50 customers race for one seat; exactly one may win.
@Test
void exactlyOneHolderWins() throws Exception {
    resetSeat("show-42", "A7");                                   // state = 'free'
    ExecutorService pool = Executors.newFixedThreadPool(50);
    List<UUID> customers = Stream.generate(UUID::randomUUID).limit(50).toList();
    List<Future<SeatHolds.Outcome>> results = new ArrayList<>();
    for (UUID c : customers) results.add(pool.submit(() -> holds.hold("show-42", "A7", c)));

    long winners = 0;
    for (Future<SeatHolds.Outcome> f : results) {
        SeatHolds.Outcome o = f.get();
        if (o == SeatHolds.Outcome.HELD || o == SeatHolds.Outcome.UNKNOWN_RESOLVED_HELD) winners++;
    }
    assertEquals(1, winners);
    Row row = serialRead("show-42", "A7");
    assertTrue(customers.contains(row.getUuid("holder")));       // the stored holder is a real racer
}

Then add faults. Stop one replica during the race and confirm the test still finds exactly one winner. Pause the network to one node long enough to trigger write timeouts and confirm the timeout path resolves correctly. Run the race thousands of times in CI, since a one-in-ten-thousand interleaving is precisely what production finds. Jepsen-style histories, where every operation's invocation and result is logged and checked for linearizability afterwards, are the gold standard if the invariant carries money.

Operating LWT tables

The metrics to dashboard are listed in the mechanism article linked above. Two habits are specific to the patterns here. Keep conditional-only tables out of bulk loaders and repair-by-rewrite scripts, which issue plain writes, and say in the table comment that the table is conditional-only. When a hold or claim is unexpectedly slow, query tracing shows which Paxos phase took the time. The broader consistency trade-offs are in Cassandra consistency, with worked cases in tunable consistency examples.

Where Accord fits

Accord, the protocol proposed in CEP-15 and merged into trunk for Cassandra 6.0, is designed to make ACID transactions across several partitions practical. As of mid-2026, 6.0 was available only as an alpha, while the 5.0 line remained the production line. Until Accord reaches a release you are prepared to run, design as if each conditional write covers one partition. Keep cross-partition workflows as sequences of idempotent single-partition steps with a durable record of progress, so that adopting Accord later simplifies code rather than rescuing it.

Failure modes

SymptomCauseFix
Two seats sold to two customersSome path wrote the table without a conditionMake the table conditional-only; audit every writer
Customer told "failed" but holds the seatTimeout treated as failure, or blind retry saw its own writeResolve CAS timeouts with a SERIAL read; mark statements non-idempotent
Latency climbs during on-salesMany LWTs on one partitionOne partition per item; queue demand for the hottest items
Cross-DC latency on every holdSERIAL used where LOCAL_SERIAL would doHome each partition in one DC and use LOCAL_SERIAL
Duplicate payments after long outagesIdempotency claims expired before retries stoppedSet the TTL longer than the maximum retry window

What to do next

  1. List each place you plan to use an LWT and write down the invariant it protects; drop any without one.
  2. Run the budget calculation with your measured round-trip times and expected writes per key.
  3. Put each contended item in its own partition and move display data to a separate plain-write table.
  4. Set serial and regular consistency explicitly on every conditional statement, and mark it non-idempotent.
  5. Handle all three outcomes in code, resolving CAS timeouts with a SERIAL read.
  6. Add a concurrency test against a real multi-node cluster, then repeat it with a node stopped.
  7. Add an idempotency-key table in front of every external side effect that can be retried.
  8. Track Accord's release status, but design today as if each LWT covers one partition.
Key takeaway: Use a Cassandra LWT when correctness depends on the current state of one partition and a lost race would break an invariant: unique claims, idempotency keys, holds and state transitions. Budget it first, because each operation takes several round trips and every LWT on a partition is serialised. Give each contended item its own partition, keep conditional tables conditional-only, set both consistency levels explicitly, and handle applied, not applied and timed out as three different outcomes. Prove the invariant with concurrent tests that inject faults, and design cross-partition work as idempotent steps until Accord is production-ready.