Few CQL features are misunderstood as often as BATCH. Developers coming from relational databases see the keyword and assume it does what batching does elsewhere: groups many writes into one round trip so a bulk load finishes faster. In Cassandra that assumption is usually wrong. A batch that spans many partitions is generally slower than sending the same writes one by one in parallel, and it can put a single coordinator node under heavy load.

What batches are actually for is narrower and more valuable. A logged batch guarantees that a set of writes to different partitions will all eventually be applied, or none will, even if the client crashes halfway. A batch that targets a single partition is applied as one atomic, isolated mutation on each replica. This page is about using batches correctly from application code: the syntax, what each kind guarantees, driver code for the most common legitimate use, the timestamp rules that cause strange bugs, conditional batches, and the errors you will see when you get it wrong. For the coordinator and batchlog internals, read the companion Cassandra BATCH architecture page.

What a batch is, and what it is not

Start from what Cassandra already gives you without a batch. A single INSERT, UPDATE or DELETE that touches one partition is atomic at the partition level on each replica: all the cells it writes land together. What Cassandra does not give you is atomicity across partitions. If your application writes a row to users_by_id and a matching row to users_by_email, those live in different partitions, usually on different nodes, and a client crash between the two writes leaves one table without the other.

A batch addresses exactly that gap, and nothing else. It is a guarantee mechanism, not a performance mechanism. Keep three separate properties in mind, because the batch kinds differ on each:

  • Atomicity: either all statements eventually take effect or none do.
  • Isolation: whether a concurrent reader can observe some statements applied and others not.
  • Cost: how much extra work the cluster does compared with sending the statements individually.

A multi-partition logged batch gives atomicity without isolation, at a real cost. A single-partition batch gives both atomicity and isolation on each replica, cheaply. An unlogged multi-partition batch gives neither guarantee and only saves the client some round trips while moving the fan-out work onto one coordinator.

The CQL syntax

The grammar is small. A batch wraps modification statements; it cannot contain SELECT.

BEGIN [ UNLOGGED | COUNTER ] BATCH
    [ USING TIMESTAMP <microseconds> ]
    <insert | update | delete> ;
    <insert | update | delete> ;
    ...
APPLY BATCH;

With no modifier the batch is logged. UNLOGGED skips the batchlog. COUNTER is required when the batch contains counter updates, and a counter batch may contain only counter updates; you cannot mix counter and regular columns in one batch. A concrete example that keeps two lookup tables in step:

BEGIN BATCH
  INSERT INTO users_by_id    (user_id, email, name) VALUES (42, 'ana@example.com', 'Ana');
  INSERT INTO users_by_email (email, user_id, name) VALUES ('ana@example.com', 42, 'Ana');
APPLY BATCH;

In application code you almost always build batches from prepared statements through the driver rather than concatenating CQL strings. Prepared statements are parsed once, bind values safely, and let the driver route correctly. The batch kind and consistency level are set on the batch object, not on the statements inside it.

Batch kinds and their guarantees

ClientdriverCoordinatorreceives batchBatchlog node Asystem.batchesBatchlog node Bother rack1 batch2 log2 logReplicas of P1users_by_idReplicas of P2users_by_emailReplicas of P3audit3 apply3 apply3 apply4 remove log entry on successIf the coordinator dies after step 2, a batchlog node replays the batch later.Readers may see P1 updated before P2: atomic eventually, never isolated across partitions.
A logged multi-partition batch: the coordinator first persists the batch on two batchlog nodes, then applies each mutation to its own replicas, then deletes the log entry.
KindAtomicIsolatedExtra costUse it for
Logged, one partitionYesYes, per replicaNegligible: written as one mutationSeveral rows or columns in one partition that must change together
Logged, many partitionsYes, eventuallyNoBatchlog write to two nodes, then fan-out, then log removalKeeping denormalized tables or index tables consistent
Unlogged, one partitionYesYes, per replicaNoneSame as logged single-partition; the log is skipped anyway
Unlogged, many partitionsNoNoCoordinator fan-out and memoryRarely justified; usually replace with concurrent single writes
CounterNo replay guaranteePer partition onlyAs unloggedGrouping counter increments on one partition

The single-partition rows are the cheapest and most useful case, and many teams never realise it. When every statement in a batch shares the same partition key, Cassandra combines them into one mutation for that partition, so the replicas apply them atomically and in isolation, and no batchlog is needed. Clustering rows and static columns do not matter for this rule as long as the table and partition key are identical; rows in two different tables are always two different partitions.

Worked example: keeping lookup tables in step

The canonical legitimate use is a denormalized model where one entity is written to several query tables. Suppose a user can be looked up by id and by email, and an email change must move the user between partitions of users_by_email. Without a batch, a crash after deleting the old email row but before inserting the new one loses the user from the email table. With a logged batch, the batchlog guarantees the whole change completes.

from cassandra.cluster import Cluster
from cassandra.query import BatchStatement, BatchType
from cassandra import ConsistencyLevel

cluster = Cluster(["10.0.0.11", "10.0.0.12"])
session = cluster.connect("app")

upd_id   = session.prepare("UPDATE users_by_id SET email = ? WHERE user_id = ?")
del_mail = session.prepare("DELETE FROM users_by_email WHERE email = ?")
ins_mail = session.prepare(
    "INSERT INTO users_by_email (email, user_id, name) VALUES (?, ?, ?)")

def change_email(user_id, name, old_email, new_email):
    if old_email == new_email:
        return    # delete + insert of one row would tie on timestamp
    batch = BatchStatement(batch_type=BatchType.LOGGED,
                           consistency_level=ConsistencyLevel.LOCAL_QUORUM)
    batch.add(upd_id,   (new_email, user_id))
    batch.add(del_mail, (old_email,))
    batch.add(ins_mail, (new_email, user_id, name))
    session.execute(batch)

Walk through what happens. The driver sends one request to a coordinator. Because the batch touches three partitions, the coordinator first writes the serialized batch to the batchlog on two other nodes in the local data centre, preferring different racks. Only then does it send each mutation to its replicas at LOCAL_QUORUM. When all three succeed, it removes the batchlog entry. If the coordinator fails after logging, a batchlog node notices the entry has not been removed and replays it, so the change completes without the client doing anything.

Two properties of this code matter in production. First, every statement is idempotent: setting a column or inserting a row with the same values twice gives the same result, so a replay or a client retry is harmless. Second, the batch is small: three rows of a few hundred bytes. Both are deliberate. The early return for an unchanged email is deliberate too: deleting and re-inserting the same row inside one batch triggers the timestamp tie described in the next section. Batches that replay non-idempotent operations, or that carry megabytes, turn the guarantee into a liability.

Shared timestamps and the delete tie

Every cell in Cassandra carries a write timestamp, and conflicts are resolved by last-write-wins on that timestamp. The rule that surprises people: all statements in a batch get the same timestamp unless you set them explicitly. The coordinator, or the client if client-side timestamps are on, assigns one timestamp to the whole batch.

That makes the order of statements inside a batch meaningless. Consider a batch that deletes a row and then re-inserts it with new values, a common way to reset a row. Both the tombstone and the new cells carry the same timestamp, and when a tombstone and a live cell tie, the tombstone wins. The row stays deleted, and the application sees its reset silently fail.

-- Looks like "clear then rewrite". Actually leaves the row deleted.
BEGIN BATCH
  DELETE FROM sessions WHERE user_id = 42;
  INSERT INTO sessions (user_id, token) VALUES (42, 'new-token');
APPLY BATCH;

-- Fix: give the insert a strictly later timestamp (microseconds).
BEGIN BATCH
  DELETE FROM sessions USING TIMESTAMP 1759500000000000 WHERE user_id = 42;
  INSERT INTO sessions (user_id, token) VALUES (42, 'new-token')
    USING TIMESTAMP 1759500000000001;
APPLY BATCH;

A batch-level USING TIMESTAMP applies one explicit value to every statement; when statements need different timestamps, set them per statement and leave the batch level clear. Often the cleaner fix is to avoid the pattern entirely: overwrite the columns you care about instead of deleting and recreating. Tombstones also have their own costs, described in the tombstones guide.

Conditional batches

A batch may include lightweight-transaction conditions such as IF NOT EXISTS or IF balance = 100. The whole batch is then applied only if every condition holds, using Paxos to serialize it against other conditional writes. The important restriction is that a conditional batch must target a single partition; Cassandra rejects conditional batches that span partitions because Paxos runs per partition.

BEGIN BATCH
  INSERT INTO seats (show_id, seat, holder) VALUES ('s1', 'A7', 'ana') IF NOT EXISTS;
  UPDATE seats SET holder = 'ana' WHERE show_id = 's1' AND seat = 'A8' IF holder = null;
APPLY BATCH;

The result is a row whose [applied] column tells you whether the batch went through, plus the current values of the columns that failed their condition when it did not. Treat it like any compare-and-set: read the flag, and on failure decide whether to retry with fresh state or report a conflict. Conditional batches inherit the cost of Paxos, which takes several round trips between replicas, so use them for the few operations that truly need compare-and-set semantics. The trade-offs are covered in lightweight transactions.

Why batching for throughput backfires

The most damaging misuse is the bulk-load batch: a client collects a thousand unrelated rows and sends them as one unlogged batch to save round trips. Consider what the cluster does with it. One coordinator must hold the entire batch in memory, work out the replicas for each of the thousand partitions, send each mutation, and wait for all of them before replying. If the batch is logged, the coordinator also writes the whole payload to two batchlog nodes first.

The faster pattern is many small, concurrent, single-partition writes. A token-aware driver sends each one straight to a replica, which acts as coordinator for its own data, so the work spreads evenly and no single request carries a large payload.

from cassandra.concurrent import execute_concurrent_with_args

ins = session.prepare("INSERT INTO events (device_id, ts, value) VALUES (?, ?, ?)")
rows = [(d, ts, v) for (d, ts, v) in readings]       # many partitions
results = execute_concurrent_with_args(session, ins, rows, concurrency=64)
failed = [r for r in results if not r.success]

There is one exception where grouping helps: many rows that share a partition key, such as a batch of readings for one device in one time bucket. Those go into one mutation on one replica set, so a single-partition batch is efficient. Cassandra protects itself with configurable size thresholds. In Cassandra 4.1 and later the settings are batch_size_warn_threshold: 5KiB, batch_size_fail_threshold: 50KiB and unlogged_batch_across_partitions_warn_threshold: 10; older releases spell the first two batch_size_warn_threshold_in_kb and batch_size_fail_threshold_in_kb. Treat the warning in the server log as a code-review finding, not noise, and do not raise the fail threshold to make it go away.

Failure modes

  • Batch too large. A batch above the fail threshold is rejected with an invalid-request error, and one above the warn threshold logs a warning naming the keyspace and tables. The fix is to split the work, usually into concurrent single-partition writes, not to raise the limit.
  • WriteTimeout with write type BATCH_LOG. The coordinator could not persist the batchlog in time. The batch may not have been applied, and retrying an idempotent batch is safe.
  • WriteTimeout with write type BATCH. The batchlog was written but some mutations did not reach enough replicas in time. The batch will be replayed and will complete, so do not treat this as a failure to undo; a read immediately afterwards may simply not see every part yet.
  • Partial visibility. Readers of a multi-partition logged batch can see some tables updated before others. Design readers to tolerate this, or keep the data that must be read together in one partition.
  • Unlogged partial application. An unlogged multi-partition batch that times out may have applied any subset of its statements, and nothing will finish the rest. Only use it with idempotent statements you are prepared to retry.
  • Counter double-counting. Counter increments are not idempotent. A client retry after a timeout can apply an increment twice, so counter batches need application-level tolerance for small overcounts.

Trade-offs

ChoiceYou gainYou pay
Logged multi-partition batchCross-table atomicity with no client bookkeepingBatchlog writes on two nodes, higher latency, no isolation
Single-partition batchAtomic, isolated multi-row change; efficient groupingRequires a data model that colocates the rows
Concurrent single writesBest throughput and even loadNo cross-partition atomicity; retries are your job
Conditional batchCompare-and-set over several rows in one partitionPaxos round trips and contention under load
Application-level reconciliationNo batch cost; repairs itself over timeCode to detect and fix drift between tables

A useful rule of thumb: if the statements share a partition key, batch freely; if they span partitions, batch only when an inconsistency between those partitions would be a real bug, and keep the batch small and idempotent. Everything else belongs in concurrent single writes. Good partition design, described in data modeling basics, is what makes the cheap single-partition case available.

What to do next

  1. Search your codebase for BatchStatement and BEGIN BATCH. For each, write down whether it spans partitions and why it needs atomicity.
  2. Replace bulk-load batches that span many partitions with concurrent prepared single writes through a token-aware driver.
  3. Grep the server logs for batch-size warnings and for warnings about unlogged batches across many partitions, and trace each back to the code path.
  4. Check every batch that mixes a delete with a write of the same row for the shared-timestamp tie, and fix it with explicit timestamps or a different model.
  5. Make sure your retry policy distinguishes BATCH_LOG timeouts, which are safe to retry, from BATCH timeouts, which will complete by replay.
  6. Read the write and read path to understand what each mutation costs on a replica.
Key takeaway: A Cassandra batch is a guarantee, not a speed-up. Batch freely when every statement shares a partition key, use small idempotent logged batches only where cross-partition inconsistency would be a bug, remember that every statement shares one timestamp, and send everything else as concurrent single-partition writes.