Cassandra will not return a large result set in one response, and you would not want it to: a coordinator that materialised a million rows in memory before replying would fall over long before your client did. Instead, queries are paged. The client asks for at most N rows, the cluster returns up to N rows plus an opaque marker that says where it stopped, and the client sends the same query again with that marker to get the next page.

The mechanism is simple, but it interacts with almost everything else in Cassandra: data modelling, consistency, tombstones, timeouts and API design. This article starts from what a page actually is on the wire, then covers driver usage in the DataStax Java driver 4.x, exposing pages through a web API without trusting the client, the alternatives (keyset paging and token-range scans), and the failure modes you will meet in production. A shorter overview is in Cassandra paging.

Advertisement

What a page is

Every CQL query carries a page size, the maximum number of rows the client wants in one response; the drivers call it page size or fetch size. It is not the same as a LIMIT clause. LIMIT caps the total result; the page size caps each response. A query with LIMIT 10000 and a page size of 500 returns twenty pages.

When a response has more rows behind it, the coordinator includes a paging state: a small blob that records the position of the last row returned, its partition key and clustering position, and counters for any remaining limit. The coordinator then forgets the query. Nothing is held open on the server between pages: no cursor, no snapshot, no locks. To fetch the next page the client sends the same statement with the same bound values plus the paging state, and any coordinator can decode it and resume the read just after that position.

Paging through a partition: each page is a separate read that resumes where the last one stoppedApplicationdriver sessionCoordinatorstateless between pagesReplica AReplica Bquery, page size 100read up to 100 rows100 rows + paging statePaging state (opaque bytes)encodes the last partition key and clustering position returned, plus remaining-row countersApplicationsame query + stateAny coordinatordecodes state, resumesReplicasread rows after positionpage 2 requestNo snapshot is held between pages: rows written or deleted in between may or may not appear.Each page is its own read at its own consistency level, so tombstones and timeouts are per page.
Pages are independent reads stitched together by the paging state. Any coordinator can serve the next page because the state travels with the request.

Two consequences follow from that design. First, a paged read is not a consistent snapshot. Rows inserted after your position, in a part of the result you have not reached yet, will appear; rows deleted from that part will not. Rows changed behind your position are simply not revisited. For most listing use cases this is fine, but jobs that need a point-in-time view must get it some other way, for example by writing with a version or timestamp column and filtering on it. Second, each page is a full read: it runs at the statement's consistency level, goes through the replicas' read path, and can time out or fail independently of the pages before it. The read path itself is described in the Cassandra read path.

Treat the paging state as opaque. Its format is internal to the server and has changed between native protocol versions, so do not parse it, and do not store it for long periods or across cluster upgrades. Use it to continue a listing now, not as a durable bookmark.

Paging in the Java driver 4.x

The default page size comes from configuration, datastax-java-driver.basic.request.page-size, which is 5000 in the driver's reference configuration, and can be overridden per statement. The synchronous ResultSet pages transparently: iterating it returns the rows of the current page, and when they run out it blocks while it fetches the next page. That is convenient for batch jobs, and it hides the cost, since a loop that looks like it is walking an in-memory list is issuing a network read every N rows.

SimpleStatement stmt = SimpleStatement.builder(
        "SELECT event_id, ts, payload FROM events_by_device WHERE device_id = ? AND day = ?")
    .addPositionalValues(deviceId, day)
    .setPageSize(500)
    .build();

ResultSet rs = session.execute(stmt);
for (Row row : rs) {               // fetches page 2, 3, ... transparently, blocking each time
    process(row);
}

For services, the asynchronous API makes the page boundary explicit. executeAsync returns an AsyncResultSet. currentPage() gives only the rows already fetched, hasMorePages() says whether to continue, and fetchNextPage() returns a CompletionStage for the next page. No thread blocks between pages:

CompletionStage<Long> countRows(AsyncResultSet rs, long soFar) {
    long n = soFar;
    for (Row row : rs.currentPage()) {
        process(row);
        n++;
    }
    if (rs.hasMorePages()) {
        long total = n;
        return rs.fetchNextPage().thenCompose(next -> countRows(next, total));
    }
    return CompletableFuture.completedFuture(n);
}

session.executeAsync(stmt).thenCompose(rs -> countRows(rs, 0));

Choosing the page size is a trade-off between round trips and per-request cost. Small pages, say 50 to 100 rows, keep each read cheap and latency predictable but multiply round trips. Large pages cut round trips but make each read longer, raise coordinator memory per request, and make it more likely that a single page hits a timeout. Think in bytes, not rows: 5,000 rows of 100 bytes is half a megabyte, while 5,000 rows carrying 50 KB payloads is 250 MB in one response. Set page size per query from the row size.

Advertisement

Exposing pages through a web API

A REST endpoint that lists a user's events cannot keep a driver ResultSet open between HTTP requests. Instead it returns the paging state to the client as a cursor and accepts it back on the next call. The driver provides two forms. getExecutionInfo().getPagingState() returns the raw ByteBuffer. getExecutionInfo().getSafePagingState() returns a PagingState object that can be serialised with toString() and parsed with PagingState.fromString(...), and whose matches(statement) method checks that the state was produced by the same query with the same bound values.

That check matters. A raw paging state handed to a browser is untrusted input. A client can replay a cursor from a different query, or tamper with the bytes, and the server will try to resume from whatever position they decode to. Validate before you use it, and encrypt or sign the cursor if the position itself is sensitive:

public Page<Event> listEvents(String deviceId, LocalDate day, String cursor) {
    SimpleStatement stmt = SimpleStatement.builder(
            "SELECT event_id, ts, payload FROM events_by_device WHERE device_id = ? AND day = ?")
        .addPositionalValues(deviceId, day)
        .setPageSize(50)
        .build();

    if (cursor != null) {
        PagingState state = PagingState.fromString(cursor);     // throws on malformed input
        if (!state.matches(stmt)) {
            throw new BadRequestException("cursor does not belong to this query");
        }
        stmt = stmt.setPagingState(state);
    }

    AsyncResultSet rs = session.executeAsync(stmt).toCompletableFuture().join();
    List<Event> items = new ArrayList<>();
    for (Row row : rs.currentPage()) items.add(Event.from(row));   // exactly one page, no extra fetch

    String next = rs.hasMorePages()
        ? rs.getExecutionInfo().getSafePagingState().toString()
        : null;
    return new Page<>(items, next);
}

Use currentPage() here rather than iterating a synchronous ResultSet: iterating past the end of the first page would silently fetch the next one, and the cursor you return would then skip a page.

Keyset paging: cursors you control

An alternative to the opaque paging state is keyset paging, sometimes called seek paging. Instead of a server-generated cursor, the client sends the last clustering key it saw and the next query starts strictly after it. It works whenever the listing is a slice of one partition in clustering order, which in a well-modelled Cassandra table is the common case.

-- Table: events for one device per day, newest first.
CREATE TABLE events_by_device (
    device_id text, day date, ts timestamp, event_id timeuuid, payload blob,
    PRIMARY KEY ((device_id, day), ts, event_id)
) WITH CLUSTERING ORDER BY (ts DESC, event_id DESC);

-- First page
SELECT ts, event_id, payload FROM events_by_device
 WHERE device_id = 'd-17' AND day = '2026-10-03' LIMIT 50;

-- Next page: strictly after the last (ts, event_id) returned, using a tuple comparison
SELECT ts, event_id, payload FROM events_by_device
 WHERE device_id = 'd-17' AND day = '2026-10-03'
   AND (ts, event_id) < ('2026-10-03 09:14:02+0000', 8d6c2e10-a0f1-11f0-8de9-0242ac120002)
 LIMIT 50;

The tuple comparison on both clustering columns is what makes this correct: comparing only ts would skip or repeat events sharing a timestamp. Because the table is ordered descending, "after" means <. Keyset cursors are readable, stable across upgrades and easy to validate, since they are just typed values, and they let the client jump to a time. They do not work across partitions, so a listing that spans days needs the application to move to the next day's partition when one is exhausted, which is usually what you want anyway.

What neither approach gives you is "go to page 37". The driver includes an OffsetPager that emulates it by re-running the query and skipping rows client-side: new OffsetPager(20).getPage(rs, 37) reads and discards 720 rows to return 20. Its cost is linear in the page number, and the driver documentation warns to enforce a maximum page number so a client cannot request page one million as a cheap denial of service. Prefer next/previous navigation in the UI.

Full-table reads: page by token range, not by one giant query

A query without a partition key, such as SELECT * FROM events_by_device, does page, but it walks the entire ring in token order through one coordinator, one page at a time, at the speed of a single client thread. If one page fails at hour three, you restart from the beginning unless you saved the state. For exports, backfills and migrations, split the ring into token ranges and scan them in parallel, each with its own paged query:

SELECT device_id, day, ts, event_id, payload FROM events_by_device
 WHERE token(device_id, day) > ? AND token(device_id, day) <= ?;

Take the ranges from the driver's token map, session.getMetadata().getTokenMap(), split each into smaller sub-ranges, and process them with a bounded worker pool. Record each finished range in a checkpoint table so a failed job resumes at range granularity instead of starting over. Routing each range's query to a replica that owns it avoids an extra hop. How tokens map to nodes is explained in the token ring and vnodes.

Tombstones, consistency and timeouts per page

Page size counts live rows returned, not rows scanned. If a partition holds 100,000 deleted rows before the first 50 live ones, the replica reads through all of those tombstones to fill your first page, and the warning and failure thresholds configured on the server apply to that one read. A listing that used to be fast becomes slow, then fails, as deletes accumulate in front of the live data, typically in queue-like tables where consumers delete what they have read. Smaller pages do not help, because the tombstones sit before the first live row. Fix the model, for example by bucketing by time so readers start from a partition without the dead history, as discussed in Cassandra tombstones.

Consistency is per page too. A listing at LOCAL_QUORUM gets quorum guarantees for each page, but different pages may be served by different replica sets as nodes come and go, and the listing as a whole is still not a snapshot. Per-page timeouts mean a long scan needs retry logic at page granularity: on a read timeout, retry the same statement with the same paging state rather than restarting the scan.

Worked example: an activity feed

A mobile app shows each device's events, newest first, 50 per screen, and users scroll back up to a week. Devices emit up to 20,000 events a day, each about 400 bytes. The design: partition by (device_id, day) so no partition exceeds about 8 MB, cluster by (ts, event_id) descending, and page with a keyset cursor encoded as day|ts|event_id. The API returns 50 rows; when the current day's partition is exhausted, the service moves to the previous day and continues, stopping after seven days. Each request is one cheap single-partition read of about 20 KB.

For a nightly export of all events to object storage, the same table is scanned by token range: 64 ranges per node, 16 workers, page size 1,000 (about 400 KB per page), and a checkpoint row per finished range. A node restart during the export costs a retry of the few ranges in flight, not the whole job.

Failure modes

  • A cursor that skips a page. A web handler iterated a synchronous result set past the page boundary before reading the paging state. Fix: use currentPage().
  • Tampered or mismatched cursor. A client sends a cursor from another query. Fix: PagingState.fromString plus matches, and a 400 response on mismatch.
  • Deep offset paging. Page-number URLs with OffsetPager make late pages slow and expensive. Fix: cap the page number or switch to next/previous navigation.
  • Tombstone walls. Pages slow down over weeks as deletes pile up in front of live rows. Fix: time-bucketed partitions and TTLs instead of explicit deletes.
  • Restarted full scans. One failed page restarts a multi-hour export. Fix: token-range splits with checkpoints and page-level retry.

What to do next

  1. Audit queries that use the default page size and set an explicit size based on average row bytes.
  2. In every API that returns a cursor, switch to getSafePagingState() and validate incoming cursors with matches.
  3. Check that web handlers read rows with currentPage() so they never fetch past one page.
  4. Where a listing is a slice of one partition, consider a keyset cursor on the full clustering key.
  5. Remove page-number navigation, or cap it, wherever OffsetPager or an equivalent skip is used.
  6. Rewrite full-table jobs as parallel token-range scans with a checkpoint per range.
  7. Trace one slow page with TRACING ON in cqlsh and look at how many tombstones it read; see Cassandra consistency for choosing the level each page runs at.
Key takeaway: Cassandra pages every large result: the client asks for at most N rows, gets them with an opaque paging state, and sends the same query back with that state for the next page. The server holds nothing between pages, so a paged read is not a snapshot, and each page is its own read with its own consistency, tombstone and timeout behaviour. Size pages in bytes, read exactly one page per web request with currentPage, validate client cursors with PagingState and matches, prefer keyset cursors for single-partition listings, avoid offset paging, and scan whole tables by token range with checkpoints.