Kafka's ordering guarantee is short: records in one partition are stored and delivered in the order the partition leader appended them. Everything else, including the ordering people actually care about, such as all events for one order or one account arriving in sequence, is something your producers, keys, consumers and operations have to preserve on top of that guarantee.

Most ordering bugs are not in Kafka. They come from a producer retry, a partition added to a busy topic, a thread pool in a consumer, or a retry topic that sends one event around the back of the queue. This article explains the guarantee from first principles, then walks through the six places where per-key order is made or broken, with code for each fix. Key design is covered in Kafka partitioning strategies and transactional pipelines in exactly-once stream processing; this page is about order.

Advertisement

What Kafka actually promises

A topic is split into partitions, and each partition is an append-only log whose records get consecutive offsets. The leader replica appends records, followers copy them, and consumers read by offset. So within a partition, offset order is the order of appends, and every consumer sees the same order. Across partitions there is no order at all: two records in different partitions can be read in either order, by different consumers, at very different times.

The practical consequence is that order is scoped by the key. The default partitioner hashes a record's key (murmur2 of the serialized key, modulo the partition count), so all records with the same key land in the same partition and are ordered relative to each other. Choose the key as the entity whose events must stay in sequence, such as order ID or account ID, and accept that events for different entities are unordered. If you need a total order across everything, you need one partition, and you get one partition's throughput.

The six places per-key order breaks

Where per-key order is made, and where it can break (red markers)Producer appsend() call orderPartitionerhash(key) mod NAccumulatorbatch per partitionIn-flightup to 5, seq numbersPartition leadercheck seq, appendFollowersISR copiesConsumer polloffset orderWorker lanesone lane per keySide effectsDB, API, topic1234561 concurrent send() calls 2 partition count or key change 3 retries without idempotence, expired batches4 unclean leader election 5 thread pools that ignore keys 6 rebalance replay, retry topics, mirroring
The path of a keyed record from producer to side effect. Each red marker is a place where per-key order can be lost.

Order is made at the producer and preserved or lost at every later hop. The sections below follow the diagram's markers: concurrent sends, partitioning changes, producer retries, leader failover, consumer concurrency, and redelivery through rebalances, retry topics and mirroring.

Advertisement

1. The producer application

The producer preserves the order of send() calls for records going to the same partition. It cannot preserve an order your application never had. If two request threads both emit events for order 42, which one calls send() first decides the log order, and that may not match the order in which the underlying state changed. The fix belongs in the application: emit events for one entity from one place, for example after the database commit that changed it, or from the database log with change data capture, which serializes events in commit order.

2. Partitioning and key changes

Because the partition is hash(key) mod N, adding partitions changes where existing keys go. Suppose order-42 hashes to a value that is 7 modulo 12 and 3 modulo 16. Before an expansion from 12 to 16 partitions its events go to partition 7; afterwards they go to partition 3. A consumer of partition 3 can now process the new events while older events for the same order are still waiting in partition 7's backlog.

Treat the partition count of an ordered topic as fixed. Choose it for several years of growth, and if you must expand, drain the topic first or migrate to a new topic with a cut-over. The same applies to changing the key, and to records without a key: the producer spreads keyless records across partitions (since Kafka 3.3 with a built-in sticky, load-aware strategy), so they have no per-entity order at all. partitioner.ignore.keys exists and does exactly what it says, so check that nobody has set it on an ordered topic.

3. Retries and the idempotent producer

The producer batches records per partition and can have several requests in flight on one connection, up to max.in.flight.requests.per.connection, default 5. Without idempotence this reorders data: batch A fails with a transient error, batch B behind it succeeds, then A is retried and lands after B. A retry after a lost acknowledgement also writes A twice.

The idempotent producer fixes both. The broker gives the producer a producer ID and epoch, and the producer numbers each batch with a per-partition sequence. The leader appends a batch only if its sequence is the next one expected, rejects out-of-order batches so the client resends them in order, and recognises duplicates of recent batches. The broker tracks the last five batches per producer and partition, which is why idempotence requires at most five in-flight requests.

Current producer documentation lists enable.idempotence=true, acks=all, retries=2147483647 and delivery.timeout.ms=120000 as defaults. There are two traps. Idempotence became the default in 3.0, but a bug meant it was not actually applied until 3.0.1, 3.1.1 and 3.2.0. And since those releases, if you do not set enable.idempotence explicitly but set a conflicting value, such as acks=1 or more than five in flight, the producer quietly disables idempotence rather than failing. Set it explicitly.

Properties p = new Properties();
p.put("bootstrap.servers", "broker-1:9092,broker-2:9092");
p.put("key.serializer", StringSerializer.class.getName());
p.put("value.serializer", StringSerializer.class.getName());
// State these explicitly: setting acks=1 alone silently turns idempotence off.
p.put("enable.idempotence", "true");
p.put("acks", "all");
p.put("max.in.flight.requests.per.connection", "5");  // must be 5 or fewer
p.put("delivery.timeout.ms", "120000");
KafkaProducer<String, String> producer = new KafkaProducer<>(p);

// The key decides the partition, so it decides what is ordered relative to what.
producer.send(new ProducerRecord<>("orders", order.id(), json(event)), (md, ex) -> {
    if (ex != null) {
        // Later records for this key may already be written. Do not "resend later":
        // fail the entity or stop the producer, and let an operator decide.
        poisonedKeys.add(order.id());
        log.error("send failed for key {}", order.id(), ex);
    }
});

Idempotence does not survive everything. A plain idempotent producer gets a new producer ID when it restarts, so its guarantees cover one producer session; a transactional.id carries identity across restarts. And when a batch fails for good, for example after delivery.timeout.ms expires, later batches for the same key may already be written. Resending the failed record later appends it after them, which is a reordering. The callback above marks the key as poisoned instead; what to do next is a business decision, not a retry loop.

4. Leader failover

An append is only durable once it is on enough replicas. With acks=all and min.insync.replicas=2 on a replication factor of 3, an acknowledged record survives the loss of any single broker. If unclean leader election is enabled, an out-of-sync replica can become leader and the log is truncated back to what it had, so acknowledged records vanish and new ones take their offsets. It is disabled by default; keep it that way for ordered topics. See ISR replication for the mechanics.

5. Consumer concurrency

Each partition is assigned to one consumer in a group, and poll() returns its records in offset order. Most consumer-side ordering bugs come from what happens next: records handed to a thread pool are processed in whatever order the threads get to them. Two events for the same order, a creation and a cancellation, run in parallel, and the cancellation can finish first.

The fix is to parallelise across keys but not within a key. Route each record to a single-threaded lane chosen by its key, and commit offsets only when every record up to them is done. The loop below uses a simple barrier per poll; libraries such as Confluent's open-source Parallel Consumer implement the same idea with per-key ordering and finer-grained offset tracking.

ExecutorService[] lanes = new ExecutorService[16];
for (int i = 0; i < lanes.length; i++) lanes[i] = Executors.newSingleThreadExecutor();

while (running) {
    ConsumerRecords<String, String> records = consumer.poll(Duration.ofMillis(500));
    List<Future<?>> pending = new ArrayList<>();
    for (ConsumerRecord<String, String> r : records) {
        // same key -> same lane -> processed in offset order; keys run in parallel
        int lane = Math.floorMod(Objects.hashCode(r.key()), lanes.length);
        pending.add(lanes[lane].submit(() -> handle(r)));
    }
    for (Future<?> f : pending) f.get();   // barrier: everything polled is done
    consumer.commitSync();                  // only then move the committed offsets
}

The barrier trades throughput for simplicity: one slow key holds up the next poll. If that matters, track completed offsets per partition and commit the highest offset below which everything is done. Never commit an offset whose records are still in flight, or a crash will skip them.

6. Redelivery: rebalances, retry topics and mirroring

Kafka consumers are at-least-once by default. After a crash or a rebalance, the new owner of a partition starts from the last committed offset and replays everything after it. Order is preserved in the replay, but your handler sees events it has already applied, which to downstream state looks like going back in time. Make handlers idempotent and order-aware by storing a version or source offset per entity and ignoring anything older, as the SQL below does. Cooperative rebalancing reduces how often this happens; see consumer rebalancing.

-- Apply an event only if it is newer than what we already hold for this entity.
UPDATE orders
   SET status = :status, version = :version, updated_at = now()
 WHERE order_id = :order_id
   AND version < :version;      -- replayed or stale events update zero rows

Retry topics and dead-letter queues break order by design. If event 5 for an order fails and is sent to a retry topic, event 6 is processed first. For ordered entities, either block the partition and retry in place with backoff, or park the whole key: once one event for a key is in the retry path, route that key's later events there too until it clears. Cross-cluster replication keeps the order of each source partition, but a failover resumes from translated offsets and can replay recent records, so the same version checks are needed there. Transactions make writes to several partitions atomic; they do not create an order between them.

Worked example: an order lifecycle

An order service emits Created, Paid, Shipped and Cancelled to a 24-partition topic keyed by order ID. The service writes the order row and an outbox row in one database transaction, and a relay publishes outbox rows in commit order with an explicitly idempotent producer. Each event carries the order's version number.

Downstream, a fulfilment consumer uses 16 key lanes, commits after each poll's barrier, and applies events with the conditional update above. During a deploy, a rebalance replays 300 events: every one updates zero rows. During a payment-provider outage, Paid for one order fails in the handler; the consumer parks that order's key, and its later Shipped waits with it instead of overtaking. The partition count stays at 24, because expanding would remap keys; when throughput nearly outgrew it, the team created a 96-partition topic and cut over after draining the old one.

Trade-offs and failure modes

ChoiceOrdering gainCost
Key by entityPer-entity orderHot keys limit one partition's throughput
Explicit idempotenceNo retry reordering or duplicatesSlightly more broker state
Fixed partition countStable key placementMust overprovision up front
Key lanes in consumersParallelism without reorderingHead-of-line blocking per lane
Park key on failureNo overtakingOne bad event stalls an entity
  • Idempotence silently off after someone sets acks=1 for latency.
  • Partitions added under load, splitting in-flight keys across two partitions.
  • Commit before processing in an async consumer, so a crash skips records.
  • Retry topic for an ordered stream, letting later events overtake a failed one.
  • Share groups (queue-style consumption from KIP-932) deliver records from one partition to several consumers, so they do not give per-key order; do not use them for ordered streams.

What to do next

  1. List every topic whose consumers depend on order, and write down the entity each is ordered by; make that the key.
  2. Set enable.idempotence=true and acks=all explicitly on those producers, and check client versions are 3.2 or later.
  3. Freeze partition counts on ordered topics and add a review step before any change.
  4. Audit consumers for thread pools and async commits; route by key into lanes and commit only completed offsets.
  5. Add a version or source offset to events and make every handler ignore stale ones.
  6. Replace retry topics on ordered streams with in-place retries or per-key parking, and keep unclean.leader.election.enable=false.
Key takeaway: Kafka orders records within a partition and nowhere else. Per-key order is something you build: key by the entity, keep the partition count fixed, run an explicitly idempotent producer and never resend a failed record behind later ones, avoid unclean leader election, process each key in one lane, and make handlers reject stale or replayed events. Most real ordering bugs live in the application code around Kafka, so audit that code first.