Every Kafka cluster runs two planes. The data plane moves records: producers append to partition leaders, followers replicate, consumers fetch. The metadata plane decides who leads what: which brokers are alive, which replica leads each partition, which replicas are in sync, what each topic's configuration and ACLs are. For most of Kafka's life the metadata plane lived in Apache ZooKeeper. Since Kafka 4.0 it lives only in KRaft, a quorum of Kafka controllers that store metadata in a replicated log of their own. Kafka 3.9 is the last release that can run in ZooKeeper mode, and it is the release you migrate from.

This article explains both designs from first principles, shows why the change matters for failover time and operations, walks a migration step by step with the configuration it needs, and lists the failure modes on each side. Configuration names and tool flags here were checked against the Kafka 3.9 operations documentation on 2026-10-01; check the docs for your exact version before running anything, because tooling around dynamic quorums is still evolving.

Advertisement

The two designs in one picture

Where Kafka keeps its metadata: ZooKeeper mode versus KRaft modeZooKeeper mode (3.x and earlier)KRaft mode (only mode from 4.0)ZooKeeper ensembleznodes: brokers, topics, configs, ACLsBroker 1 = controllerwon the /controller znodeBroker 2Broker 3Broker 4reads ALL metadata on failoverLeaderAndIsr / UpdateMetadataController Aactive (leader)Controller Bhot standbyController Chot standby__cluster_metadata logRaft-replicated, snapshottedBroker 1Broker 2Broker 3fetch log, heartbeat
In ZooKeeper mode the active controller is an ordinary broker that rebuilds its view from ZooKeeper and pushes changes to brokers by RPC. In KRaft mode dedicated controllers replicate a metadata log with Raft; brokers fetch and replay that log.

The difference is where the truth lives and how it reaches brokers. In ZooKeeper mode the truth is a tree of znodes in a separate system, and one broker at a time acts as controller, reading that tree and pushing the results to every other broker. In KRaft mode the truth is an ordered log of metadata records, replicated among a small set of controller nodes using a Raft-style consensus protocol, and every broker keeps a local copy of the same log. Everything else in this article follows from that one change.

How the ZooKeeper-era controller worked

Each broker registered itself with an ephemeral znode under /brokers/ids; when its ZooKeeper session expired, the znode vanished and the cluster treated the broker as dead. Topic assignments, partition state, configurations and ACLs were stored as znodes too. The first broker to create the ephemeral /controller znode became the controller, and a counter called the controller epoch let brokers ignore orders from a deposed one.

On election, the new controller read the full metadata state from ZooKeeper, then sent LeaderAndIsr and UpdateMetadata requests to brokers so each knew what it led and followed. Every later change, such as a broker dying or a topic being created, went the same way: write to ZooKeeper, then push RPCs.

Three properties of this design caused most of the operational pain. First, failover cost grew with cluster size: a new controller had to load every partition's state before it could act, so clusters with very many partitions saw long windows where leadership changes stalled. Second, there were two sources of state: the controller's in-memory view, ZooKeeper's znodes and the brokers' caches could diverge after bugs or partial failures, and the push model offered no simple way to tell how far behind a broker was. Third, you ran two distributed systems, each with its own configuration, security model (SASL and TLS for ZooKeeper were separate from Kafka's), monitoring, upgrades and failure modes. KIP-500, the proposal behind KRaft, set out to remove all three.

Advertisement

KRaft: the metadata log and the quorum

In KRaft, controllers are Kafka processes started with process.roles=controller. Together they form a quorum that replicates a single-partition internal topic, __cluster_metadata. Each record in it is a metadata event: a broker registered, a topic was created, a partition's leader or in-sync replica set changed, a config was altered. The current state of the cluster is simply the result of replaying the log from the start, or from the latest snapshot.

One controller is the Raft leader and acts as the active controller; the others are followers. Leadership is decided by voting with a monotonically increasing epoch, so a stale leader is fenced the same way a stale ZooKeeper controller was, but without an external system. A metadata record is committed once a majority of controllers have it, which is why quorum sizes are odd: three controllers tolerate one failure, five tolerate two.

KRaft's replication differs from textbook Raft in one useful way: followers pull from the leader using Kafka's fetch mechanism, the same machinery partition followers use, rather than the leader pushing entries. Brokers also fetch the metadata log, but as observers that never vote. A broker's position in the log, its offset, tells you exactly how current its metadata is, which replaces the guesswork of the push model.

Because standby controllers already hold the full log in memory, failover no longer means reading the world from scratch. A new leader is elected, and it can act on state it already has. Log growth is bounded by periodic snapshots: a controller or broker writes its current in-memory state to a snapshot file and older log segments can be dropped. A node that falls far behind loads a snapshot and then fetches only the tail.

Brokers in KRaft: registration, heartbeats and fencing

A KRaft broker starts, registers with the active controller, and sends periodic heartbeats. Liveness is decided by those heartbeats rather than by a ZooKeeper session. A broker that registers but has not caught up with the metadata log stays fenced: the controller will not make it a partition leader until it reports that it has replayed enough of the log. A broker that stops heartbeating is fenced again, and the controller moves leadership of its partitions to in-sync replicas elsewhere, writing those leader changes as records in the log.

Controlled shutdown follows the same path. The broker asks the controller to move leadership away, the controller writes the changes, other brokers see them by fetching the log, and only then does the broker stop.

Feature levels are versioned through the log too. A record carries the cluster's metadata.version, and the kafka-features.sh tool reports and upgrades it after all nodes run a new release; keep that step separate from the binary upgrade so you can roll binaries back first.

Deploying a KRaft cluster

There are two deployment modes. In isolated mode controllers and brokers are different processes, ideally on different machines. In combined mode one process has process.roles=broker,controller. The Kafka documentation says combined mode can be used in development but should be avoided in critical deployments, because a broker under load can starve the controller sharing its process, and you cannot scale or restart one role without the other.

Quorums come in two flavours. A static quorum lists every voter in controller.quorum.voters and membership changes require coordinated restarts. A dynamic quorum, added by KIP-853 in Kafka 3.9, uses controller.quorum.bootstrap.servers for discovery and lets you add and remove controllers with kafka-metadata-quorum.sh or the Admin API. In 3.9 a static quorum cannot be converted to a dynamic one, so choose dynamic when you create a new cluster unless your tooling forbids it.

# controller.properties (one of three controllers)
process.roles=controller
node.id=1
listeners=CONTROLLER://ctrl-1.internal:9093
controller.listener.names=CONTROLLER
controller.quorum.bootstrap.servers=ctrl-1.internal:9093,ctrl-2.internal:9093,ctrl-3.internal:9093
log.dirs=/var/lib/kafka/metadata

# broker.properties
process.roles=broker
node.id=101
listeners=PLAINTEXT://broker-101.internal:9092
controller.listener.names=CONTROLLER
controller.quorum.bootstrap.servers=ctrl-1.internal:9093,ctrl-2.internal:9093,ctrl-3.internal:9093
log.dirs=/data/kafka

Storage must be formatted before first start, which writes the cluster ID and initial quorum into the log directories. Generate the ID once and use it on every node:

CLUSTER_ID=$(bin/kafka-storage.sh random-uuid)

# first controller of a new dynamic quorum
bin/kafka-storage.sh format --cluster-id "$CLUSTER_ID" --standalone \
    --config config/controller.properties

# controllers and brokers that join later
bin/kafka-storage.sh format --cluster-id "$CLUSTER_ID" --no-initial-controllers \
    --config config/broker.properties

# on ctrl-2 and ctrl-3, once each has caught up: promote it from observer to voter
bin/kafka-metadata-quorum.sh --command-config config/controller.properties \
    --bootstrap-server broker-101.internal:9092 add-controller

# health of the quorum: leader, epoch, high watermark, lag of each voter and observer
bin/kafka-metadata-quorum.sh --bootstrap-server broker-101.internal:9092 describe --status
bin/kafka-metadata-quorum.sh --bootstrap-server broker-101.internal:9092 describe --replication

A controller formatted with --no-initial-controllers only follows the log until it is added, so this sequence gives a one-voter quorum until you run add-controller on the other two; confirm all three are listed as voters before you rely on it. If you know all initial controllers up front, the --initial-controllers flag formats them as a group instead. Give controllers fast, dedicated disks: every metadata commit waits on an fsync at a majority of them, so a slow controller disk shows up as slow topic creation and slow leader elections cluster-wide.

Side by side

ConcernZooKeeper modeKRaft mode
Source of truthznodes in a separate ensembleRaft-replicated __cluster_metadata log
Controller failoverNew controller loads all state from ZooKeeperStandby already holds state in memory
How brokers learn changesPushed RPCs from the controllerBrokers fetch and replay the log; offset shows staleness
LivenessZooKeeper session and ephemeral znodeHeartbeats to the active controller; fencing
Systems to operateKafka plus ZooKeeperKafka only
Security configurationSeparate for ZooKeeperOne Kafka listener and auth model
Supported releasesUp to 3.9Production-ready in later 3.x; the only mode in 4.x

KRaft is not free of trade-offs. Controllers are a new tier you must size and monitor, the tooling (quorum description, metadata shell, feature levels) is different from what ZooKeeper-era runbooks assumed, and some third-party tools that read znodes directly simply stop working. Those are one-time costs; the ZooKeeper costs were permanent.

Migrating from ZooKeeper to KRaft

A ZooKeeper-mode cluster cannot be upgraded to 4.0 directly. It must first be migrated to KRaft while on 3.9, which the docs describe as the final and most complete iteration of the migration feature. The migration copies metadata into a new controller quorum while the cluster keeps serving traffic. The steps, in order:

  1. Prepare. Run every broker on 3.9.x (the docs ask for 3.9.1 or later) with inter.broker.protocol.version=3.9. Read the cluster ID from ZooKeeper; the new controllers must be formatted with the same ID.
  2. Start controllers in migration mode. Deploy the KRaft quorum with zookeeper.metadata.migration.enable=true and the zookeeper.connect string. They wait for brokers.
  3. Enable migration on brokers. Add zookeeper.metadata.migration.enable=true, controller.quorum.bootstrap.servers (or the static voters list) and controller.listener.names, plus the controller listener's security mapping, then do a rolling restart. When every broker has registered, the KRaft controller copies the metadata from ZooKeeper into the log and becomes the active controller. From now on it writes each change to the log and to ZooKeeper, so ZooKeeper-mode brokers keep working. This is the dual-write phase.
  4. Move brokers to KRaft. For each broker set process.roles=broker, keep node.id equal to its old broker.id, remove the ZooKeeper settings and the migration flag, and restart, one at a time.
  5. Finalise. Remove the migration flag and zookeeper.connect from the controllers and restart them. The controllers stop writing to ZooKeeper.

Step five is the point of no return. The documentation is explicit: after finalisation it is not possible to revert to ZooKeeper mode. Before it, a revert is possible, and the exact procedure depends on which step you completed last, so print the revert table from the docs for your version and rehearse it on a staging cluster. Keep the ZooKeeper ensemble running, and backed up, until you have run in KRaft for long enough to trust it. Then upgrade to 4.x.

Worked example: a 12-broker cluster

Consider a 3.7 cluster: 12 brokers, a three-node ZooKeeper ensemble, about 40,000 partitions, rolling restarts taking roughly 90 seconds per broker. A safe plan looks like this.

Week one: upgrade brokers to 3.9.x in a normal rolling restart, then bump inter.broker.protocol.version in a second roll. Week two: deploy three dedicated controllers on separate hosts (and separate racks or zones, matching the ZooKeeper layout), formatted with the existing cluster ID, in migration mode. Roll the brokers with migration settings, about 18 minutes for 12 brokers, and watch the controller logs and the quorum's describe --status until migration completes. Leave the cluster in dual-write for several days of normal traffic, including at least one deliberate controller failover, and compare leader-election times against your ZooKeeper-era baseline.

Week three: roll brokers into KRaft-only mode. Only after a further quiet period, finalise the controllers, then plan the 4.x upgrade and decommission ZooKeeper. Each phase is a stable state you can sit in, and only the last one is irreversible.

Failure modes

SymptomLikely causeWhat to do
Brokers stay fenced after startCannot reach the controller listener, or still replaying metadataCheck controller.listener.names and security mapping; watch the broker's metadata offset approach the leader's
Slow topic creation and electionsSlow controller fsync or an overloaded combined-mode nodeDedicated controllers on fast disks; isolated mode
Quorum loses its leaderMajority of controllers down or partitionedRestore a majority; never run two controllers across a single failure domain
Migration never startsA broker not registered or missing migration configEvery broker must run migration settings before metadata copies
Wrong cluster IDControllers formatted with a fresh IDReformat with the ID read from ZooKeeper before brokers connect
Tooling broke after migrationScripts read znodes directlyMove them to the Admin API and kafka-metadata-quorum.sh

Related reading

For the log model and client settings see the Kafka overview; for how data partitions replicate and track in-sync replicas see ISR replication. ZooKeeper itself is explained in the ZooKeeper deep dive, and managed clusters are covered in Amazon MSK.

What to do next

  1. Inventory your clusters by version and mode; anything in ZooKeeper mode needs 3.9.x before it can move to 4.x.
  2. Find every script, dashboard or tool that reads ZooKeeper directly and plan its replacement.
  3. For new clusters, use isolated mode with three or five dedicated controllers on separate failure domains and a dynamic quorum.
  4. Add quorum health to monitoring: active controller present, leader epoch changes, and each node's lag from describe --replication.
  5. Rehearse the migration and its revert on a staging copy, timing each rolling restart.
  6. Finalise only after a dual-write soak that includes a deliberate controller failover, and keep ZooKeeper backed up until then.
Key takeaway: ZooKeeper mode kept Kafka's metadata in a second system and rebuilt the controller's view from it on every failover; KRaft keeps it in a Raft-replicated log that standby controllers and brokers already hold, so failover is fast, staleness is measurable as an offset, and there is one system to secure and run. Kafka 4.0 supports only KRaft, and the way there is a 3.9 migration that is reversible until the final controller restart and permanent after it.