For years, Apache Kafka required ZooKeeper to manage broker metadata, leader election, and cluster configuration. In Kafka 4.0, ZooKeeper is gone. The KRaft consensus protocol, built into Kafka itself, handles all of this natively. For enterprise teams, the implications are significant.
What KRaft Actually Changes
KRaft (Kafka Raft) replaces ZooKeeper with a built-in metadata quorum using the Raft consensus algorithm. A set of dedicated controller nodes, separate from broker nodes in production deployments, manages all cluster metadata: partition assignments, ISR lists, broker registrations, topic configurations, and ACLs.
This is not a minor plumbing change. ZooKeeper was a separate distributed system with its own failure modes, scaling characteristics, operational requirements, and security model. Removing it reduces operational complexity substantially.
Operational Benefits for Enterprise Teams
- Single system to operate: no separate ZooKeeper cluster to provision, monitor, patch, and secure
- Faster metadata operations: controller failover drops from 30–60 seconds to under 10 seconds
- Larger partition counts: KRaft supports millions of partitions per cluster vs ZooKeeper's practical limit of ~200K
- Simplified security: one auth model (SASL/SCRAM + mTLS) across the entire cluster
- Cleaner Kubernetes deployment: Strimzi Operator KRaft support is GA in Kafka 4.x
The Migration Question
For teams running Kafka 2.x or 3.x in production, migration to KRaft is the most common question. The answer depends on your current version, partition count, and operational model.
Kafka 3.x clusters can migrate to KRaft mode using the built-in migration tool without a full cluster rebuild, provided the Kafka version is 3.6 or later. Kafka 2.x clusters require a rolling upgrade to 3.x first. FluxNode handles this migration path as a managed engagement, our platform engineers have executed this process across multiple enterprise Kafka clusters in the UAE.
FluxNode deploys Kafka 4.x in KRaft mode with dedicated controller and broker node pools, distributed across 3 availability zones. ZooKeeper is not required and is not installed.
Controller and Broker Separation in Production
For production deployments, the KRaft best practice is separating controller nodes from broker nodes. Controller nodes manage metadata and consensus; brokers handle data. This separation ensures that a broker failure does not affect the control plane, and controller nodes can be sized differently, typically less memory-intensive than brokers.
FluxNode's reference architecture uses 3 dedicated controller nodes in a quorum plus 6 broker nodes (configurable), deployed across 3 availability zones. Replication factor defaults to 3. This topology provides controller failover under 10 seconds and broker failover under 30 seconds.
What Has Not Changed
KRaft does not change the Kafka client API. All existing producers, consumers, and Kafka Connect connectors are binary-compatible. Debezium CDC, Apache Camel connectors, and Flink Kafka sources/sinks all work without modification. The migration path is an operational exercise, not a code rewrite.