
11/8/2018 · Jun Rao
What this post added
Introduced significant improvements to the Kafka controller for handling a larger number of partitions per cluster. This involved optimizing controlled broker shutdowns by using asynchronous ZooKeeper writes and batched leader communication, reducing shutdown time from 6.5 minutes to 3 seconds in tests. Controller failover time was also improved by switching to asynchronous ZooKeeper APIs for state reloading, resulting in a 100% improvement in observed reloading time. The post recommends a limit of 4,000 partitions per broker and 200,000 partitions per cluster.