Confluent Cloud Availability and Resilience
Scaling Kafka to 10+ GB/Second in Confluent Cloud

Scaling Kafka to 10+ GB/Second in Confluent Cloud

6/12/2020 · Dan Rosanova

What this post added

This post details Confluent Cloud's ability to scale Kafka clusters to over 10 GBps of aggregate throughput without downtime. It explains the underlying mechanisms for elastic scaling, including the concept of Confluent Units for Kafka (CKUs) which encapsulate cloud resources. The process involves provisioning new CKUs, which adds brokers and storage, followed by an automated, sophisticated partition rebalancing strategy. This strategy involves identifying candidate partitions for movement based on broker utilization, incrementing ISR counts, placing new replicas on new brokers, and then carefully managing failovers to new leaders. The post highlights the use of Kubernetes and the Confluent Operator for infrastructure automation and discusses the algorithm for rebalancing, including optimizations and the impact of leader changes on client connections.

Read the original post ↗