
6/16/2023 · Laurent Bernaille
What this post added
This post details the platform-level recovery process following a major incident, focusing on the EU1 region. It describes the phased restoration of Kubernetes control planes (parent and child clusters) and application nodes, highlighting the importance of rebooting nodes to regain network connectivity. It also details the challenges encountered when scaling clusters to process backlogged data, specifically hitting limits on the maximum number of VM instances in a peering group and the maximum number of IPs available in subnets for high-traffic clusters. The post explains how these limits delayed full recovery and the steps taken to address them, including engaging with the cloud provider.