
4/8/2024 · Matthew Prince, John Graham-Cumming, Jeremy Hartman
What this post added
This post details Cloudflare's response to a second major data center power failure in Portland, Oregon, within five months. Following the first incident, the company initiated 'Code Orange' to prioritize high availability of its control plane and analytics infrastructure. The post highlights the successful automated failover of critical services like APIs and Dashboards within minutes during the second outage, contrasting with significant downtime experienced previously. Key improvements include moving configuration databases to a Highly Available (HA) topology, pre-provisioning capacity for automatic failover, making Logpush infrastructure HA with an active failover option in Amsterdam, and enhancing the resiliency of services like Stream and Zero Trust by distributing functionality across more data centers. The time to bring a data center back online was also reduced from 72 hours to approximately 10 hours.