
2/12/2018
What this post added
This post details the Production Engineering team's approach to supporting global events on Facebook, specifically focusing on the planning and infrastructure required for Facebook Live during New Year's Eve. It outlines the three categories of load variance (routine, spontaneous, planned), the architecture of Facebook Live, key resource metrics (network, CPU, storage), and the load metrics used for planning (total broadcasts, peak concurrent broadcasts, dependent system load). The post describes the process of estimating traffic using historical data and predictive modeling, scaling strategies like hardware allocation and segment consolidation, and load testing methods (artificial load, synthetic tests, shadow traffic). A specific example of an unexpected issue with health checks impacting routing hosts is also highlighted.