Site Reliability and Outage Management
How Facebook deals with PCIe faults to keep our data centers running reliably

How Facebook deals with PCIe faults to keep our data centers running reliably

6/2/2021 · Ashwin Poojary, Bill Holland, Makan Diarra, Ray Park

What this post added

This post details Meta's approach to managing PCIe faults in its data centers, introducing a workflow and suite of tools (PCIcrawler, MachineChecker, PCIe Error Logging Service, FBAR) to detect, diagnose, remediate, and repair hardware issues. It highlights the importance of analyzing PCIe error rates (corrected and uncorrected), link speeds, and link widths, and describes automated remediation strategies such as reseating components, swapping hardware, and identifying firmware-related issues through data analysis in Scuba. The post also advocates for industry-wide adoption of PCIe AER functionality.

Read the original post ↗