BlogsMetaData Center Energy Efficiency and Optimization

Data Center Energy Efficiency and Optimization

Data Center Energy Efficiency and Optimization

64
posts
2010–2026

Meta's commitment to optimizing data center energy usage and improving overall efficiency has evolved from initial strategies like improved airflow management and adjusted temperature set points to a fundamental redesign of hardware and data center infrastructure. This includes the development of custom-designed servers, power supplies, racks, and battery backup systems, leading to significant reductions in energy consumption and cost. A key milestone was the formation of the Open Compute Project. This evolution now includes the deployment of the StatePoint Liquid Cooling (SPLC) system, a novel evaporative cooling technology that uses water instead of air to cool data halls, enabling the construction of highly energy- and water-efficient data centers in a wider range of environmental conditions and further reducing water usage.

2026

Lights Out, Systems On: Validating Instant Power Loss Readiness

6/3/2026

This post introduces 'Instantaneous PowerLoss Storm,' a new testing paradigm for validating data center readiness against zero-notice power loss. It details the defense-in-depth strategies implemented across the data center stack, including solutions for bootstrapping challenges like circular dependencies and the 'boomerang' problem. The post also outlines the tradeoffs made between reliability and growth velocity, defining tolerable impacts, and describes the incremental validation process through exercises in pre-production and production regions, culminating in the simulation of large-scale power loss events.

Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale

4/16/2026

This post introduces Meta's Capacity Efficiency Program, which utilizes a unified AI agent platform to automate the detection and resolution of performance issues across its infrastructure. The platform employs AI agents that encode domain expertise into reusable skills, leveraging standardized tool interfaces for both proactive optimization ('offense') and regression detection/mitigation ('defense'). This has resulted in the recovery of hundreds of megawatts of power and a significant reduction in manual engineering investigation time, enabling the program to scale efficiently.

Investing in Infrastructure: Meta’s Renewed Commitment to jemalloc

3/2/2026

This post details Meta's renewed commitment to jemalloc, a high-performance memory allocator, emphasizing its role as a foundational component for reliable and performant infrastructure. It highlights efforts to reduce technical debt, modernize the codebase, and collaborate with the open-source community to evolve jemalloc for current and future hardware and workloads, including specific improvements for huge-page allocation, memory efficiency, and AArch64 optimizations.

2025

Design for Sustainability: New Design Principles for Reducing IT Hardware Emissions

10/14/2025

This post introduces 'Design for Sustainability,' a new set of technical design principles for IT hardware aimed at reducing emissions and cost. Key contributions include: detailing modular rack designs like ORv3 and ORv3N with specific features for power and flexibility; emphasizing the reuse and retrofitting of existing rack designs to reduce e-waste and costs; advocating for the use of green steel produced with renewable energy and recycled steel, aluminum, and copper in hardware components; and highlighting the importance of improving hardware reliability to extend useful life.

How Meta Is Leveraging AI To Improve the Quality of Scope 3 Emission Estimates for IT Hardware

10/14/2025

This post details Meta's new methodology for improving Scope 3 emission estimates for IT hardware by leveraging AI. It introduces the use of NLP for identifying similar components to existing Product Carbon Footprints (PCFs) and LLMs (specifically Llama 3.1) for extracting data from heterogeneous sources and creating a standardized taxonomy for IT hardware carbon footprints. The post highlights the collaboration with the OCP PCR workstream to open-source this methodology, aiming to foster industry-wide sustainable manufacturing practices.

OCP Summit 2025: The Open Future of Networking Hardware for AI

10/14/2025

This post details Meta's advancements in open networking hardware for AI training clusters, presented at OCP Summit 2025. It introduces the evolution of Disaggregated Scheduled Fabric (DSF) to a dual-stage architecture supporting larger AI clusters, a new Non-Scheduled Fabric (NSF) architecture for massive AI clusters like Prometheus, and new OCP switch platforms like Minipack3N based on NVIDIA's Spectrum-4 ASIC. It also highlights advancements in optical interconnects (2x400G FR4 LITE, 400G/2x400G DR4) and the launch of the Ethernet for Scale-Up Networking (ESUN) initiative, underscoring Meta's leadership in driving open, scalable, and interoperable networking solutions for AI infrastructure.

A case for QLC SSDs in the data center

3/4/2025

This post introduces the exploration and adoption of Quad-Level Cell (QLC) Solid State Drives (SSDs) as a new storage tier in Meta's data centers. It highlights QLC's advantages in higher density, improved power efficiency, and better cost compared to existing TLC SSDs, positioning it as a middle ground between HDDs and TLC SSDs. The post details the hardware and software considerations for integrating QLC, including form factor challenges and adapting storage software for high-density, read-bandwidth-intensive workloads, aiming to significantly increase byte density and reduce power consumption at the server and drive level.

How Precision Time Protocol handles leap seconds

2/3/2025

This post details Meta's approach to handling leap seconds within data centers using Precision Time Protocol (PTP). It contrasts PTP's nanosecond precision with NTP's millisecond precision and explains Meta's 'self-smearing' algorithmic approach via the fbclock library. The post also discusses the trade-offs of this approach, the challenges of integrating PTP with NTP, and advocates for the elimination of leap seconds in favor of UTC for simplified and more precise timekeeping.

2024

OCP Summit 2024: The open future of networking hardware for AI

10/15/2024

This post details Meta's latest contributions to the Open Compute Project (OCP) Summit 2024, focusing on next-generation network hardware for AI training clusters. It introduces two new disaggregated network fabrics (DSF) and a new NIC (FBNIC), highlighting their role in enabling more flexible, scalable, and efficient AI infrastructure. The post also showcases specific hardware like the Arista 7700R4 series, Minipack3, and Cisco 8501 switches, along with advancements in optics and the evolution of Meta's FBOSS and OCP-SAI software, all aimed at improving AI data center performance and sustainability through open hardware.

Simulator-based reinforcement learning for data center cooling optimization

9/10/2024

This post details Meta's application of simulator-based reinforcement learning (RL) to optimize data center cooling. It explains how RL models the cooling control system as a sequential state machine, using environmental variables as states and control setpoints (e.g., supply airflow) as actions. The approach uses a physics-based simulator to train the RL agent offline, exploring potential actions and their rewards to learn an optimal policy. This has led to an average reduction of 20% in supply fan energy consumption and 4% in water usage in pilot regions, while maintaining data center temperature conditions within specifications. The post also highlights the applicability of this methodology to future AI-optimized data center designs.

RETINAS: Real-Time Infrastructure Accounting for Sustainability

8/26/2024

Introduced the RETINAS initiative and a new metric: 'real-time server fleet utilization effectiveness.' This metric measures server resource usage (compute, storage) and efficiency in near real-time to reduce emissions from embodied carbon. It applies depreciation concepts from finance to server reliability, efficiency, and useful life, integrating embodied carbon into infrastructure metrics for informed server fleet management decisions and to reduce Scope 3 emissions. The post details the metric's formula, its application through examples comparing static and dynamic accounting, and its characteristics for horizontal and vertical slicing to enable relative comparison of circularity strategies.

DCPerf: An open source benchmark suite for hyperscale compute applications

8/5/2024

This post introduces DCPerf, an open-source benchmark suite developed by Meta to accurately represent the diverse workloads found in hyperscale cloud deployments. It addresses the limitations of existing benchmarks by analyzing production workloads and capturing their characteristics, enabling better design and evaluation of server and datacenter hardware and software. DCPerf has been used internally at Meta for product evaluation and capacity planning, and is now being shared with the broader industry to foster collaborative innovation in compute platform designs.

2022

OCP Summit 2022: Open hardware for AI infrastructure

10/18/2022

This post announces Grand Teton, Meta's next-generation GPU-based hardware platform for AI at scale, which offers 4x host-to-GPU bandwidth, 2x compute and data network bandwidth, and 2x the power envelope compared to Zion. It also introduces Open Rack v3 (ORv3) with flexible power shelf placement and 48VDC output, and an improved battery backup unit. The post details advancements in cooling strategies, including Air-Assisted Liquid Cooling (AALC) and facility water cooling, and introduces Grand Canyon, a new HDD storage platform for AI infrastructure. Additionally, it announces the launch of the PyTorch Foundation under the Linux Foundation.

How thermal simulation helps optimize Meta’s data centers

9/14/2022

This post details the development and application of a dynamic thermal simulator for Meta's data centers. The simulator combines physics-based models (using equations for thermal processes and building-modeling languages like Modelica) with statistical data science to replicate cooling and thermal behavior. It enables testing of control policies (AI and model predictive control) and accurately predicts performance in extreme scenarios, including for unbuilt facilities. The simulator was validated against real-world data, including a severe winter storm in Texas, showing a mean absolute error of 0.5°F for supply air temperature over a 15-day period and 1.3°F across 24 random dates. Future research aims to couple these models with machine learning methods like reinforcement learning for real-time optimization of energy and water consumption, and to test new facility and equipment designs.

Transparent memory offloading: more memory at a fraction of the cost and power

6/20/2022

This post introduces Transparent Memory Offloading (TMO), Meta's data center solution for addressing the growing memory needs and high cost of DRAM. TMO utilizes cheaper, higher-capacity memory technologies like NVMe SSDs and compressed memory by transparently offloading less frequently accessed memory pages. It employs a new Linux kernel mechanism (PSI) to measure workload pressure and a userspace agent (Senpai) to dynamically adjust offloading, achieving significant memory savings (20-32%) across millions of servers and contributing to reduced cost and power consumption.

2021

OCP Summit 2021: Open networking hardware lays the groundwork for the metaverse

11/9/2021

This post details Meta's adoption of open networking hardware and the migration of its data center network to the Open Compute Project (OCP) Switch Abstraction Interface (SAI). It introduces the Wedge 400/400C next-generation Top-of-Rack (TOR) switches, highlighting their increased switching capacity (12.8 Tbps) and performance improvements over previous models. The post also describes the migration of FBOSS (Facebook Open Switching System) to support SAI, enabling easier integration with multiple ASIC vendors. Furthermore, it announces the deployment of 200G optics and next-generation 200G/400G fabric switches like Minipack2 and Arista 7388X5, which offer higher bandwidths (25.6 Tbps) and reduced power consumption per bit, laying the groundwork for future computing platforms like the metaverse.

Open-sourcing a more precise time appliance

8/11/2021

This post introduces the open-sourcing of a new Open Compute Time Appliance, including a PCIe Time Card that transforms commodity servers into precise time appliances. This innovation significantly improves the accuracy of Meta's timing infrastructure to 100 microseconds, enhancing data center management and distributed database performance. The post details the technical challenges and solutions involved in building this appliance, emphasizing its advantages over off-the-shelf alternatives in terms of security, reliability, and cost, and highlights the collaborative effort with the Open Compute Project community.

2020

How Facebook keeps its large-scale infrastructure hardware up and running

12/9/2020

This post details Meta's methodologies for maintaining high hardware availability in its large-scale data centers. It introduces systems like MachineChecker and FBAR for hardware failure detection and remediation, Cyborg for lower-level fixes, and a hybrid mechanism for minimizing performance impact during error reporting. Furthermore, it highlights the use of machine learning to predict repairs for undiagnosed/misdiagnosed tickets and an automated root cause analysis tool leveraging Scuba and FP-Growth for log analysis.

FioSynth: A representative I/O benchmark and data visualizer for data center workloads

11/18/2020

Introduced FioSynth, an open-source tool for automating storage workload execution and parsing results. FioSynth synthesizes production I/O traces to simulate diverse Facebook production services, focusing on latency outliers rather than peak IOPS. It uses fio, offers scalable workloads, parses results into CSV, supports client/server mode for parallel execution, optionally collects device health logs for write amplification factor calculation, and includes preconditioning and run cycle definition for repeatability. This tool aids in identifying storage performance inefficiencies and reproducing functional issues.

The next decade: How Facebook is stepping up the fight against climate change

9/15/2020

This post details Meta's updated climate change goals, committing to net zero emissions for its value chain by 2030. It outlines a strategy involving renewable energy, hyperefficient operations, supplier engagement, and carbon removal. The post also details the tracking of Scope 1, 2, and 3 emissions, the expansion of renewable energy procurement to power data centers, water stewardship initiatives, and the integration of circular economy principles into hardware and product lifecycles. It also highlights efforts to partner with suppliers to reduce their emissions and the development of a Climate Science Information Center to share reliable climate information.

Pcicrawler: A Python-based command-line interface tool to debug PCI issues at scale

8/5/2020

This post introduces Pcicrawler, a Python-based command-line tool designed to diagnose and debug Peripheral Component Interconnect (PCI) and PCI Express (PCIe) issues at scale. It provides detailed hardware information, visualizes PCI topology, and offers machine-parsable output for automation, contributing to the monitoring, diagnosis, and resolution of hardware-related performance and reliability issues within Meta's data centers.

Accelerometer and SoftSKU: Improving hardware platform performance for diverse microservices

5/11/2020

Introduced SoftSKU, a novel mechanism to optimize server processors for specific microservices by exploiting OS and hardware configuration knobs, achieving up to 7.2% performance improvement without new hardware. Developed Accelerometer, an analytical model to predict hardware acceleration gains with less than 3.7% error rate, aiding in informed hardware investment decisions.

2019

Systems @Scale Tel Aviv 2019 recap

12/20/2019

This post discusses advancements in scaling Facebook's data center infrastructure to support billions of users, focusing on technological innovations for reliability, scalability, and efficiency. It touches upon the physical infrastructure aspects of running large-scale services.

Facebook announces next-generation Open Rack frame

3/15/2019

This post announces the next-generation Open Rack frame, a collaborative initiative with Microsoft and the Open Compute Project (OCP) community. It addresses increasing power demands from AI and networking by introducing a new architecture based on Open Rack. Key technical contributions include enabling greater sharing between Microsoft and Facebook through a common OCP rack architecture, providing a flexible frame and power infrastructure for diverse solutions, and enabling new thermal solutions like liquid cooling manifolds and door-based heat exchangers. The post details the advantages of 48V power delivery over 12V for efficiency and handling high dI/dt, and outlines the goals for the Rack & Power project and Advanced Cooling Solutions subproject, including a converged rack frame, flexible power shelf, universal AC power interconnect, pluggable DC power shelf output interconnect, and battery backup systems.

Reinventing Facebook’s data center network

3/14/2019

This post details the reinvention of Facebook's data center network to address increasing demand and physical constraints. Key technical contributions include the F16 fabric design, which uses 100G CWDM4-OCP optics to achieve 4x capacity with fewer fabric chips and reduced power consumption compared to 400G alternatives. The development of the Minipack switch, a modular and flexible building block consuming 50% less power and space, is also highlighted, along with its contribution to OCP. The post also introduces HGRID, an evolution of Fabric Aggregator, designed to scale aggregation to six buildings per region and flatten the regional network for improved East-West traffic and uplink bandwidth.

Sharing a common form factor for accelerator modules

3/14/2019

This post introduces the OCP Accelerator Module (OAM) as a common form factor specification for hardware accelerators (GPUs, FPGAs, ASICs, etc.) used in AI and high-performance computing. It addresses the limitations of existing form factors like PCIe CEM by providing a more optimized solution for high bandwidth, interconnect flexibility, and power scalability. The OAM specification includes details on power input (12V/48V), TDP support, dimensions, link configurations, and support for various interconnect topologies like hybrid cube mesh and fully connected. Meta has contributed this specification to the Open Compute Project community to foster open collaboration and innovation.

Open-sourcing homomorphic hashing to secure update propagation

3/1/2019

This post introduces the application of homomorphic hashing (LtHash) to secure update propagation within Meta's configuration management systems, specifically the Location Aware Distribution system. It details the challenges of maintaining update integrity at scale in distributed systems and presents LtHash as an efficient solution that allows for updates to database hashes without recomputing the entire hash from scratch. The post also announces the open-sourcing of the LtHash implementation in the Folly library.

Opening our newest data center in Los Lunas, New Mexico

2/7/2019

This post announces the opening of Meta's new data center in Los Lunas, New Mexico. It highlights the facility's design for energy efficiency, evidenced by its LEED Gold certification. The post also details Meta's commitment to supporting the local grid with new wind and solar energy resources, contributing to the company's goal of operating on 100% renewable energy. This represents an expansion and operationalization of Meta's efforts in data center energy efficiency and renewable energy integration.

Rethinking data center design for Singapore

1/14/2019

This post details the design and construction of Meta's first custom-built data center in Asia, located in Singapore. Key technical contributions include: a vertical, 11-story building design to minimize land footprint in a dense urban environment; integration of 100% renewable energy sources, including rooftop solar arrays; and the implementation of the novel StatePoint Liquid Cooling (SPLC) system, developed with Nortek Air Solutions, which uses indirect evaporative cooling to reduce water and power consumption by over 20% in hot, humid climates. The design targets an annual PUE of 1.19, significantly lower than the regional average.

Data centers year in review

1/1/2019

This post details Meta's 2018 data center expansion and efficiency efforts, including the development of the Fabric Aggregator for network scalability, the implementation of the StatePoint Liquid Cooling (SPLC) system for energy-efficient cooling in diverse environments, and the design of the first multistory data center in Singapore. It also highlights the open-sourcing of StateService and sharing of FBOSS details.

2018

StatePoint Liquid Cooling system: A new, more efficient way to cool a data center

6/5/2018

Introduced the StatePoint Liquid Cooling (SPLC) system, a new indirect evaporative cooling technology developed with Nortek Air Solutions. The SPLC system uses water to cool data halls, operating in three modes (economizer, adiabatic, and super-evaporative) to optimize water and power consumption based on external conditions. It utilizes a patented membrane energy exchanger that allows water to evaporate through a membrane separation layer, cooling the water without cross-contamination. This system is more water-efficient than previous indirect cooling systems (reducing water usage by over 20% in hot/humid climates and almost 90% in cooler climates) and enables data center deployment in locations previously unsuitable for direct cooling due to environmental challenges like dust, humidity, or salinity. The SPLC system also offers flexibility in cooling delivery systems and reduces the required square footage for effective cooling.

2017

2017 Year in review: Data centers

12/11/2017

This post details the expansion of Meta's data center fleet in 2017 with four new locations and expansions at existing sites. It highlights the continued innovation in hardware, including the introduction of Bryce Canyon (storage chassis), Big Basin (AI hardware), Tioga Pass (CPU server), and Yosemite v2 (compute platform). A major technical achievement was the migration of the user database from InnoDB to MyRocks, which reduced server and storage requirements by half. The post also reiterates the commitment to powering data centers with 100% clean and renewable energy.

Hardware Analytics and Lifecycle Optimization (HALO) at Facebook

3/21/2017

This post introduces the Hardware Analytics and Lifecycle Optimization (HALO) system at Facebook. HALO is designed to monitor, process, and visualize data from millions of hardware components to improve global hardware health and inform future hardware design. It leverages Lumberjack for data collection (SMART attributes, RAID logs, network interface logs, etc.), integrates with asset and ticketing systems, and uses Facebook's Operational Data Store (ODS). Data processing involves analyzing failure rates and error categories at a fleet level, with tools like Attribot for identifying known hardware issues and Systemic Issue Detection (SID) for identifying emerging systemic hardware problems. The outcomes include improved global hardware health through cross-functional remediation efforts with vendors and internal teams, and design improvements based on field experience, such as toolless serviceability and airflow design adjustments.

OCP Summit 2017 — Facebook news recap

3/20/2017

This post details Meta's contributions to the Open Compute Project (OCP) Summit 2017, focusing on new hardware and software for data center infrastructure. Key contributions include Big Basin (next-gen AI GPU server with increased memory and arithmetic throughput for larger ML models), Bryce Canyon (high-density storage platform with improved HDD density and thermal/power efficiency), Tioga Pass (successor to Leopard compute platform with dual-CPU and OpenBMC support), Yosemite v2 (multi-node compute platform with hot-service capability), CWDM4-OCP specification (datacenter-focused optimization for 100G optical transceivers), Backpack (second-gen modular switch platform), and Wedge 100S (top-of-rack network switch with enhanced boot security and OpenBMC/FBOSS software). The post also highlights collaborations with Microsoft and Canonical to bring SAI and SONiC to the Wedge 100 platform, expanding the software ecosystem for open networking hardware.

The end-to-end refresh of our server hardware fleet

3/8/2017

This post details the end-to-end refresh of Meta's server hardware fleet, introducing new chassis designs: Bryce Canyon for high-density storage (72 HDDs in 4 OU, 20% higher density than Open Vault, improved thermal/power efficiency), Big Basin for AI GPU training (30% larger models, 100% throughput improvement over Big Sur, disaggregated GPU design), Tioga Pass for general compute (dual-socket, M.2 NVMe support, doubled PCIe bandwidth, 100G NIC), and Yosemite v2 (modular multi-node chassis with hot-service capability). These designs leverage Open Compute Project standards and technologies like OpenBMC for system management, aiming for increased efficiency, performance, and modularity.

Designing 100G optical connections

3/8/2017

This post details the development and deployment of 100G single-mode optical transceivers, a crucial step in upgrading data center network speeds. It highlights the challenges of moving to higher data rates, the decision to use single-mode fiber for future-proofing, and the optimization of transceiver specifications (CWDM4-OCP) to reduce power consumption and cost. The contribution to the Open Compute Project (OCP) is also emphasized, making this technology accessible to the wider industry.

The growing ecosystem around open networking hardware

1/24/2017

This post details the growing ecosystem around Meta's open networking hardware initiatives, such as Wedge 100 and Backpack. It highlights how various companies are building software and hardware solutions that integrate with these open designs, fostering collaboration and accelerating innovation in data center networking. This expands the scope of the feature thread to include the impact of open hardware on the broader networking industry and the collaborative development of infrastructure.

Sustainable materials in the data center

1/18/2017

This post details the technical evaluation and implementation of natural fiber-filled polypropylene (NFFPP) as a sustainable alternative to petroleum-based plastics (PCABS) in data center hardware. It covers material property adjustments (flammability rating, shrinkage), design considerations for fiber-filled composites (gate placement, draft angles, ribs), molding process tuning (temperature, drying), and cosmetic aspects. Specific examples of NFFPP integration into OCP hardware like Yosemite, Honey Badger, and Open Rack are provided, along with a comparative analysis of the upstream carbon footprint reduction.

2016

Introducing Backpack: Our second-generation modular open switch

11/8/2016

This post introduces Backpack, Meta's second-generation modular open switch platform, designed to support the migration to 100G data centers. It highlights the hardware design challenges, including improved cooling systems for higher power ASICs and optics, and a disaggregated architecture for scalability and thermal performance. The post also details the software stack (FBOSS, OpenBMC), testing methodologies, and the contribution of Backpack's specifications to the Open Compute Project, emphasizing its role in building a more efficient and scalable data center infrastructure.

Wedge 100: More open and versatile than ever

10/18/2016

This post details the Wedge 100, a second-generation top-of-rack network switch designed for 100G data centers. It highlights the hardware updates for improved serviceability, thermal management, and Open Rack compatibility, as well as software updates to the FBOSS stack to support the new platform and enhance operational flexibility. The post also announces the acceptance of the Wedge 100 specification into the Open Compute Project and discusses the broader hardware and software ecosystem, including commercial availability and third-party software integrations.

Facebook in Los Lunas: Our newest data center

9/14/2016

Announced the location of a new data center in Los Lunas, New Mexico, scheduled to come online in late 2018. This facility will be the seventh data center overall and the fifth US site. It will incorporate the latest Open Compute Project (OCP) compute, storage, and networking technologies and will be powered by 100-percent renewable energy, aiming for high efficiency and sustainability.

Facebook opens lab to others to validate infrastructure software

8/30/2016

This post details the creation of a lab space at Meta's headquarters dedicated to allowing software vendors to test their software on open hardware components developed by Meta and the Open Compute Project (OCP). This initiative aims to address the challenges of integrating commercial and open-source software with open hardware, thereby increasing the adoption of OCP technologies. The lab has already seen successful validation of software from Canonical and Red Hat on Meta's Leopard servers, Honey Badger storage enclosures, and Knox JBODs, demonstrating compatibility with MAAS, Juju, Gluster Storage, Ansible, CloudForms, and OpenStack.

Inside Facebook’s hardware labs: Moving faster with more collaboration

8/3/2016

This post details the creation and capabilities of 'Area 404', a new 22,000-square-foot collaborative hardware lab. It highlights the lab's state-of-the-art machine tools (9-axis mill-turn lathe, 5-axis vertical milling machine, 5-axis water jet, sheet metal shear and folder, CNC fabric cutter, CMM, electron microscope, CT scanner) and their application in accelerating hardware development cycles for infrastructure, Connectivity Lab, Oculus, and Building 8 teams. The post emphasizes the shift from isolated labs to a collaborative environment, reducing iteration time from weeks to days and enabling in-house modeling, prototyping, and failure analysis for projects like custom server racks, optical detectors, and VR camera rigs.

Open Compute Project U.S. Summit 2016 — Facebook news recap

3/10/2016

This post details Facebook's contributions and collaborations announced at the Open Compute Project U.S. Summit 2016. Key technical contributions include Lightning (an NVMe-based storage platform), OpenBMC (board management software), Yosemite (open server design), Wedge 100 and 6-pack (network switches), and Big Sur (open AI hardware platform). The post also highlights collaborations on 48V rack power distribution and optical interconnect standards, all aimed at improving data center hardware efficiency, flexibility, and performance.

Facebook’s new front-end server design delivers on performance without sucking up power

3/9/2016

This post details the redesign of Facebook's web servers to improve performance per watt. Key contributions include a collaborative effort with Intel to develop a new one-processor Xeon-D CPU (Broadwell-D) with a lower TDP (65W), and a redesigned server infrastructure (Mono Lake boards and Yosemite sleds) that doubles the number of CPUs per rack within the existing 11kW power budget. This approach eliminates NUMA issues and improves overall data center efficiency by achieving higher compute density without exceeding power constraints. The post also highlights the potential for LLC cache partitioning to further optimize memory bandwidth and performance.

Introducing Lightning: A flexible NVMe JBOF

3/9/2016

This post introduces Lightning, Meta's first NVMe JBOF (Just a Bunch of Flash) solution designed for disaggregated storage. It details the hardware components, including a PCIe retimer card, PCIe Expansion Board (PEB), and PCIe Drive Plane Board (PDPB), enabling end-to-end PCIe Gen 3 connectivity. The design addresses challenges like hot-plug support, management, signal integrity, external cabling, and power consumption, aiming to provide a scalable and flexible flash building block for data centers. The solution is contributed to the Open Compute Project.

OpenBMC: One board management software for all hardware at Facebook

3/9/2016

This post details the extension of OpenBMC to support Facebook's NVMe-based storage platform, Lightning. It describes the unique interface requirements for managing Lightning's BMC (USB, I2C, PCIe) due to the absence of a traditional NIC, and introduces new features like 'fscd' for fan control, 'log-util' for capturing secondary device logs, and 'flash-util' for reading SSD sideband status and managing the PCIe switch. The post highlights how this brings uniform board management across all Facebook data center hardware, simplifying provisioning and improving fault analysis.

Opening designs for 6-pack and Wedge 100

3/9/2016

This post details the opening of designs for the 6-pack and Wedge 100 network switches to the Open Compute Project (OCP). It highlights the evolution of Meta's data center networking from hierarchical systems to a data center fabric, with Wedge and 6-pack as key components. The post emphasizes the modularity, 100G connectivity support, and extensive hardware and network-level testing performed, showcasing a commitment to open hardware and collaborative development in advancing data center infrastructure efficiency.

Facebook in Ireland: Our newest data center

1/25/2016

Announces the construction of a new data center in Clonee, Ireland, which will be powered by 100% renewable energy. Highlights the use of OCP server and storage hardware, including Yosemite for compute and fabric, Wedge, and 6-pack in the network stack, as part of Meta's global infrastructure expansion.

2015

Open networking advances with Wedge and FBOSS

11/19/2015

This post details the advancement of Meta's open networking initiatives with the introduction and deployment of the Wedge network switch and FBOSS software. It highlights the disaggregation of networking hardware and software, enabling customization and faster innovation. The post discusses the operationalization of Wedge and FBOSS at scale, including custom deployment tools, monitoring strategies, and the implementation of nonstop forwarding. It also shares lessons learned from treating network switches like servers, focusing on protecting the CPU and control plane, and outlines future plans for scaling open networking solutions with Wedge 100 and 6-pack.

OpenBMC for server: Porting and supporting new features for “Yosemite”

8/14/2015

This post details the porting and enhancement of OpenBMC for the 'Yosemite' multi-node server platform. Key contributions include removing the insecure RMCP+ protocol, adding SSH and REST API support for improved management and security, implementing IPMI FRUID support, sensor monitoring, multi-node support for up to four server boards, Ethernet driver enhancements with NCSI, I2C driver secondary support, an IPMB framework for BMC and Bridge-IC communication, IPMI SDR support with caching, discrete signal monitoring, and new utilities for platform management. These advancements enhance the manageability, security, and efficiency of Meta's server infrastructure.

Using ISC Kea DHCP in our data centers

7/21/2015

This post details Meta's migration of its data center DHCP infrastructure from ISC dhcpd to ISC Kea. This migration was driven by the need for faster server provisioning and reduced downtime, especially at scale. By leveraging Kea's extensibility and hook points, Meta was able to create a stateless DHCP service that dynamically fetches configuration from its inventory system, simplifying deployment and maintenance, and ultimately contributing to faster hardware bootstrapping and repair times.

Facebook in Fort Worth: Our newest data center

7/7/2015

Announced the construction of a new data center in Fort Worth, Texas, which will be the fifth U.S. data center. This facility will incorporate the latest Open Compute server, storage, and network designs and will be powered by 100% renewable energy through a new wind deal.

Open Compute Project U.S. Summit 2015 – Facebook News Recap

3/10/2015

This post announces Facebook's contributions and announcements at the Open Compute Project U.S. Summit 2015. Key technical contributions include the introduction of 'Yosemite,' a new SoC compute server designed for high-powered microservers and disaggregated rack infrastructure; the proposal for the 'Wedge' top-of-rack network switch specification to the OCP Foundation, with efforts to create a product package for easier adoption; the release of 'OpenBMC,' an open software framework for system management; and the availability of the 'FBOSS Agent,' the software for the Wedge switch. The post also highlights significant cost and energy savings achieved through OCP and related efficiency work.

Introducing “Yosemite”: the first open source modular chassis for high-powered microservers

3/10/2015

This post introduces "Yosemite," an open-source modular chassis for high-powered microservers. It represents a shift from traditional scale-up server architectures to a scale-out approach using System-on-Chip (SoC) processors. The design focuses on modularity, power efficiency (targeting 65W TDP per SoC), and cost-effectiveness, contributing to Meta's ongoing efforts in data center energy efficiency and hardware optimization.

Facebook Open Switching System (“FBOSS”) and Wedge in the open

3/10/2015

This post announces the open-sourcing of FBOSS (Facebook Open Switching System) and Wedge, a top-of-rack switch specification. It details the disaggregation of network hardware and software, making switches more like servers for easier management and development. FBOSS is presented as a set of applications running on Linux, rather than a proprietary OS, and includes the FBOSS agent for programming ASICs and OpenBMC for system management. The contribution of Wedge to the Open Compute Project (OCP) signifies a move towards open networking standards and hardware, aiming to foster community development and broader adoption of open-source networking solutions. This directly supports the broader goal of optimizing data center infrastructure through open and flexible solutions.

Introducing “OpenBMC”: an open software framework for next-generation system management

3/10/2015

This post introduces OpenBMC, an open software framework for next-generation system management developed by Meta. It addresses the limitations of closed BMC software stacks by enabling rapid iteration, customization, and collaboration. OpenBMC is built as a flexible Linux distribution using the Yocto Project and has been deployed with Meta's Wedge switch. The post highlights the potential for community contributions to the common, SoC, and board layers, fostering innovation in system management applications and hardware support, and aligning with the Open Compute Project's goals.

Serving Facebook Multifeed: Efficiency, performance gains through redesign

3/10/2015

This post introduces the concept of 'disaggregation' as a strategy for optimizing data center infrastructure and improving efficiency and performance. It details how this concept was applied to redesign the Multifeed system, separating the CPU-intensive Aggregator component from the memory-intensive Leaf component. This disaggregation led to significant improvements in hardware utilization, scalability, performance (10% latency reduction), and reliability, with a notable 40% efficiency improvement in memory and CPU consumption. The post also highlights the potential for disaggregation in other services like Search and Operational Analytics.

2014

Favorite Hacks of 2014

12/30/2014

This post details the development of Autoscale, a power-efficient load balancing system created during a hackathon. Autoscale reduces CPU consumption through adaptable decision logic and smarter distribution, leading to significant power savings (10-15% on average for web clusters). It also mentions drone hacks using Android SDKs for experimental purposes and the @Scale app built with Parse for event navigation. The osquery tool, an open-source infrastructure insight tool, was also launched and enhanced through a hackathon challenge.

Introducing data center fabric, the next-generation Facebook data center network

11/14/2014

This post introduces the 'data center fabric,' a next-generation network architecture designed to replace traditional cluster-based designs. It details a disaggregated approach using server pods and fabric switches, connected via spine switches, to create a unified, high-performance network. The post highlights the benefits of this architecture, including rapid deployment, scalability, non-oversubscribed performance, and operational efficiency, achieved through standard BGP4 routing, layer3 connectivity, and a hybrid control plane. This represents a significant evolution in Meta's data center infrastructure, focusing on network performance and scalability to support exponential growth in machine-to-machine traffic.

Making Facebook’s software infrastructure more energy efficient with Autoscale

8/8/2014

Introduced Autoscale, a power-efficient load balancing system for Facebook web clusters. Autoscale dynamically adjusts the active server pool to concentrate workload onto a subset of servers during low-traffic periods, ensuring each active server maintains at least medium-level CPU utilization. This approach avoids running servers at low RPS, which is power-inefficient, and instead allows idle servers to be powered down or repurposed for batch processing. The system uses a feedback loop control mechanism with PI controllers and modeled relationships between CPU utilization and requests-per-second to determine the optimal active pool size. Preliminary results showed up to 27% power savings around midnight and an average of 10-15% over a 24-hour cycle in production web clusters.

Introducing “Wedge” and “FBOSS,” the next steps toward a disaggregated network

6/18/2014

This post introduces 'Wedge,' a disaggregated top-of-rack network switch, and 'FBOSS,' its Linux-based operating system. These projects represent a significant step in disaggregating network hardware and software, applying the same modular and open-source principles pioneered by the Open Compute Project (OCP) to networking. Wedge offers a modular hardware design leveraging server-like components for greater flexibility and integration into existing fleet management systems. FBOSS provides a Linux-based OS that allows network devices to be managed and controlled similarly to servers, enhancing visibility, automation, and control, and enabling faster implementation of forwarding software and hybrid control logic for optimized link utilization and faster failure recovery. This initiative aims to make network operations more efficient and scalable, aligning with Meta's broader goals of disaggregation and open hardware.

2011

Reflections on the Open Compute Summit

6/22/2011

This post details advancements in server hardware design, including new AMD and Intel motherboard designs that double compute density by using two narrow motherboards per chassis. It also introduces a new storage server platform with a variable compute-to-storage ratio, capable of supporting 50 hard drives per node. The post also provides an update on the Prineville data center's PUE of 1.077 and discusses the environmental and operational improvements planned for the new data center in Forest City, North Carolina, such as increased inlet temperatures and humidity. The formation of a non-profit foundation for the Open Compute Project is announced to encourage broader community collaboration and innovation.

How Project Triforce Prepared our Software Stack for Prineville

5/16/2011

This post details 'Project Triforce,' an internal initiative to simulate a third data center region using an existing cluster. This was crucial for testing and preparing Facebook's complex software stack for the new, custom-built data center in Prineville, Oregon, which was part of a broader strategy to build data centers from the ground up and share hardware designs through the Open Compute Project. It highlights the challenges of scaling infrastructure, software complexity, new configurations (like Flashcache with MySQL), and the need for rapid deployment, solved by tools like Kobold for automated provisioning.

Designing a Very Efficient Data Center

4/14/2011

This post details the design innovations of the Prineville data center, a key initiative within the Open Compute Project. It highlights the "less is more" philosophy leading to a facility with a PUE of 1.07 and WUE of 0.31 liters/kWH, a 45% reduction in CapEx, and improved reliability. Key innovations include eliminating centralized UPS (replacing with 48VDC at cabinet level), PDUs (using 277VAC distribution), chillers (using 100% outside air evaporative cooling), and ductwork (using dry wall supply air shafts). The post quantifies power loss reductions through these changes and details the custom reactor power panel and DC backup power system.

Inside the Open Compute Project Server

4/8/2011

This post details the development and launch of the Open Compute Project (OCP) server, a custom-designed hardware solution by Facebook engineers and partners. It outlines the technical specifications and design choices for the server's motherboard (supporting AMD and Intel CPUs, with a direct power supply interface and removed unnecessary features), power supply (achieving 94.5% efficiency with a novel 48VDC backup system to replace traditional UPS), chassis ('vanity-free' design for utility, ease of assembly, and weight reduction), rack ('triplet' design for efficient server density and networking port utilization, with serviceability features), battery cabinet (99.5% efficient DC backup system using 48VDC strings, replacing traditional UPS), and thermal design (taller chassis with larger fans for reduced energy consumption).

Building Efficient Data Centers with the Open Compute Project

4/7/2011

This post details the development of Meta's first custom-designed data center, including custom servers, power supplies, server racks, and battery backup systems. It highlights specific efficiency improvements such as a 480-volt electrical distribution system, removal of non-essential components, heat reuse, and elimination of a central UPS. The post also announces the formation of the Open Compute Project to share these innovations and best practices with the industry, aiming for collective advancement in data center efficiency.

2010

New Cooling Strategies for Greater Data Center Energy Efficiency

11/4/2010

This post details specific strategies implemented to improve data center energy efficiency in a second, smaller data center. Key technical contributions include: extending air-side economizer use, elevating rack inlet temperatures (from 51F to 67F), reducing excess air supply by replacing perforated tiles with solid ones to optimize underfloor static pressure, implementing cold aisle containment, and optimizing server fan speeds (reducing airflow and power consumption per server). These measures resulted in significant electricity savings and reduced carbon emissions.

Optimizing Data Center Energy Usage

10/20/2010

This post details specific engineering efforts to optimize data center energy usage. It outlines improvements in airflow distribution through cold aisle containment, mechanical sealing, and the optimization of server fan speed control. It also describes reducing cooling levels by shutting down CRAH units and raising set point temperatures for CRAH units and chilled water. These technical changes resulted in significant annual energy savings (2.5 million kWh) and cost reductions ($230,000).