AI Research and Development
Building Meta’s GenAI Infrastructure

Building Meta’s GenAI Infrastructure

3/12/2024 · Kevin Lee, Adi Gangidi, Mathew Oldham

What this post added

This post details the architecture and implementation of Meta's new 24k GPU AI clusters, designed for GenAI workloads like Llama 3 training. It covers hardware (Grand Teton, OpenRack), networking (RoCE with Arista 7800, InfiniBand with NVIDIA Quantum2), storage (Tectonic FUSE, Hammerspace NFS), and performance optimizations (network topology awareness, NCCL tuning, FP8 support, checkpointing). It also highlights efforts in debuggability (desync debug) and PyTorch evolution for large-scale training.

Read the original post ↗