
3/12/2024 · Kevin Lee, Adi Gangidi, Mathew Oldham
What this post added
This post details the architecture and implementation of Meta's new 24k GPU AI clusters, designed for GenAI workloads like Llama 3 training. It covers hardware (Grand Teton, OpenRack), networking (RoCE with Arista 7800, InfiniBand with NVIDIA Quantum2), storage (Tectonic FUSE, Hammerspace NFS), and performance optimizations (network topology awareness, NCCL tuning, FP8 support, checkpointing). It also highlights efforts in debuggability (desync debug) and PyTorch evolution for large-scale training.