Internet Latency and Network Performance Analysis
RoCE networks for distributed AI training at scale

RoCE networks for distributed AI training at scale

8/5/2024 · Adi Gangidi, James Hongyi Zeng

What this post added

This post details Meta's development and operation of large-scale RoCE networks for distributed AI training. It describes the network topology, including dedicated frontend and backend networks, AI Zones with two-stage Clos topologies, and aggregator training switches for inter-building connectivity. The post also discusses routing evolution, moving from ECMP and path pinning to Enhanced ECMP with queue pair scaling to address low entropy and burstiness in AI workloads. Finally, it outlines the shift in congestion control from DCQCN to receiver-driven traffic admission for 400G deployments, highlighting the use of collective libraries and RoCE transport.

Read the original post ↗