
7/21/2026
What this post added
This post details a world record set for pre-training the DeepSeek-V3 671B Mixture of Experts (MoE) model on the NVIDIA GB300 NVL72 system. It highlights the critical role of MoE architectures in pushing frontier AI capabilities and the associated communication challenges. The post emphasizes how the GB300 NVL72, through extreme co-design of silicon, interconnect (fifth-generation NVLink), networking (ConnectX-8 SuperNICs, Quantum-X800 InfiniBand/Spectrum-X Ethernet), and software (Megatron Core, TorchTitan, JAX), addresses these challenges. It showcases a 3x performance improvement over previous generations (GB200 NVL72) and a 1.5x gain in six months due to software optimizations, underscoring the continuous evolution of AI infrastructure for large-scale training.