Agentic AI Infrastructure Acceleration with BlueField DPUs
Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72 | NVIDIA Technical Blog

Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72 | NVIDIA Technical Blog

7/21/2026

What this post added

This post details a world record set for pre-training the DeepSeek-V3 671B Mixture of Experts (MoE) model on the NVIDIA GB300 NVL72 system. It highlights the critical role of MoE architectures in pushing frontier AI capabilities and the associated communication challenges. The post emphasizes how the GB300 NVL72, through extreme co-design of silicon, interconnect (fifth-generation NVLink), networking (ConnectX-8 SuperNICs, Quantum-X800 InfiniBand/Spectrum-X Ethernet), and software (Megatron Core, TorchTitan, JAX), addresses these challenges. It showcases a 3x performance improvement over previous generations (GB200 NVL72) and a 1.5x gain in six months due to software optimizations, underscoring the continuous evolution of AI infrastructure for large-scale training.

Read the original post ↗