
Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72 | NVIDIA Technical Blog
7/21/2026
This post details the world-record pre-training of DeepSeek-V3 671B on the NVIDIA GB300 NVL72 system, achieving 1,648 TFLOPs per GPU. It highlights the critical role of NVLink's fifth generation, providing 1.8 TB/s per-GPU bandwidth and 130 TB/s rack-level all-to-all bandwidth, enabling efficient MoE training. The post emphasizes the system's co-design across silicon, interconnect, networking (ConnectX-8 SuperNICs, Quantum-X800), and software (Megatron Core, TorchTitan, JAX), showcasing a 3x performance improvement over GB200 NVL72 and a 1.5x gain in six months due to software optimizations. It also mentions the integration of BlueField DPUs for infrastructure processing and the memory-semantic nature of NVLink for low-latency GPU-to-GPU communication.





