BlogsNVIDIANVLink Scale-Up Networking for AI Factories

NVLink Scale-Up Networking for AI Factories

NVLink Scale-Up Networking for AI Factories

7
posts
2026

This feature thread tracks the evolution of NVIDIA NVLink, a purpose-built scale-up networking fabric for AI factories. It focuses on enabling high-bandwidth, low-latency GPU-to-GPU communication essential for large-scale AI workloads like MoE and LLMs. The thread covers advancements in NVLink generations, including the sixth generation with NVLink 6 Switch, which provides up to 3.6 TB/s per GPU and 260 TB/s rack-level bandwidth with in-network compute capabilities. It also details the extreme scale-up capabilities of the NVIDIA Vera Rubin platform, leveraging Groq 3 LPX LPUs with LPU C2C technology for deterministic, low-latency, high-throughput inference of trillion-parameter MoE models. This includes high-radix point-to-point links, compiler-scheduled data movement, and hardware-driven plesiosynchronous timing to enable thousands of LPUs to act as a single coherent system, addressing the unique demands of agentic AI workloads.

2026

Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72 | NVIDIA Technical Blog

7/21/2026

This post details the world-record pre-training of DeepSeek-V3 671B on the NVIDIA GB300 NVL72 system, achieving 1,648 TFLOPs per GPU. It highlights the critical role of NVLink's fifth generation, providing 1.8 TB/s per-GPU bandwidth and 130 TB/s rack-level all-to-all bandwidth, enabling efficient MoE training. The post emphasizes the system's co-design across silicon, interconnect, networking (ConnectX-8 SuperNICs, Quantum-X800), and software (Megatron Core, TorchTitan, JAX), showcasing a 3x performance improvement over GB200 NVL72 and a 1.5x gain in six months due to software optimizations. It also mentions the integration of BlueField DPUs for infrastructure processing and the memory-semantic nature of NVLink for low-latency GPU-to-GPU communication.

NVIDIA NVLink: The Scale-Up Network for AI Factories | NVIDIA Technical Blog

7/20/2026

This post introduces the sixth generation of NVIDIA NVLink, highlighting its role as the scale-up network for AI factories. It details the performance improvements (3.6 TB/s per GPU, 260 TB/s rack-level bandwidth, 130 TFLOPS in-network compute) and its superiority over Ethernet for MoE and LLM workloads, citing up to 2.3X decode throughput gains. The post emphasizes the extreme co-design of NVLink with the entire AI stack (hardware, software, libraries) and its contribution to features like disaggregated inference and expert parallelism. It also outlines key evaluation metrics for scale-up networking: delivered performance, factory resiliency, and platform maturity with a proven supply chain, underscoring NVLink's robust operational features and future expansion capabilities.

Enhancing Goodput in Large-Scale LLM Training with Nonuniform Tensor Parallelism | NVIDIA Technical Blog

7/6/2026

Introduces Nonuniform Tensor Parallelism (NTP) as an experimental framework to enhance Goodput in large-scale LLM training. NTP dynamically adjusts the tensor parallelism degree in response to transient GPU unavailability, preventing training stalls and throughput loss. It also proposes dynamic power boosting to compensate for performance loss in affected scale-up domains and employs efficient, overlapped tensor resharding techniques that minimize overhead to less than 1%. This enables resilient operation even as scale-up domains grow to 72 GPUs on NVIDIA Blackwell and Blackwell Ultra systems.

Maximize AI Factory Energy Efficiency Through Full-Stack Inference and Training Optimizations | NVIDIA Technical Blog

6/23/2026

This post details how NVIDIA maximizes AI factory energy efficiency through full-stack inference and training optimizations. It highlights system co-design with power, cooling, and infrastructure, and collaboration with ecosystem partners. Key technical contributions include the use of the NVIDIA GB200 NVL72 rack-scale system with direct-to-chip liquid cooling and in-rack power smoothing, NVIDIA DSX for dynamic power allocation and real-time telemetry, and the adoption of narrow precision formats like NVFP4 for improved throughput and energy efficiency. For LLM training, it discusses energy-aware techniques such as coordinated GPU speed tuning to minimize idle time and reduce energy consumption without increasing training time, and fine-grained profiling of kernel and phase-level energy usage in collaboration with the ML.ENERGY Initiative and Megatron-LM.

NVIDIA Blackwell Tops MLPerf Training 6.0 with Industry-Leading Scale and Performance | NVIDIA Technical Blog

6/16/2026

This post details the use of NVIDIA's scale-out networking platforms, specifically Spectrum-X Ethernet and Quantum InfiniBand, to achieve unprecedented scale and throughput in MLPerf Training 6.0. It highlights advanced adaptive routing and congestion control mechanisms within Spectrum-X to manage low-entropy, bursty traffic from MoE models, ensuring effective bandwidth near theoretical capacity and balancing tail latency. The post also showcases the integration of these networking capabilities with software optimizations like full-iteration CUDA graphs, CuTe DSL kernel fusions, MXFP8 attention blocks, and router/hybrid EP optimizations to achieve record-breaking training times for large MoE models across up to 8,192 Blackwell GPUs.

Unlock Exascale Performance on NVIDIA GB200 NVL72 with Slurm Topology-Aware Job Scheduling | NVIDIA Technical Blog

5/21/2026

This post introduces the integration of Slurm's topology/block plugin with NVIDIA GB200 NVL72 systems to enable topology-aware job scheduling. It explains how this alignment with NVL72 domain boundaries minimizes fragmentation and optimizes GPU occupancy. The post details how GB200 NVL72 supports larger job segment sizes (up to 18 nodes) for high I/O workloads like MoE training, and how flexible segment sizing benefits smaller jobs. It also presents scheduling recommendations based on simulations showing high GPU occupancy and utilization, emphasizing the importance of prioritizing large jobs with segment sizes that maximize NVLink domain usage and using smaller segment sizes for smaller jobs. Continuous monitoring and adjustment of segment sizes are highlighted as key for sustained performance.

How the NVIDIA Vera Rubin Platform is Solving Agentic AI’s Scale-Up Problem | NVIDIA Technical Blog

5/14/2026

This post introduces the NVIDIA Vera Rubin platform, highlighting its solution to agentic AI's scale-up problem through the integration of NVIDIA Groq 3 LPX LPUs with LPU C2C technology. It details how LPU C2C achieves deterministic, low-latency, high-throughput inference for trillion-parameter MoE models by employing high-radix point-to-point links, compiler-scheduled data movement, and hardware-driven plesiosynchronous timing. This enables thousands of LPUs to operate as a single coherent system, addressing the unique demands of agentic workloads that require predictable scale-up networking.