
5/21/2026
What this post added
This post introduces the integration of Slurm's topology/block plugin with NVIDIA GB200 NVL72 systems to enable topology-aware job scheduling. It explains how this alignment with NVL72 domain boundaries minimizes fragmentation and optimizes GPU occupancy. The post details how GB200 NVL72 supports larger job segment sizes (up to 18 nodes) for high I/O workloads like MoE training, and how flexible segment sizing benefits smaller jobs. It also presents scheduling recommendations based on simulations showing high GPU occupancy and utilization, emphasizing the importance of prioritizing large jobs with segment sizes that maximize NVLink domain usage and using smaller segment sizes for smaller jobs. Continuous monitoring and adjustment of segment sizes are highlighted as key for sustained performance.