
7/30/2026
What this post added
This post details CoreWeave's MLPerf® Training v6.0 records, specifically achieving industry-leading time-to-train for DeepSeek-V3 671B in 2.02 minutes using an 8,192-GPU NVIDIA GB300 NVL72 cluster with NVIDIA Spectrum-X Ethernet networking. It highlights efficient performance and scale across various cluster sizes (64 to 8,192 GPUs) and models (DeepSeek-V3 671B, Llama-3.1-405B, GPT-OSS-20B, Llama 3.1 8B). The post emphasizes the full-stack engineering approach, including NVLink-domain-aware scheduling in CoreWeave Kubernetes Service (CKS) and topology-aware workload placement in SUNK, deep networking optimizations, and fleet-wide performance consistency managed by CoreWeave Mission Control. It also details the use of NVIDIA NeMo Framework Release 26.04, full CUDA Graphs, and tuned parallelism strategies.