5/20/2026
What this post added
This post details CoreWeave's achievement of new industry records in MLPerf Training v5.0 benchmarks, showcasing the performance of their largest-ever NVIDIA Blackwell GPU cluster. It highlights specific performance gains, such as training Llama 3.1 405B in 27.3 minutes, more than twice as fast as comparable Hopper GPU clusters. The post elaborates on the underlying infrastructure components that contribute to this performance, including AI-optimized data centers, performance-optimized GPU instances, bare metal access, high-performance storage (up to 2GB/s/GPU), NVIDIA Quantum InfiniBand and BlueField DPUs, SUNK (Slurm on Kubernetes), CoreWeave Kubernetes Service (CKS), and Mission Control for observability and uptime. It quantifies customer benefits like up to twice as fast training speeds, 20% higher performance on like-for-like clusters, and a 14% improvement in price-adjusted performance.