5/20/2026
What this post added
This post details the performance gains and architectural advantages of deploying NVIDIA GB300 NVL72 instances with NVIDIA Blackwell Ultra GPUs for AI inference. It quantifies a 6.5x performance improvement on the DeepSeek R1 model by comparing a 16-GPU H100 system to a 4-GPU GB300 system, attributing the uplift to the GB300's superior memory and interconnect bandwidth, enabling 4-way Tensor Parallelism instead of 16-way. The post also outlines CoreWeave's infrastructure enhancements, such as a topology-aware scheduler and automated Rack LifeCycle Controller, which optimize the utilization of the GB300's NVLink and networking capabilities for AI workloads.