4/2/2025
What this post added
This post announces CoreWeave's record-breaking AI inferencing benchmark results using NVIDIA GB200 Grace Blackwell Superchips, achieving 800 TPS on Llama 3.1 405B. It also reports a 40% throughput increase for Llama 2 70B on H200 GPUs compared to H100. The post emphasizes CoreWeave's role as the first cloud provider to offer GB200 NVL72 instances and highlights their prior achievements in deploying H100 and H200 GPUs, and demoing GB200 NVL72 systems.