ML Inference Benchmarking
CoreWeave Achieves New Record-Breaking AI Inferencing Benchmark with NVIDIA GB200 Grace Blackwell Superchips

CoreWeave Achieves New Record-Breaking AI Inferencing Benchmark with NVIDIA GB200 Grace Blackwell Superchips

4/2/2025

What this post added

This post announces CoreWeave's record-breaking AI inferencing benchmark results using NVIDIA GB200 Grace Blackwell Superchips, achieving 800 TPS on Llama 3.1 405B. It also reports a 40% throughput increase for Llama 2 70B on H200 GPUs compared to H100. The post emphasizes CoreWeave's role as the first cloud provider to offer GB200 NVL72 instances and highlights their prior achievements in deploying H100 and H200 GPUs, and demoing GB200 NVL72 systems.

Read the original post ↗