ML Inference Benchmarking
CoreWeave Delivers Breakthrough AI Performance with NVIDIA GB200 and H200 GPUs in MLPerf Inference v5.0

CoreWeave Delivers Breakthrough AI Performance with NVIDIA GB200 and H200 GPUs in MLPerf Inference v5.0

5/20/2026

What this post added

This post introduces CoreWeave's MLPerf Inference v5.0 results, showcasing performance with NVIDIA GB200 Grace Blackwell Superchips (800 TPS on Llama 3.1 405B, 2.86X per-chip improvement over H200) and H200 GPUs (33,000 TPS on Llama 2 70B, 40% improvement over H100). It details the infrastructure optimizations enabling these results, including bare-metal Kubernetes, topology-aware scheduling with SUNK, and model loading acceleration with Tensorizer.

Read the original post ↗