5/20/2026
What this post added
This post introduces CoreWeave's MLPerf Inference v5.0 results, showcasing performance with NVIDIA GB200 Grace Blackwell Superchips (800 TPS on Llama 3.1 405B, 2.86X per-chip improvement over H200) and H200 GPUs (33,000 TPS on Llama 2 70B, 40% improvement over H100). It details the infrastructure optimizations enabling these results, including bare-metal Kubernetes, topology-aware scheduling with SUNK, and model loading acceleration with Tensorizer.