ML Inference Benchmarking
CoreWeave Leads MLPerf 0.7 Endpoints | CoreWeave Blog

CoreWeave Leads MLPerf 0.7 Endpoints | CoreWeave Blog

7/30/2026

What this post added

This post introduces CoreWeave's leading results in the MLPerf 0.7 Endpoints benchmark, showcasing their ability to achieve high throughput and efficiency with large language models like DeepSeek-R1 on NVIDIA GB200 NVL72 systems. It quantifies performance metrics such as sustained output tokens per second and tokens per second per GPU, and explains the significance of these metrics for inference services in terms of cost and user experience. The post also details the specific infrastructure components (bare metal, networking, CKS, SUNK, DFS) that contribute to these benchmark achievements, emphasizing that testing was conducted on production-ready infrastructure.

Read the original post ↗