AI Object Storage Cost Optimization
CAIOS Achieves 7+ GB/s per GPU on NVIDIA Blackwell Ultra | CoreWeave

CAIOS Achieves 7+ GB/s per GPU on NVIDIA Blackwell Ultra | CoreWeave

5/20/2026

What this post added

This post details benchmark results of CoreWeave AI Object Storage (CAIOS) on 16 NVIDIA Blackwell Ultra GPU nodes, achieving an average throughput of 7+ GB/s per GPU. This was accomplished using the Warp S3 benchmarking tool on CoreWeave Kubernetes Service, with 10,000 50MB objects. Tests compared Ethernet (TCP) transport, which capped at 11.25 GB/s per node (2.81 GB/s/GPU), with RDMA (NVIDIA Quantum InfiniBand), which sustained 28.06 GB/s per node (7.02 GB/s/GPU). The post attributes the 3x performance increase over previous H100 benchmarks to: 1) fewer GPUs per node (4 vs 8), 2) the adoption of NVIDIA Quantum InfiniBand, and 3) LOTA pipeline optimizations yielding an approximate 17% improvement.

Read the original post ↗