5/20/2026
What this post added
This post details benchmark results demonstrating CoreWeave AI Object Storage (CAIOS) delivering over 2 GB/s per GPU throughput. It highlights the LOTA™ (Local Object Transport Accelerator) feature, which caches data on local NVMe disks within GPU nodes to reduce latency. Benchmarking methodology involved 20 GPU nodes performing read and write operations using the Warp S3 benchmarking tool. Results showed aggregate read throughput reaching 368 GiB/s, or 18.4 GiB/s per GPU node, with performance scaling to any number of GPUs. The post also discusses the real-world impact on AI workflows and future plans to evaluate write performance.