BlogsCoreweaveAI Object Storage Cost Optimization

AI Object Storage Cost Optimization

AI Object Storage Cost Optimization

8
posts
2026

CoreWeave AI Object Storage (CAIOS) continues to enhance its performance and cost-efficiency for AI workloads. Recent benchmarks on NVIDIA Blackwell Ultra GPU nodes, utilizing NVIDIA Quantum InfiniBand and LOTA pipeline optimizations, have achieved over 7 GB/s per GPU throughput. This represents a significant improvement over previous benchmarks on H100 nodes, driven by architectural changes (fewer GPUs per node), network fabric upgrades (InfiniBand), and ongoing LOTA optimizations. CAIOS aims to provide fast data access, quick recovery and resiliency, scalability, and airtight security for GenAI workloads. This post details the Local Object Transport Accelerator (LOTA) which provides a direct path between GPUs and data by bypassing object storage gateways and indexes, and transparently caching data on local compute node storage. It also highlights performance metrics like up to 2 GB/s per GPU throughput and 25 GB/s per 1 PB of reserved storage, 99.9% uptime, and eleven nines of durability. Security features include encryption at rest and in transit, and identity access management with MFA and role-based access controls. The solution is designed to scale to hundreds of thousands of GPUs.

2026

AI storage and LLMs: 4 critical needs to look out for

5/20/2026

This post introduces and details the Local Object Transport Accelerator (LOTA) as a key component of CoreWeave AI Object Storage. LOTA provides a direct, accelerated path between GPUs and data repositories by bypassing object storage gateways and indexes, and includes transparent local caching on compute nodes. It also quantifies performance benefits such as up to 2 GB/s per GPU throughput and 25 GB/s per 1 PB of reserved storage, and emphasizes security features like encryption and IAM. The post outlines four critical needs for AI storage: fast data access, quick recovery and resiliency, scalability, and airtight security, and positions CAIOS as fulfilling these needs.

AI Storage Without Limits: Exploring the Latest CoreWeave AI Object Storage Expansions | Blog

5/20/2026

This post details the expansion of CoreWeave AI Object Storage with cross-region and multi-cloud flexibility, enabling data access from anywhere in CoreWeave's regions, other clouds, and on-premises environments. It highlights the elimination of dataset replication needs, the reduction of data divergence risks, and the maintenance of local disk performance. The networking backbone supporting this includes private interconnects, direct cloud peering, and cross-region networking up to 400 Gbps. The post also reiterates the performance benefits of LOTA technology and the continued absence of egress, ingress, or request fees.

Break Free From Hyperscaler Storage Constraints | CoreWeave

5/20/2026

This post details CoreWeave's AI Object Storage architecture, emphasizing its design for GPU performance with an InfiniBand backbone and LOTA caching, delivering up to 7 GB/s per GPU. It introduces a unified data plane for global access, zero egress fees, and usage-based pricing with adaptive tiers (hot, warm, cold) for cost optimization. The post also highlights the Zero Egress Migration (0EM) program for data migration.

CAIOS Achieves 7+ GB/s per GPU on NVIDIA Blackwell Ultra | CoreWeave

5/20/2026

This post details benchmark results of CoreWeave AI Object Storage (CAIOS) on 16 NVIDIA Blackwell Ultra GPU nodes, achieving an average throughput of 7+ GB/s per GPU. This was accomplished using the Warp S3 benchmarking tool on CoreWeave Kubernetes Service, with 10,000 50MB objects. Tests compared Ethernet (TCP) transport, which capped at 11.25 GB/s per node (2.81 GB/s/GPU), with RDMA (NVIDIA Quantum InfiniBand), which sustained 28.06 GB/s per node (7.02 GB/s/GPU). The post attributes the 3x performance increase over previous H100 benchmarks to: 1) fewer GPUs per node (4 vs 8), 2) the adoption of NVIDIA Quantum InfiniBand, and 3) LOTA pipeline optimizations yielding an approximate 17% improvement.

CoreWeave AI Object Storage Delivers 2+ GB/s per GPU

5/20/2026

This post details benchmark results demonstrating CoreWeave AI Object Storage (CAIOS) delivering over 2 GB/s per GPU throughput. It highlights the LOTA™ (Local Object Transport Accelerator) feature, which caches data on local NVMe disks within GPU nodes to reduce latency. Benchmarking methodology involved 20 GPU nodes performing read and write operations using the Warp S3 benchmarking tool. Results showed aggregate read throughput reaching 368 GiB/s, or 18.4 GiB/s per GPU node, with performance scaling to any number of GPUs. The post also discusses the real-world impact on AI workflows and future plans to evaluate write performance.

Expanding CoreWeave AI Object Storage: Unified Dataset, Lower Costs, and Unmatched Scale

5/20/2026

Introduces unified global datasets accessible across regions, clouds, and on-prem, leveraging Local Object Transport Accelerator (LOTA) technology for accelerated performance (7 GB/s per GPU). Implements new automated usage-based pricing tiers (Hot, Warm, Cold) to reduce storage costs by over 75% for AI workloads, eliminating manual data management and access fees.

Introducing our Zero Egress Migration program | CoreWeave Blog

5/20/2026

Introduces the Zero Egress Migration ([0]EM) program, a service to facilitate large-scale data transfers from other cloud providers to CoreWeave, covering egress fees and offering zero egress fees from CoreWeave. The program orchestrates migrations at petabyte scale (over 1 PB/day) with end-to-end checksum validation and provides a real-time migration dashboard for transparency. This complements the existing AI Object Storage capabilities by addressing data mobility challenges.

Lower AI Storage Costs by 75% | CoreWeave

5/20/2026

This post details the technical implementation of CoreWeave AI Object Storage's automated usage-based billing. It explains how real-time access tracking is used to seamlessly transition data between three pricing tiers (Hot, Warm, Cold) based on inactivity periods (7 days for Hot to Warm, 30 days for Warm to Cold). The system ensures no rehydration delays, retrieval fees, egress fees, or request fees, maintaining line-rate performance across all tiers. The post also includes a competitive cost analysis comparing CoreWeave's approach to traditional hyperscaler object storage solutions, highlighting the benefits of its simplified, transparent, and cost-effective model for AI workloads.