BlogsCoreweaveLiquid Cooling Infrastructure

Liquid Cooling Infrastructure

Liquid Cooling Infrastructure

8
posts
2026

CoreWeave enhances its AI infrastructure by deploying NVIDIA Spectrum-X SN6600-LD liquid-cooled Ethernet switches, doubling per-rack network bandwidth to 1.64 Pb/s. This enables a non-blocking, multi-plane, multi-rail spine and leaf fabric for Vera Rubin NVL72 GPUs, reducing network latency and keeping GPU idle time near zero. The deployment integrates with CoreWeave Mission Control's Valvey and Racky for unified thermal management and operational control, and the Rack Lifecycle Controller (RLCC) for automated rack management. This upgrade results in a 100% increase in total switching performance per rack, a 62.5% reduction in rack footprint, and a 30% improvement in power efficiency compared to previous air-cooled solutions.

2026

A Deep Dive on CoreWeave Innovations | CoreWeave Blog

7/30/2026

This post introduces Valvey, a patent-pending programmable per-rack liquid-cooling valve assembly that transforms liquid cooling into a software-defined control surface, enabling per-rack isolation and faster failure containment. It also details Racky, a unified rack manager that aggregates sensors and provides a standardized, software-addressable control surface for each rack, working in concert with Valvey and the Rack LifeCycle Controller. The deployment of the liquid-cooled NVIDIA Spectrum-X SN6600 Ethernet Switch is highlighted as a key component for stitching NVL72 racks into a unified cluster, with tight integration of power and cooling management around the switch to maintain peak performance. The post also elaborates on the multi-rail, multi-plane networking fabric supporting both InfiniBand and Spectrum-X Ethernet with RDMA over Converged Ethernet (RoCE), detailing the use of NVIDIA ConnectX-9 SuperNICs for high-bandwidth, low-latency connectivity.

CoreWeave Brings Liquid Cooled Switching to AI Networking | CoreWeave Blog

7/30/2026

This post details the deployment of NVIDIA Spectrum-X SN6600-LD liquid-cooled Ethernet switches, highlighting their 102.4 Tb/s capacity and 1.6 Tb/s ports. It explains how this enables a non-blocking fabric for Vera Rubin NVL72 GPUs, leading to improved MFU and lower inference latency. The post also describes the integration of these switches into CoreWeave Mission Control, extending Valvey and Racky for unified thermal management and operational control, and the Rack Lifecycle Controller (RLCC) for automated rack management. It quantifies the benefits as a 100% increase in switching performance, a 62.5% reduction in rack footprint, and a 30% improvement in power efficiency.

First-Ever Measured Vera Rubin NVL72 Silicon Performance Stats | CoreWeave Blog

7/23/2026

This post provides the first measured silicon performance statistics for the NVIDIA Vera Rubin NVL72 on CoreWeave, demonstrating a 10x improvement in tokens per megawatt for the DeepSeek R1 inference workload compared to NVIDIA GB200 NVL72. It details the architectural advantages of Vera Rubin NVL72, including its GPU count, CPU count, NVLink fabric bandwidth, and NVFP4 support, and highlights the impact of these on inference performance and cost-efficiency for agentic AI workloads. The post also notes that this is an initial measurement and further optimizations are expected.

5/20/2026

This post details the performance gains and architectural advantages of deploying NVIDIA GB300 NVL72 instances with NVIDIA Blackwell Ultra GPUs for AI inference. It quantifies a 6.5x performance improvement on the DeepSeek R1 model by comparing a 16-GPU H100 system to a 4-GPU GB300 system, attributing the uplift to the GB300's superior memory and interconnect bandwidth, enabling 4-way Tensor Parallelism instead of 16-way. The post also outlines CoreWeave's infrastructure enhancements, such as a topology-aware scheduler and automated Rack LifeCycle Controller, which optimize the utilization of the GB300's NVLink and networking capabilities for AI workloads.

CoreWeave Leads the Way with First NVIDIA GB300 NVL72 Deployment

5/20/2026

This post announces the first deployment of NVIDIA GB300 NVL72 instances on CoreWeave, detailing the hardware specifications (72 Blackwell Ultra GPUs, 36 Grace CPUs, 18 BlueField-3 DPUs), performance improvements for AI reasoning (up to 10x user responsiveness, 5x throughput per watt, 50x output for reasoning model inference), and enhanced memory (21TB HBM3e per rack). It also highlights platform-level optimizations like RLCC, Cabinet Wrangler, Cabinet Details dashboard, and integration with Weights & Biases for infrastructure observability. The post emphasizes the collaboration with Dell, Switch, and Vertiv for this deployment.

Why ClusterMAX 2.0 Validates CoreWeave’s Engineering

5/20/2026

This post elaborates on CoreWeave's engineering approach to AI infrastructure, as validated by the ClusterMAX 2.0 report. It details specific technical implementations including the SUNK framework for job queuing and resource management, a custom Rack LifeCycle Controller for managing GB200/GB300 systems as unified objects, the use of NVIDIA BlueField DPUs for workload isolation and performance, end-to-end monitoring leveraging DCGM, NVML, and interconnect telemetry, AI Object Storage with LOTA caching for data proximity, and a zero-trust security architecture. It also emphasizes the direct-to-expert support model.

Why Liquid Cooling Matters for AI Scale | CoreWeave

5/20/2026

This post introduces CoreWeave's strategic shift towards liquid cooling for AI infrastructure. It quantifies the heat output of AI server racks, contrasts it with traditional cooling methods, and explains why liquid cooling is becoming essential for high-density deployments like the NVIDIA GB200 NVL72. The post details the benefits for customers, including denser GPU deployments, access to newer hardware, and improved power utilization, and outlines the operational approach for implementing liquid cooling in new data centers, including the continued role of air cooling for less intensive components.

Liquid Cooling for AI Data Centers | CoreWeave Blog

5/4/2026

This post elaborates on CoreWeave's adoption of liquid cooling for AI data centers, focusing on direct-to-chip cooling for AI equipment and closed-loop systems for coolant recycling. It also introduces the use of AI agents for proactive thermal management and highlights the sustainability benefits of these approaches. The post frames liquid cooling as a critical architectural pillar for performance and reliability in AI infrastructure.