Blogs›Coreweave Feature Trails
See how major capabilities shipped, upgraded, and evolved across Coreweave's engineering blog.
Publishing pulse
2025–2026 · peak 2026
83 posts mapped

CoreWeave enhances its AI infrastructure for scaling complex agentic workloads by integrating serverless reinforcement learning (RL), production inference at scale, and fleet-wide observability. This post details the "superintelligence loop" where production data autonomously improves agent reliability over time. It highlights Serverless RL for post-training LLMs, CoreWeave Inference for reliable execution, and W&B Weave for end-to-end observability and signal extraction. W&B Skills and MCP serv. The platform now emphasizes the critical role of inference in production AI, detailing challenges in agentic inference such as GPU reservation, unpredictable I/O, high token generation, downstream dependencies, and concurrent agent coordination. It addresses the long lead times for GPU infrastructure and the systemic problems in moving from POC to production, including cold starts, traffic spikes, observability gaps, and cost drift. CoreWeave is building a full-stack infrastructure layer for AI inference, focusing on performance from bare-metal GPUs and high-speed networking to AI Object Storage with local GPU cache. The platform supports a migration path from Serverless Inference to Dedicated Inference and Inference on CKS, with framework-agnostic interfaces and included engineering support. The post also highlights the importance of accessing new GPU generations quickly to maintain a competitive advantage.
Timeline

CoreWeave is enhancing its AI infrastructure security by integrating NVIDIA BlueField-4 DPUs. This provides hardware-enforced tenant isolation using EVPN VXLAN, offloads networking and security tasks from CPUs to DPUs, and extends the zero-trust infrastructure model with finer-grained, hardware-based isolation via the BlueField Advanced Secure Trusted Resource Architecture. This aims to deliver extreme performance, scale, next-gen security, and maximum efficiency for AI workloads.
Timeline

CoreWeave is enhancing its AI infrastructure by providing detailed guidance on evaluating the Total Cost of Ownership (TCO) for AI workloads. This involves a holistic approach that goes beyond simple GPU hour pricing to consider the full stack of compute, storage, networking, and orchestration. The focus is on translating GPU efficiency into economic efficiency through metrics like Model FLOPs Utilization (MFU) and Goodput, aligning storage architecture with GPU throughput requirements, and prioritizing cost-effective pricing models based on workload patterns. This post specifically addresses the 'token pricing illusion' by introducing the concept of a 'useful token' and advocating for cost-per-useful-token as a more accurate metric for certain workloads, while also identifying scenarios where token pricing remains optimal.
Timeline

CoreWeave enhances its cloud solutions to cater to specific customer needs, focusing on scalability, security, and high-performance compute for demanding workloads. This involves providing tailored infrastructure, direct connections, dedicated storage, and upgrade options to enable customers to design, execute, and scale proprietary models efficiently. The goal is to deliver timely access to cutting-edge technologies in a fully managed fashion, leading to rapid deployment, increased compute. The Conductor platform is expanded to support LoRA training and inference for image generation, with APIs for integration into asset management and artist tools. Video model support and other cloud infrastructures are planned.
Timeline
.jpg)
CoreWeave enhances its AI infrastructure for large-scale model training by focusing on optimizing NVIDIA GPU clusters for both performance and reliability. This involves a multi-layered approach including bare-metal hardware, a dual-fabric network architecture (InfiniBand for compute, Ethernet for storage), topology-aware scheduling with automated node eviction and job rescheduling via SUNK (Slurm on Kubernetes), and optimized asynchronous checkpointing using Tensorizer. The SUNK system is designed to mitigate issues like stragglers, synchronization stalls, and stalled GPUs by providing better visibility into per-rank performance, communication bottlenecks, and data pipeline issues, aiming to reduce variance and prevent coordination failures that lead to wasted compute.
Timeline
.jpg)
CoreWeave enhances its AI infrastructure by deploying NVIDIA Spectrum-X SN6600-LD liquid-cooled Ethernet switches, doubling per-rack network bandwidth to 1.64 Pb/s. This enables a non-blocking, multi-plane, multi-rail spine and leaf fabric for Vera Rubin NVL72 GPUs, reducing network latency and keeping GPU idle time near zero. The deployment integrates with CoreWeave Mission Control's Valvey and Racky for unified thermal management and operational control, and the Rack Lifecycle Controller (RLCC) for automated rack management. This upgrade results in a 100% increase in total switching performance per rack, a 62.5% reduction in rack footprint, and a 30% improvement in power efficiency compared to previous air-cooled solutions.
Timeline
.jpg)
CoreWeave continues to enhance its AI infrastructure for scaling complex workloads, focusing on production inference and training performance. This post details the selection of NVIDIA GPUs for inference workloads, mapping specific GPU platforms (GB300 NVL72, GB200 NVL72, HGX B300, HGX B200, HGX H200, HGX H100, RTX PRO 6000 Blackwell Server Edition) to common inference patterns based on model size, context window, concurrency, batching behavior, latency SLOs, and deployment shape. It highlights the optimization of Kimi K2.7 Code using NVFP4 quantization and a DFlash speculative decoder on Blackwell GPUs, achieving leading price-performance. The optimization process involved mirroring weights to CoreWeave AI Object Storage with LOTA, baseline measurements with various benchmarks, custom 3-pass calibration for NVFP4 quantization using NVIDIA Model-Optimizer, and training a DFlash speculative decoding model with a modified D-PACE loss function. This configuration is deployed on CoreWeave Inference via vLMM by default, providing customers with immediate throughput and cost benefits.
Timeline

CoreWeave ARENA provides a practical, production-like environment for evaluating AI workloads, moving beyond theoretical benchmarks to real-world performance, scaling, and cost analysis. It integrates with existing tools like Weights & Biases and CoreWeave Mission Control for operational visibility, supporting Kubernetes-native and Slurm-based orchestration. This allows teams to make informed infrastructure decisions with verifiable evidence before committing to production deployments. This post demonstrates the ability to run 10,000 Large-Eddy Simulations (LES) for CFD in 32 hours on CoreWeave's GPU cloud infrastructure, achieving NASA's CFD Vision 2030 target four years early. This showcases the platform's capability for high-throughput, high-fidelity engineering simulations.
Timeline

CoreWeave AI Object Storage (CAIOS) continues to enhance its performance and cost-efficiency for AI workloads. Recent benchmarks on NVIDIA Blackwell Ultra GPU nodes, utilizing NVIDIA Quantum InfiniBand and LOTA pipeline optimizations, have achieved over 7 GB/s per GPU throughput. This represents a significant improvement over previous benchmarks on H100 nodes, driven by architectural changes (fewer GPUs per node), network fabric upgrades (InfiniBand), and ongoing LOTA optimizations. CAIOS aims to provide fast data access, quick recovery and resiliency, scalability, and airtight security for GenAI workloads. This post details the Local Object Transport Accelerator (LOTA) which provides a direct path between GPUs and data by bypassing object storage gateways and indexes, and transparently caching data on local compute node storage. It also highlights performance metrics like up to 2 GB/s per GPU throughput and 25 GB/s per 1 PB of reserved storage, 99.9% uptime, and eleven nines of durability. Security features include encryption at rest and in transit, and identity access management with MFA and role-based access controls. The solution is designed to scale to hundreds of thousands of GPUs.
Timeline

CoreWeave enhances its AI infrastructure with Flexible Capacity Plans, introducing Flex Reservations and Spot instances to better align with fluctuating AI workload demands. Flex Reservations offer guaranteed access up to a defined ceiling with a lower holding fee and usage-based pricing, addressing overprovisioning. Spot instances provide lower-cost, interruptible compute for fault-tolerant or batch workloads, with enhanced preemption signaling and notice windows. This portfolio approach aims to provide predictable capacity for AI innovation and production inference.
Timeline