BlogsCoreweaveAgentic Workflow Infrastructure

Agentic Workflow Infrastructure

Agentic Workflow Infrastructure

14
posts
2026

CoreWeave enhances its AI infrastructure for scaling complex agentic workloads by integrating serverless reinforcement learning (RL), production inference at scale, and fleet-wide observability. This post details the "superintelligence loop" where production data autonomously improves agent reliability over time. It highlights Serverless RL for post-training LLMs, CoreWeave Inference for reliable execution, and W&B Weave for end-to-end observability and signal extraction. W&B Skills and MCP serv. The platform now emphasizes the critical role of inference in production AI, detailing challenges in agentic inference such as GPU reservation, unpredictable I/O, high token generation, downstream dependencies, and concurrent agent coordination. It addresses the long lead times for GPU infrastructure and the systemic problems in moving from POC to production, including cold starts, traffic spikes, observability gaps, and cost drift. CoreWeave is building a full-stack infrastructure layer for AI inference, focusing on performance from bare-metal GPUs and high-speed networking to AI Object Storage with local GPU cache. The platform supports a migration path from Serverless Inference to Dedicated Inference and Inference on CKS, with framework-agnostic interfaces and included engineering support. The post also highlights the importance of accessing new GPU generations quickly to maintain a competitive advantage.

2026

Co-designed AI factories for production AI | CoreWeave Blog

7/31/2026

This post details the co-design and validation process with NVIDIA for AI factories, focusing on the integration of GPUs, CPUs, networking, and storage to achieve high performance and efficiency for production AI workloads. It highlights the NVIDIA DSX Platform, Exemplar Cloud validation for both training and inference, and the extension of co-design principles to network and storage layers (NVIDIA Spectrum-X Ethernet, NVIDIA Quantum InfiniBand, CoreWeave AI Object Storage with LOTA). It also discusses the rapid bring-up of new hardware generations (GB300 NVL72) and the integration with CoreWeave's Kubernetes Service (CKS), SUNK, and observability tools.

AI Agents Decode F1® Radio in Near Real Time | CoreWeave Blog

7/30/2026

This post details the application of CoreWeave's AI infrastructure to process real-time Formula 1® radio communications for Aston Martin Aramco. It describes the engineering challenge of managing 40 simultaneous audio streams with a sub-five-second reaction time, requiring a robust data orchestration and inference pipeline. The solution involves a fine-tuned transcription model (75 iterations, 7 hours of hand-annotated data) achieving production-grade accuracy, diarization, and topic categorization. The system leverages Weights & Biases Serverless Inference on CoreWeave Kubernetes Service with TensorRT-LLM for efficient inference. The platform enables near real-time analysis of competitor and internal team communications, transforming raw audio into actionable strategic insights.

Autonomous agents | CoreWeave Blog

7/30/2026

This post introduces CoreWeave's unified agentic AI capabilities, focusing on closing the loop between training and inference. It details the integration of serverless reinforcement learning (RL) for post-training LLMs, production inference at scale, and fleet-wide observability using W&B Weave. The "superintelligence loop" concept is explained, emphasizing how production data drives autonomous agent improvement. Specific technical aspects include Serverless RL's elastic scaling and per-token billing, CoreWeave Inference's predictable performance, and W&B Weave's signal extraction and evaluation APIs.

CoreWeave Sandboxes for Agentic AI | CoreWeave Blog

7/30/2026

Introduces CoreWeave Sandboxes as a new execution layer for agentic AI workloads, including RL, agent tool use, and model evaluation. Details two deployment options: on CoreWeave Kubernetes Service (CKS) clusters with per-sandbox Kubernetes pods governed by profiles, and a serverless runtime via Weights & Biases using Kata Containers for strong isolation. Highlights integrated debugging and observability through W&B Weave, correlating sandbox events with training metrics and LLM traces. Discusses efficient compute utilization by scheduling sandboxes across existing CoreWeave capacity, including idle CPUs on GPU nodes.

GLM 5.2 on CoreWeave Inference | CoreWeave Blog

7/30/2026

This post announces the availability of GLM 5.2 on CoreWeave Inference, emphasizing its performance and price-performance for agentic workflows and software engineering tasks. It details the engineering effort involved in optimizing the inference stack for GLM 5.2, including model-level tuning, GPU and networking optimization, and serving runtime improvements, to achieve fast inference speeds and cost-efficiency. The post also highlights the benefits of using open-weight models like GLM 5.2 for production workflows and the ease of transitioning from evaluation to production on CoreWeave's managed platform.

Infrastructure for Agentic Workflows | CoreWeave Blog

7/30/2026

This post defines the unique infrastructure requirements for agentic workflows, contrasting them with traditional inference workloads. It details how agentic applications involve multi-turn tool calls and sequential reasoning steps, leading to different latency, reliability, and cost considerations. The post outlines three core infrastructure needs: performance that holds across the entire chain, scalability that responds to bursty demand, and predictable economics for dynamic workloads.

Introducing NVIDIA HGX B300 on the Essential Cloud for AI

7/30/2026

This post announces the general availability of NVIDIA HGX B300 on CoreWeave Cloud, highlighting its doubled interconnect speed with NVIDIA Quantum-X800 InfiniBand networking, BlueField-3 DPUs, and 800 Gbps ConnectX-8 SuperNICs. It details benchmarks showing significant improvements in token generation and end-to-end request latency compared to HGX H200. The post emphasizes the HGX B300's increased memory capacity (270 GB), enhanced NVFP4 performance (14 PFLOPs), and system-level bandwidths for agentic AI workloads. It also details the integration of HGX B300 with CoreWeave Mission Control™ for rapid provisioning and monitoring, CoreWeave Kubernetes Service (CKS) and Slurm on Kubernetes (SUNK) for orchestration, and CoreWeave AI Object Storage for data access.

Production AI Runs on Inference | CoreWeave Blog

7/30/2026

This post shifts the focus to production-grade inference as the operational backbone of AI, particularly for agentic workflows. It highlights the challenges of scaling inference for production AI applications, emphasizing the need for predictable latency, reliability, and cost. The post introduces CoreWeave Inference with distinct deployment paths (Serverless and Dedicated) to address these needs, moving beyond just model deployment to operationalizing AI at scale. It discusses the evolution of inference demands from simple request-response to continuous reasoning and tool execution, and the importance of infrastructure, cost predictability, and control for mature workloads.

Why Inference Is the Defining Layer of AI | CoreWeave

7/30/2026

This post shifts the focus from AI training to AI inference, highlighting its growing importance and the unique challenges posed by agentic AI workloads. It details how agentic inference compounds complexity through persistent GPU reservation, unpredictable I/O, high token generation, downstream dependencies, and concurrent agent coordination. The post also addresses systemic issues in production inference, such as cold starts, traffic spikes, observability gaps, and cost drift, and positions CoreWeave's full-stack infrastructure approach as a solution, emphasizing bare-metal GPUs, high-speed networking, AI Object Storage, and a flexible platform migration path from prototype to production.

Announcing distributed AI on CoreWeave with fully managed Ray on Anyscale

5/20/2026

This post announces the integration of Anyscale, powered by Ray, into CoreWeave's platform via CoreWeave Kubernetes Service (CKS). It details how Anyscale provides a fully managed solution for distributed AI workloads, abstracting the complexity of Ray. Key benefits highlighted include enhanced performance and reliability through features like cluster validation, proactive health checking, and world-class observability. The post also contrasts Anyscale with self-managed Ray on CKS, emphasizing Anyscale's advantages such as Anyscale Workspaces for interactive development, fully managed Ray clusters with autoscaling, RayTurbo for performance optimizations, budget/cost dashboards, and SLA-backed support. Specific use cases like scaling RL workloads with RLlib and integrating with Weights & Biases are also mentioned.

Breaking the Bottlenecks: Scaling AI Without Stalling | CoreWeave Blog

5/20/2026

This post elaborates on the challenges of scaling AI infrastructure, including compute starvation, utilization inefficiencies, networking constraints, and storage drag. It critiques traditional scaling approaches as fragile and introduces strategies for building resilient AI infrastructure. Key strategies include designing for GPU elasticity, architecting for inevitable bursts, unifying observability, bringing compute closer to data, and designing for resilient multi-cloud/multi-AZ inference. It also discusses the benefits of purpose-built AI clouds and outlines future trends in AI infrastructure.

CoreWeave Mission Control: CoreWeave’s AI Operating Standard

5/20/2026

This post introduces two new capabilities to CoreWeave Mission Control: Telemetry Relay for enhanced transparency by forwarding audit and observability signals to SIEM/monitoring tools, and GPU Straggler Detection for deep bottleneck analysis. It also announces the preview of the CoreWeave Mission Control Agent, which offers conversational workflows for accessing telemetry and remediation guidance. The post emphasizes how these additions deepen transparency, improve bottleneck analysis, and streamline operational insights for AI workloads.

Scaling Reinforcement Learning with torchforge on CoreWeave Cloud

5/20/2026

This post introduces and details the integration of the torchforge Reinforcement Learning framework with CoreWeave's Slurm-on-Kubernetes (SUNK) infrastructure. It highlights how torchforge simplifies RL by separating algorithm design from distributed infrastructure, enabling researchers to scale complex RL workloads to thousands of GPUs. The post elaborates on the technical aspects of SUNK, including its ability to manage thousands of GPUs, ensure high availability, and scale compute nodes on demand, replacing the Slurm controller API to handle hundreds of thousands of jobs. It also details how SUNK's features like priorities, preemption, quotas, gang scheduling, and topology-aware scheduling address the challenges of running large-scale RL jobs, leading to faster job startup, more consistent cluster utilization, and higher end-to-end throughput. The post also touches upon the researcher-centric experience provided by SUNK, including secure isolated environments and IdP-federated cluster access.

Powering Production-Ready Agentic AI with RAG | CoreWeave Blog

1/21/2026

This post details the integration of vector databases (Milvus, Dragonfly, Pinecone) as a knowledge retrieval layer for agentic AI, enabling production-ready RAG by providing low-latency, scalable context retrieval alongside GPU-accelerated inference. It outlines two deployment patterns: running open-source vector databases (Milvus, Dragonfly) directly on CoreWeave Kubernetes Service (CKS) for maximum control and data locality, and pairing CoreWeave with managed services like Pinecone over high-performance cloud interconnects. The post also describes how this enables specific agent patterns like retrieval-augmented assistants, long-running agent workflows, and multimodal/recommendation-oriented agents.