Agentic Workflow Infrastructure
Scaling Reinforcement Learning with torchforge on CoreWeave Cloud

Scaling Reinforcement Learning with torchforge on CoreWeave Cloud

5/20/2026

What this post added

This post introduces and details the integration of the torchforge Reinforcement Learning framework with CoreWeave's Slurm-on-Kubernetes (SUNK) infrastructure. It highlights how torchforge simplifies RL by separating algorithm design from distributed infrastructure, enabling researchers to scale complex RL workloads to thousands of GPUs. The post elaborates on the technical aspects of SUNK, including its ability to manage thousands of GPUs, ensure high availability, and scale compute nodes on demand, replacing the Slurm controller API to handle hundreds of thousands of jobs. It also details how SUNK's features like priorities, preemption, quotas, gang scheduling, and topology-aware scheduling address the challenges of running large-scale RL jobs, leading to faster job startup, more consistent cluster utilization, and higher end-to-end throughput. The post also touches upon the researcher-centric experience provided by SUNK, including secure isolated environments and IdP-federated cluster access.

Read the original post ↗