BlogsTogether AIInstant GPU Clusters

Instant GPU Clusters

Instant GPU Clusters

9
posts
2025–2026

Together GPU Clusters (formerly Instant Clusters) now integrate autoscaling powered by Kubernetes Cluster Autoscaler, Role-Based Access Control (RBAC) for structured multi-team governance, full-stack observability with dedicated Grafana instances, and self-serve node repair with active health checks and acceptance testing to reduce Mean Time To Repair (MTTR) for hardware failures. These enhancements aim to provide the elasticity of a virtualized stack with the performance of bare metal, enabling production-ready managed infrastructure for large-scale distributed training and variable inference workloads.

2026

Together AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community

7/20/2026

This post announces a partnership with Y Combinator to provide a dedicated GPU cluster for YC's AI-native startups. It highlights the challenge of compute access for startups and positions the dedicated cluster as a solution offering flexible, cost-effective access to GPUs for both inference and training. The cluster is managed via Together's self-service portal, allowing startups to provision and manage their own resources. The partnership aims to support founders by providing resources comparable to larger companies.

New in Together GPU Clusters: Reliability and control for production GPU clusters

7/15/2026

Introduced passive health checks and auto node repair for detecting and recovering from hardware failures. Upgraded the Slurm-on-Kubernetes stack (Slinky 1.0) for improved reliability, including self-healing worker daemons, no zombie processes, durable job accounting, reliable process cleanup, and accurate GPU state after reschedules. Added a new cluster details view with node health, usage metrics, and an event timeline. Implemented external OIDC for Kubernetes RBAC to enable per-user authentication and authorization. Introduced startup scripts for cluster customization at various lifecycle events.

What is an AI Native Cloud?

4/7/2026

This post defines the concept of an 'AI Native Cloud' and outlines its key characteristics: a full AI stack from hardware to software, a fast path from research to production, reliability at massive scale, a builder-centric approach, and a partnership model that moves at AI-native pace. It positions Together AI as building this AI Native Cloud, emphasizing its vertical integration and continuous evolution with research innovations. While it references existing infrastructure like GPU clusters, its primary contribution is conceptualizing and defining the overarching cloud strategy for AI-native companies.

New in Together GPU Clusters: Autoscaling, observability, and self-healing

3/10/2026

This post introduces major enterprise enhancements to Together GPU Clusters, including the integration of autoscaling (powered by Kubernetes Cluster Autoscaler), Role-Based Access Control (RBAC) with 'Admin' and 'Member' roles scoped to 'Projects', full-stack observability via dedicated Grafana instances with pre-built dashboards for GPU, networking, storage, and orchestration metrics, and self-serve node repair with active health checks (DCGM Diag, NCCL, InfiniBand tests) and automated acceptance tests during provisioning. These features aim to improve elasticity, governance, debugging, and reliability for GPU cluster management.

Key research and product announcements at the AI Native Conf

3/5/2026

This post introduces several significant technical advancements that enhance AI infrastructure and performance. Key contributions include: FlashAttention-4, a new attention algorithm optimized for NVIDIA Blackwell GPUs, offering substantial speedups over existing solutions. Together Megakernel, a single-kernel implementation for running entire models, achieving significant latency reduction for real-time voice agents. together.compile, an automation tool for generating optimized GPU kernels, improving generation speed for video and image models. The Reinforcement Learning API is enhanced with infrastructure for RL training, focusing on efficient weight distribution and rollout optimization. ThunderAgent is introduced as a program-aware abstraction for agentic workloads, addressing KV cache thrashing, cross-node memory imbalance, and tool lifecycle issues, leading to improved throughput and memory savings. ATLAS-2 provides an online training flywheel for speculative decoding, continuously updating speculators from live traffic for sustained performance gains. Cache-aware prefill–decode disaggregation (CPD) introduces a three-tier serving stack to optimize long-context inference by intelligently routing requests based on cache hit rates, significantly increasing sustainable throughput.

Inside multi-node training: How to scale model training across GPU clusters

1/12/2026

This post details the infrastructure, techniques, and practical steps for scaling model training across GPU clusters, focusing on multi-node distributed training. It explains parallelism strategies (data, tensor, pipeline), network interconnects (NVLink, InfiniBand), and checkpointing/fault tolerance. A production example of training Qwen2.5-72B is provided, along with guidance on getting started, infrastructure verification, framework configuration, checkpointing implementation, scaling tests, and monitoring. It also covers common failure modes and troubleshooting for multi-node training.

2025

How to run TorchForge reinforcement learning pipelines in the Together AI Native Cloud

12/3/2025

This post details the integration of TorchForge reinforcement learning (RL) pipelines with Together AI's Instant Clusters. It highlights the use of Instant Clusters' low-latency GPU communication, preconfigured drivers, and heterogeneous scheduling for RL workloads. The post also describes the integration of Together CodeSandbox and Together Code Interpreter as environments for RL agents, enabling tool-use and code execution. A demo is provided for training a BlackJack agent using GRPO, vLLM, and TorchStore, and a standalone integration for using Code Interpreter as an OpenEnv environment is also released.

Announcing General Availability of Together Instant Clusters, offering ready to use, self-service NVIDIA GPUs

9/9/2025

This post announces the General Availability of Together Instant Clusters, a new self-service, API-first offering for provisioning multi-node NVIDIA GPU clusters. It details the automated setup, pre-configured components (GPU Operator, NVIDIA Network Operator, Cert Manager), and optimizations for distributed training (InfiniBand, NVLink) and inference. The post also highlights reliability measures, including burn-in tests and continuous monitoring, and introduces flexible pricing models.

Bringing 100,000 GPUs to Europe

6/12/2025

This post announces the expansion of Together AI's GPU cluster infrastructure to Europe, detailing a partnership with Hypertec and 5C Group to deploy up to 2 gigawatts of AI-dedicated data center capacity and nearly 100,000 NVIDIA Blackwell and future-generation GPUs. The deployment will prioritize France, the UK, Italy, and Portugal, with initial rollouts beginning in late 2025 and large-scale buildouts planned through 2028. The focus is on providing sovereign, regulation-ready AI infrastructure to meet European demand for frontier model training and inference workloads, while also emphasizing sustainability and compliance with EU regulations.