Instant GPU Clusters
New in Together GPU Clusters: Autoscaling, observability, and self-healing

New in Together GPU Clusters: Autoscaling, observability, and self-healing

3/10/2026

What this post added

This post introduces major enterprise enhancements to Together GPU Clusters, including the integration of autoscaling (powered by Kubernetes Cluster Autoscaler), Role-Based Access Control (RBAC) with 'Admin' and 'Member' roles scoped to 'Projects', full-stack observability via dedicated Grafana instances with pre-built dashboards for GPU, networking, storage, and orchestration metrics, and self-serve node repair with active health checks (DCGM Diag, NCCL, InfiniBand tests) and automated acceptance tests during provisioning. These features aim to improve elasticity, governance, debugging, and reliability for GPU cluster management.

Read the original post ↗