.png)
3/10/2026
What this post added
This post introduces major enterprise enhancements to Together GPU Clusters, including the integration of autoscaling (powered by Kubernetes Cluster Autoscaler), Role-Based Access Control (RBAC) with 'Admin' and 'Member' roles scoped to 'Projects', full-stack observability via dedicated Grafana instances with pre-built dashboards for GPU, networking, storage, and orchestration metrics, and self-serve node repair with active health checks (DCGM Diag, NCCL, InfiniBand tests) and automated acceptance tests during provisioning. These features aim to improve elasticity, governance, debugging, and reliability for GPU cluster management.