Instant GPU Clusters
New in Together GPU Clusters: Reliability and control for production GPU clusters

New in Together GPU Clusters: Reliability and control for production GPU clusters

7/15/2026

What this post added

Introduced passive health checks and auto node repair for detecting and recovering from hardware failures. Upgraded the Slurm-on-Kubernetes stack (Slinky 1.0) for improved reliability, including self-healing worker daemons, no zombie processes, durable job accounting, reliable process cleanup, and accurate GPU state after reschedules. Added a new cluster details view with node health, usage metrics, and an event timeline. Implemented external OIDC for Kubernetes RBAC to enable per-user authentication and authorization. Introduced startup scripts for cluster customization at various lifecycle events.

Read the original post ↗