
Keeping 20,000 GPUs healthy
12/28/2025
This post details Modal's GPU reliability system, covering instance type testing and selection based on performance and reliability differences across cloud providers, machine image preparation with automated testing, instance boot checks, and lifetime management via passive (dmesg, dcgmi health) and active (dcgmi diag, GPUBurn, NCCL tests) healthchecking. It also discusses observability through GPU metrics like memory usage, utilization, temperature, and power. The system automatically marks unhealthy hosts and initiates disposal or reinstallation.