GPU Health and Reliability System
Keeping 20,000 GPUs healthy

Keeping 20,000 GPUs healthy

12/28/2025

What this post added

This post details Modal's GPU reliability system, covering instance type testing and selection based on performance and reliability differences across cloud providers, machine image preparation with automated testing, instance boot checks, and lifetime management via passive (dmesg, dcgmi health) and active (dcgmi diag, GPUBurn, NCCL tests) healthchecking. It also discusses observability through GPU metrics like memory usage, utilization, temperature, and power. The system automatically marks unhealthy hosts and initiates disposal or reinstallation.

Read the original post ↗