BlogsModalGPU Health and Reliability System

GPU Health and Reliability System

GPU Health and Reliability System

1
posts
2025

Modal has developed a comprehensive system for maintaining the health and reliability of its globally distributed GPU worker pool, encompassing over 20,000 GPUs. This system includes rigorous instance type testing and selection, robust machine image preparation and automated testing, lightweight instance boot checks, and continuous lifetime management through passive and active healthchecking. Observability is provided via detailed GPU metrics for customers and internal dashboards. The system aims to proactively identify and mitigate GPU hardware issues across multiple cloud providers.

2025

Keeping 20,000 GPUs healthy

12/28/2025

This post details Modal's GPU reliability system, covering instance type testing and selection based on performance and reliability differences across cloud providers, machine image preparation with automated testing, instance boot checks, and lifetime management via passive (dmesg, dcgmi health) and active (dcgmi diag, GPUBurn, NCCL tests) healthchecking. It also discusses observability through GPU metrics like memory usage, utilization, temperature, and power. The system automatically marks unhealthy hosts and initiates disposal or reinstallation.