
What does 99.9% uptime mean for inference?
7/16/2026
This post defines the engineering requirements and architectural considerations for achieving different levels of inference uptime (99%, 99.9%, 99.99%). It details failure domains at node, data center, and regional levels, and explains the corresponding engineering solutions such as automated health checking, multi-facility deployment with live traffic routing, and multi-region deployment with reserved capacity. The post also highlights the importance of infrastructure ownership and full-stack visibility for reliable inference, and outlines key questions customers should ask providers about their SLA definitions and architecture.