
11/10/2016 · Pablo Carranza
What this post added
This post details the decision to move GitLab's infrastructure from the cloud to bare metal due to performance limitations and cost inefficiencies encountered with CephFS on shared cloud resources. It highlights the challenges of achieving consistent high IOPS performance in the cloud, the concept of 'noisy neighbors,' and the trade-offs between cloud scalability and actual performance at scale. The post also emphasizes the importance of building observable systems with detailed monitoring (using Prometheus) to proactively identify and address infrastructure issues, using OSD Journal Latency as a key metric that indicated downtime.