
6/25/2020 · Olena Horal-Koretska
What this post added
This post details learnings from shadowing an SRE, emphasizing the interconnectedness of systems, the importance of resource management (e.g., CPU allocation for PostgreSQL vs. Redis), and how SREs use tools like Prometheus and Kibana for monitoring and incident detection. It highlights that incidents are essentially bugs in site infrastructure and that SREs proactively scale services based on resource usage trends. The post also touches on the role of SREs in supporting and developing GitLab infrastructure beyond just incident response.