Log Management and Observability
Shadowing a Site Reliability Engineer

Shadowing a Site Reliability Engineer

4/13/2020 · Laura Montemayor

What this post added

This post provides a first-hand account of shadowing Site Reliability Engineers (SREs) at GitLab, detailing their daily workflow, tooling, and challenges. It highlights the SRE's role in managing alerts, distinguishing between noise and actionable incidents, and responding to actual incidents. Key takeaways include the importance of streamlined tooling (GitLab issues, Slack, Zoom), the challenge of managing alert noise, the relative infrequency of major incidents, the difficulty of effective monitoring, the critical role of communication (especially in an all-remote, async environment), the love for documenting everything in issues for handover and root cause analysis, the collaborative nature of incident resolution, and the growing field of monitoring with the creation of a Scalability team to curate alerting criteria.

Read the original post ↗