
12/16/2019 · Tristan Read
What this post added
This post details a frontend engineer's experience shadowing an on-call SRE, observing incident response activities. Key takeaways include the utility of change-based alerting, the distinction between alerts and incidents, the diverse toolset used by SREs (PagerDuty, Slack, Grafana, Kibana, Zoom, internal GitLab projects), the potential single point of failure in monitoring GitLab.com with GitLab itself, and feature proposals for multi-user issue editing and more webhooks for issues. The post highlights the complexity of SRE work and the ongoing development of features to support these workflows.