Observability Platform
Minimizing on-call burnout through alerts observability

Minimizing on-call burnout through alerts observability

3/29/2024 · Monika Singh

What this post added

This post details the evolution of Cloudflare's alert observability by implementing a new data pipeline using Vector.dev to aggregate all alert states (firing, silenced, inhibited, resolved) from Prometheus Alertmanager into ClickHouse. This addresses limitations of previous tools by providing complete state information, enabling better troubleshooting, reporting, and analysis of alert noise to mitigate on-call burnout. New dashboards for alert overview, alertname breakdown, receiver-specific insights, state timelines, and silence analysis have been developed.

Read the original post ↗