Site Reliability and Outage Management
SLICK: Adopting SLOs for improved reliability

SLICK: Adopting SLOs for improved reliability

12/13/2021 · A Posten, Dávid Bartók, Filip Klepo, Vatika Harlalka

What this post added

This post introduces SLICK, a dedicated SLO store built to standardize and centralize Service Level Indicator (SLI) and Service Level Objective (SLO) definitions across Meta. SLICK provides tooling for defining SLOs, offers high-retention, full-granularity data (up to two years) for key service metrics, and integrates SLOs into daily workflows and incident management. The architecture includes a DSL for configuration, a syncer, a UI for dashboards and index, a service for query abstraction, and data pipelines for ingestion into a sharded MySQL database. SLICK has seen widespread adoption, with over 1,000 services onboarded, and has demonstrably helped teams like LogDevice and backend ML services identify and fix reliability regressions.

Read the original post ↗