
3/3/2023 · Lukasz Mierzwa
What this post added
This post details how Cloudflare operates a large-scale Prometheus deployment (916 instances, 4.9 billion time series) for network monitoring. It explains the concepts of metrics, labels, cardinality, samples, and time series, and discusses the challenges of high cardinality and memory consumption in Prometheus. It also outlines the initial steps in the Prometheus lifecycle: HTTP scrape and TSDB storage.