Data Infrastructure & Analytics
How Cloudflare runs Prometheus at scale

How Cloudflare runs Prometheus at scale

3/3/2023 · Lukasz Mierzwa

What this post added

This post details how Cloudflare operates a large-scale Prometheus deployment (916 instances, 4.9 billion time series) for network monitoring. It explains the concepts of metrics, labels, cardinality, samples, and time series, and discusses the challenges of high cardinality and memory consumption in Prometheus. It also outlines the initial steps in the Prometheus lifecycle: HTTP scrape and TSDB storage.

Read the original post ↗