BlogsBackblazeNetwork Traffic Analysis and Capacity Planning

Network Traffic Analysis and Capacity Planning

Network Traffic Analysis and Capacity Planning

4
posts
2026

This post continues the analysis of network traffic variance, focusing on the impact of AI workloads. It details the creation of a new time-series dataset, the use of statistical baselines, and the application of Python's SciPy library for generating variance signals. The analysis covers total traffic volume, magnitude (bits per IP), and communication uniqueness across different network types (CDN, Hosting, Hyperscaler, ISP-regional, Neocloud) and regions (US-West, U. The post also delves into t. This post discusses the challenges of data silos in AI projects, emphasizing the need for centralized data management and infrastructure planning. It highlights the importance of inventorying data assets, establishing governance, and planning storage infrastructure for future needs, particularly concerning multimodal data and hyperscaler egress fees. The competitive advantage in AI is identified as proprietary data, and aligning AI strategy with data strategy is crucial for scalable AI programs.

2026

Network Stats for Q2 2026: All Eyes on Neocloud Traffic Variance

7/23/2026

Introduced a new methodology for analyzing network traffic variance by building a time-series dataset and modeling traffic behavior against statistical baselines. Utilized SciPy for statistical analysis to generate variance signals. Presented new analysis on traffic variance, focusing on the impact of neocloud and hyperscaler traffic, and its implications for network architecture and capacity planning. Updated existing charts with Q2 2026 data and introduced new heatmaps for magnitude and uniqueness metrics.

Backblaze Drive Stats for Q1 2026

7/9/2026

This post provides Q1 2026 drive statistics, including quarterly and lifetime Annualized Failure Rates (AFR) for various drive models. It details the methodology for drive selection and exclusion criteria. A significant portion of the post discusses an edge case involving two different mechanical issues affecting specific drive models, impacting writes and power cycling. The post highlights mitigation strategies, such as reducing power cycling frequency and using a 'no-upload' mode for affected Vaults, and emphasizes the role of active management in mitigating drive failure risks. It also notes the continued investment in higher capacity drives and their impressive AFR.

Why AI Projects Stall: Data Silos

6/15/2026

This post identifies data silos as a major impediment to AI project success. It advocates for centralizing multimodal data, establishing data governance early, and planning storage infrastructure with long-term costs and retrieval latencies in mind. The post emphasizes that proprietary data is the key differentiator for AI initiatives and that aligning AI strategy with data strategy leads to better infrastructure decisions.

What Network Data Can and Can’t Tell Us About AI Infrastructure

6/10/2026

This post analyzes network telemetry data to understand the operational demands of AI workloads. It discusses how traffic volume, connection persistence, endpoint concentration, and geographic clustering provide insights into data movement patterns. The post also emphasizes the limitations of interpreting a single quarter of data, highlighting the need for longer-term trend analysis and flexible infrastructure design in the rapidly evolving AI landscape. It explains how network telemetry can inform infrastructure planning by revealing practical questions about sustained high-throughput networking, storage and compute coupling, and regional infrastructure concentration.