BlogsBackblaze Feature Trails

Backblaze logo

Backblaze Feature Trails

See how major capabilities shipped, upgraded, and evolved across Backblaze's engineering blog.

Feature trails

2

AI Data Pipeline Ingest and Archive

Active

This post introduces the foundational stages of an AI data pipeline, focusing on data ingest and archiving. It emphasizes the critical role of storage infrastructure in handling streaming ingestion, the importance of robust taxonomy for data organization and searchability, and the necessity of retaining comprehensive metadata (both as object tags and sidecar files) for reproducibility and downstream processing. The post also discusses how to expose metadata to training pipelines via manifests and highlights the specific storage requirements for GenAI workloads, including sustained high throughput, an always-hot architecture with no tiering, and free data movement, as exemplified by Backblaze B2 Overdrive.

2 posts

Timeline

Network Traffic Analysis and Capacity Planning

Active

This post continues the analysis of network traffic variance, focusing on the impact of AI workloads. It details the creation of a new time-series dataset, the use of statistical baselines, and the application of Python's SciPy library for generating variance signals. The analysis covers total traffic volume, magnitude (bits per IP), and communication uniqueness across different network types (CDN, Hosting, Hyperscaler, ISP-regional, Neocloud) and regions (US-West, U. The post also delves into t. This post discusses the challenges of data silos in AI projects, emphasizing the need for centralized data management and infrastructure planning. It highlights the importance of inventorying data assets, establishing governance, and planning storage infrastructure for future needs, particularly concerning multimodal data and hyperscaler egress fees. The competitive advantage in AI is identified as proprietary data, and aligning AI strategy with data strategy is crucial for scalable AI programs.

4 posts

Timeline