Cassandra Data Movement
The Evolution of Cassandra Data Movement at Netflix

The Evolution of Cassandra Data Movement at Netflix

6/20/2026 · Netflix Technology Blog

What this post added

This post details the evolution of Netflix's Cassandra data movement, moving from the monolithic Casspactor engine to a new layered architecture. Key technical contributions include: 1. Replacing fragile metadata dependencies with direct S3 reads as the single source of truth. 2. Introducing a 'Connector Factory' model built on Spark DataFrames, allowing data abstractions to create model-aware connectors. 3. Shifting mutation compaction and processing to Spark Executors to handle skewed partitions and avoid out-of-memory errors. 4. Eliminating intermediate Iceberg tables to reduce storage bloat. 5. Implementing robust time travel by processing schema, topology, and data as a cohesive unit. 6. Introducing auto-sizing capabilities for jobs. 7. Achieving significant performance and cost improvements.

Read the original post ↗