
4/22/2025 · Mark Sakurada
What this post added
This post introduces strategies for achieving data efficiency by "shifting left" data governance and processing. It highlights the importance of schemas for data consistency and quality, and the potential of stream processing (using Apache Flink) for efficient data transformation. It also emphasizes the role of open table formats like Apache Iceberg for data reusability and consistency. Three key strategies are presented: implementing no-copy and zero ETL solutions, shifting data governance left, and reducing data waste by processing data at the source.