
6/18/2020
What this post added
This post introduces foundational approaches to data warehousing and analysis at Shopify. It details the adoption of dimensional modeling (Kimball methodology) for data schemas, a unified data modeling platform built on Spark within a single GitHub repo, and open access to modelled data via a Presto cluster. The post also highlights rigorous ETL processes with unit testing, centralized vetted dashboards, reproducible vetted data points, a culture of peer review for all data work, deep product understanding within specialized data teams, effective communication of insights with recommendations, collaboration across data teams, and a positive philosophy about data's impact. The core technical contributions lie in establishing standardized data modeling practices, ensuring data consistency and accessibility, and implementing robust data pipeline testing and validation.