Data Warehousing and Analytics Platform
How Meta understands data at scale

How Meta understands data at scale

4/28/2025 · Vasileios Lakafosis, Hannes Roth, Benjamin Renard, Wenlong Dong, Zhonghu Gao, Dave Kurtzberg

What this post added

This post details Meta's advancements in data understanding as part of its Privacy Aware Infrastructure (PAI). Key contributions include the development of a "shift-left" approach integrating data schematization and annotations early in product development, the creation of a universal privacy taxonomy for standardized data privacy management, and the establishment of OneCatalog for discovering, registering, and enumerating data assets. The post also outlines solutions to challenges such as understanding data at scale through a shared asset schema format, ensuring consistent definitions with a unified taxonomy of semantic types, improving annotation quality by combining schematization with code annotations and multiple classification signals, and overcoming organizational barriers through collaboration and intuitive tooling. A walkthrough of understanding user data for the "Beliefs" feature in Facebook Dating illustrates the five-step approach: schematizing data using DataSchema (based on Thrift IDL), predicting metadata at scale through a universal privacy taxonomy and data classification, and applying these principles across distributed systems and data warehouses.

Read the original post ↗