
Scaling Grab's Data Lake: Our journey to Apache Iceberg adoption
7/24/2026
This post details Grab's migration from Hive Parquet to Apache Iceberg for its data lake. It outlines the challenges faced with the previous directory-based Hive Metastore approach, including catalog latency, the small file problem, and operational overhead. The post explains the strategic decision to adopt Iceberg, highlighting its benefits for consistency and performance. It showcases the substantial efficiency gains achieved through Iceberg adoption, such as 10x query performance improvement with Z-ordering and significant S3 API cost reductions. A major contribution is the introduction and open-sourcing of the UnifiedSparkCatalog, a Spark catalog that abstracts away table format differences, enabling seamless coexistence and migration of tables across Iceberg, Delta, Hudi, and Hive. The post also discusses lessons learned, including handling Hive lock contention and timestamp compatibility issues, and outlines future directions like Storage Partitioned Joins and Apache XTable.
