
7/24/2026
What this post added
This post details the evolution of Grab's Hugo data ingestion platform, migrating from a Spark-based batch processing system to a unified, self-service platform powered by Apache Flink. It describes the architectural shift from siloed workflows using Kafka Connect and Sprinkler to a streamlined Flink-based approach for both MySQL CDC and Kafka ingestion. Key technical contributions include the elimination of intermediary Kafka hops for CDC, automated schema detection and dynamic validation for Kafka streams, and a reduction in pipeline components from four to two. The post highlights significant improvements in onboarding time (from days to minutes) and increased adoption rates.