Data Warehousing and Analytics Platform
Data @Scale – Boston Recap

Data @Scale – Boston Recap

11/19/2014 · Ryan Mack

What this post added

This post details presentations from the Data @Scale – Boston conference, highlighting advancements in large-scale data processing and systems engineering. Key technical contributions include: Twitter's use of the Lambda architecture and probabilistic algorithms for real-time high-volume analytics processing; Facebook's technical challenges and infrastructure enablers for the 'Lookback videos' project, focusing on compute, network, storage, distribution, projections, and modeling; Vertica's Database Designer (DBD) for automated projection design in columnar databases and Live Aggregate Projections (LAPs) for faster querying through pre-aggregation; Wayfair's scaling of Redis and Memcached using composable tools like Ketama and Twemproxy for a resilient distributed caching system; TripAdvisor's lessons learned in handling operations at scale with Hadoop, log shipping, and anomaly detection; Facebook's implementation of spatial indexing on RocksDB for efficient geo-spatial data storage and optimization for various workloads; Facebook's use of graph partitioning for optimizing distributed systems serving graph-based datasets; Constant Contact's scaling strategies using Cassandra for key-value data and sharded MySQL for relational data, and their analytics platform leveraging Hadoop; and VoltDB's architecture for high-volume transactional and stream processing.

Read the original post ↗