Data Warehousing and Analytics Platform
Data @Scale – Boston recap

Data @Scale – Boston recap

11/13/2018

What this post added

This post summarizes presentations from the Data @Scale conference in Boston. Key technical contributions include: discussions on protecting patient privacy while using large-scale health care data sets; balancing flexibility and control in database deployments; lessons learned from scaling a timeseries database (InfluxDB), including failure conditions and trade-offs between monolithic and service-oriented implementations; leveraging sampling to reduce data warehouse resource consumption and manage uncertainty propagation in aggregated metrics; exploring Transient Replication and Cheap Quorums in Apache Cassandra for disk space and compute savings; detailing Facebook's Deletion Framework for managing data deletion at scale across distributed systems; describing Wayfair's transformation of data plumbing components into a scalable infrastructure, including the development of Tremor as a replacement for logstash; presenting Kubeflow for portable machine learning on Kubernetes, addressing scalability, portability, and composability; outlining DataXu's journey to a cloud-native warehouse using AWS services and spot instances; detailing HubSpot's process for migrating Elasticsearch instances at scale to enhance security and reduce migration time; presenting Presto's Cost-Based Optimizer (CBO) for improved join efficiency and a new mechanism for seamless statistics collection; introducing Palisade and Akamill for overload protection and stream processing in data analytics and storage systems; and discussing best practices for building highly reliable data pipelines at Datadog, including ephemeral clusters, job isolation, and rapid recovery mechanisms.

Read the original post ↗