
6/5/2008
What this post added
This post details Facebook's early adoption and scaling of Hadoop for large-scale data storage and processing. It highlights the use of Hadoop's distributed file system and map-reduce paradigm to enable previously impossible projects, leading to features like the Facebook Lexicon and improved search relevance. The post describes the deployment of multiple Hadoop clusters, daily data loading volumes, and the diverse range of projects utilizing this infrastructure. Key decisions aiding adoption include language flexibility for map-reduce programs and the embrace of SQL via the in-house data warehousing layer called Hive, which offers classic data warehouse features and is planned for open-source release.