
3/8/2017 · Kevin Lee
What this post added
Introduced Big Basin, a next-generation AI GPU server, as the successor to Big Sur. Big Basin enables training of 30% larger machine learning models due to increased arithmetic throughput and memory (16 GB vs 12 GB). It features a modular, scalable design with disaggregated CPU and GPU compute, allowing independent scaling of components and integration with existing OCP infrastructure like the Tioga Pass server platform. The GPU tray is swappable for future upgrades. The system uses external PCIe cables for head node to GPU connection, functioning as a 'just a bunch of GPUs' (JBOG). It offers improved serviceability and thermal efficiency by positioning GPUs in front of cool air intake. Big Basin is equipped with eight NVIDIA Tesla P100 GPUs connected via NVLink in a hybrid cube mesh, offering improved performance per watt with single-precision floating-point arithmetic increasing from 7 to 10.6 teraflops and introducing half-precision for higher throughput. Tested with ResNet-50, it achieved nearly 100% throughput improvement over Big Sur. The design specifications are open-sourced via the Open Compute Project.