
5/15/2025 · Exa Labs
What this post added
This post details the infrastructure powering Exa's neural network search engine, including a large GPU cluster (Exacluster) built with 144 H200 GPUs and 3,456 CPUs, managed by Pulumi for IaC, Ansible/Kubespray for Kubernetes deployment, NVIDIA operators for hardware management, Alluxio for distributed caching of local NVMe storage, and Flyte for workflow orchestration. This setup enables rapid training and deployment of large-scale ML models.