
8/3/2026 · Darren Liu, Huayu Li, Raghav Boinepalli, Yuzhen Huang, Jackie (Jiaqi) Xu, Richard Qiu, Chunzhi Yang, Rich Zhu, Dev (Devashish) Shankar, Huaqing Xiong, Lei Tian, Ruilin Chen, Xiaoyi (Leo) Liu, Yasmine Badr, Wentao Duan
What this post added
This post details Meta's advancements in optimizing the training efficiency of its LLM-scale ads foundation model, GEM. Key contributions include: 1. Co-designing compute and scaling efficiency innovations to double end-to-end training efficiency to 20-25% MFU and scale training FLOPs 4x in 12 months. 2. Developing a customized recommendation kernel library, including Jagged Flash Attention (JFA) to eliminate padding waste on jagged inputs and improve backward pass efficiency, and Generalized Dot-Product Attention (GDPA) to unify and accelerate diverse, asymmetric attention modules. 3. Implementing mixed ultra-low precision training (MXFP8 attention and MLP) optimized for recommendation workloads. 4. Employing topology-aware 5D parallelism with SM-free collectives (2D FSDP + Expert Parallelism for dense parameters, Fully Sharded 2D Model Parallelism for sparse parameters) co-designed with Meta's multi-tiered network hierarchy to reduce communication overhead.