
GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model
8/3/2026
This post details Meta's advancements in optimizing the training efficiency of its LLM-scale ads foundation model, GEM. Key contributions include: 1. Co-designing compute and scaling efficiency innovations to double end-to-end training efficiency to 20-25% MFU and scale training FLOPs 4x in 12 months. 2. Developing a customized recommendation kernel library, including Jagged Flash Attention (JFA) to eliminate padding waste on jagged inputs and improve backward pass efficiency, and Generalized Dot-Product Attention (GDPA) to unify and accelerate diverse, asymmetric attention modules. 3. Implementing mixed ultra-low precision training (MXFP8 attention and MLP) optimized for recommendation workloads. 4. Employing topology-aware 5D parallelism with SM-free collectives (2D FSDP + Expert Parallelism for dense parameters, Fully Sharded 2D Model Parallelism for sparse parameters) co-designed with Meta's multi-tiered network hierarchy to reduce communication overhead.

































































































































