Generative Recommender Training Efficiency
Faster than Light: Optimizing Generative Recommender Training Efficiency at LinkedIn

Faster than Light: Optimizing Generative Recommender Training Efficiency at LinkedIn

5/28/2026

What this post added

Introduced a C++-based fused data loader to reduce I/O bottlenecks and average training step time by 50%. Replaced standard attention kernels with FlashAttention-3 and FlexAttention, achieving up to 25% faster training for Ads GR and 2x faster for 3D mask training. Developed a custom CUDA kernel for fused metrics calculation, reducing update time from ~40ms to ~0.5ms and contributing to 22% GPU hour savings. Enabled fused optimizer operations by turning on the fused flag in Adam, reducing optimizer time by 50% and contributing to 15% GPU hour savings. Implemented fused embedding table lookups to reduce kernel launches and memory traffic.

Read the original post ↗