5/28/2026
What this post added
Introduced a C++-based fused data loader to reduce I/O bottlenecks and average training step time by 50%. Replaced standard attention kernels with FlashAttention-3 and FlexAttention, achieving up to 25% faster training for Ads GR and 2x faster for 3D mask training. Developed a custom CUDA kernel for fused metrics calculation, reducing update time from ~40ms to ~0.5ms and contributing to 22% GPU hour savings. Enabled fused optimizer operations by turning on the fused flag in Adam, reducing optimizer time by 50% and contributing to 15% GPU hour savings. Implemented fused embedding table lookups to reduce kernel launches and memory traffic.