
3/5/2026
What this post added
This post details FlashAttention-4, an optimization for attention kernels on asymmetric hardware like Blackwell GPUs. Key contributions include: 1. New forward and backward software pipelines for maximum overlap of tensor cores, softmax exponential, and memory operations. 2. Forward pass: software emulation of exponential using FMA units and conditional online softmax rescaling. 3. Backward pass: storing intermediates in tensor memory, using 2-CTA MMA mode to reduce shared memory traffic and atomic operations, and supporting deterministic execution. 4. New tile scheduler for causal masks and variable sequence lengths.