FlashAttention-4 Algorithm and Kernel Co-Design
FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling

FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling

3/5/2026

What this post added

This post details FlashAttention-4, an optimization for attention kernels on asymmetric hardware like Blackwell GPUs. Key contributions include: 1. New forward and backward software pipelines for maximum overlap of tensor cores, softmax exponential, and memory operations. 2. Forward pass: software emulation of exponential using FMA units and conditional online softmax rescaling. 3. Backward pass: storing intermediates in tensor memory, using 2-CTA MMA mode to reduce shared memory traffic and atomic operations, and supporting deterministic execution. 4. New tile scheduler for causal masks and variable sequence lengths.

Read the original post ↗