
9/26/2025
What this post added
Reverse-engineered and explained the internal architecture of Flash Attention 4, focusing on its asynchronous pipeline, warp specialization, and the 'life of a tile' through the GPU memory hierarchy. Detailed the implementation of faster approximate exponentials and an efficient online softmax using CUDA Cores and a cubic polynomial approximation, contrasting it with traditional SFU usage. Explained the producer/consumer model with barriers and the mapping of pipeline steps onto warps.