FlashAttention-4 Algorithm and Kernel Co-Design
Inside the Together AI kernels team

Inside the Together AI kernels team

4/1/2026

What this post added

This post details the development and impact of the Kernels Lab, highlighting their work on FlashAttention and the ThunderKittens library for optimizing NVIDIA Blackwell GPUs. It also showcases the Together Megakernel implementation for real-time voice agent workloads, achieving significant latency reductions. The team's approach emphasizes academic-industry symbiosis and customer-facing collaboration for custom kernel optimization.

Read the original post ↗