FlashAttention-4 Algorithm and Kernel Co-Design
Together AI at ICML 2026: frontier research across the full stack

Together AI at ICML 2026: frontier research across the full stack

6/30/2026

What this post added

This post details the Aurora paper on adaptive speculative decoding, which is shipped as the ATLAS speculator in production. It highlights how research across the full stack, from frontier agents to GPU kernels, contributes to the Together platform and production workloads.

Read the original post ↗