Speculative Decoding Acceleration
Together AI delivers fastest inference for the top open-source models

Together AI delivers fastest inference for the top open-source models

12/1/2025

What this post added

This post details significant advancements in Together AI's inference platform, focusing on achieving up to 2x faster serverless inference for open-source LLMs. Key technical contributions include: 1. Next-gen GPU hardware optimization for NVIDIA Blackwell architecture (GB200 NVL72), including FP8/FP4 compute and low-overhead scheduling. 2. Development of 'Together Kernels' for Blackwell, including FlashAttention-4 and fused MoE kernels. 3. Turbo optimization suite featuring near-lossless quantization to FP8/FP4 with architecture-aware calibration and block-wise scaling. 4. Production-grade speculative decoding with training-efficient algorithms, high-accuracy custom draft models, adaptive acceptance strategies, and fail-safe fallback mechanisms. 5. A scalable draft-model training pipeline supporting models up to 1T parameters, utilizing curriculum-based training, post-training recipes, data-mixing, and an alignment evaluation framework.

Read the original post ↗