
6/24/2026
What this post added
This post details the application of speculative decoding for achieving state-of-the-art inference latencies, specifically highlighting its use with Modal Auto Endpoints and Modal Servers. It elaborates on the low-latency playbook, emphasizing optimizations in client-server communication, host overhead, prefill latency, and decode latency. A significant contribution is the detailed explanation of applying speculative decoding with custom speculator models, including 'mid-training' on task-specific synthetic data using techniques like DFlash. The post also discusses the integration with open-source engines (SGLang) and kernels (FlashAttention-4) and provides a case study with Decagon Voice, demonstrating a 100ms reduction in p50 latency.