
11/13/2025
What this post added
This post details Decagon's collaboration with Modal to improve speculative decoding for LLM inference in real-time voice AI. Key contributions include: 1. Forking and building a custom training pipeline for EAGLE3 draft models using SpecForge, pre-training on broad conversational data and fine-tuning on Decagon's domain dataset, resulting in 38% higher accept lengths. 2. Re-engineering the SGLang inference engine by building a new asynchronous scheduler for speculative decoding, eliminating CPU stalls and enabling parallel GPU kernel execution, which improved throughput by up to 12% and achieved p90 latency of 342ms.