
7/27/2026
What this post added
This post details the integration of Moonshot's Kimi K3 model with Modal's platform, specifically highlighting the application of a custom-trained DFlash speculator tuned to K3's architecture. This integration resulted in a 360% faster interactivity (from 100 to 460 tokens per second) and 88% higher throughput (from 800k to 1.5 million TPM per GPU) on agentic workloads. The post also touches upon the engineering challenges and solutions involved in making Kimi K3, a 2.8 trillion parameter multimodal model with a 1M token context window, servable at scale, including its mixture-of-experts architecture, quantization-aware training, and contributions to vLLM for prefix caching.