
7/27/2026
What this post added
This post details the technical implementation of Kimi Delta Attention, FP4 MoE kernels, and adapted decode kernels to optimize Kimi K3 performance. It highlights the model architecture's new Kimi Delta Attention and refined attention residuals, and the scaling of Mixture of Experts (MoE) to 16 out of 896 experts with Stable Latent MoE, resulting in 2.5 times the scaling efficiency compared to Kimi K2. The post also emphasizes the cost-effectiveness for high-volume production workloads and the benefits for large, long-horizon tasks.