
5/20/2026
What this post added
This post details CoreWeave's achievement of leading inference performance for the Kimi K2.6 model, as benchmarked by Artificial Analysis. It highlights the technical optimizations employed, including training a custom NVFP4 quantized model and implementing EAGLE3 speculative decoding on NVIDIA GB300 and GB200 NVL72 clusters. The post also mentions the validation of these optimizations across various benchmarks and the integration of these performance enhancements throughout the CoreWeave Inference stack, from hardware access to custom tuning.