
3/19/2026 · Michelle Chen, Kevin Flansburg, Ashish Datta, Kevin Jain
What this post added
This post announces the integration of large language models (LLMs) into Workers AI, starting with Moonshot AI's Kimi K2.5. It details the technical advancements made to support these models, including custom kernels for the Infire inference engine, optimizations for GPU utilization, and the implementation of parallelization techniques. The post also introduces platform improvements for agentic workloads, such as prefix caching with surfaced metrics and discounts, a new `x-session-affinity` header for improved cache hit rates, and redesigned asynchronous APIs for durable inference processing. It highlights significant cost savings achieved by using Kimi K2.5 for internal security review agents.