Artificial Intelligence Integration & Impact
Powering the agents: Workers AI now runs large models, starting with Kimi K2.5

Powering the agents: Workers AI now runs large models, starting with Kimi K2.5

3/19/2026 · Michelle Chen, Kevin Flansburg, Ashish Datta, Kevin Jain

What this post added

This post announces the integration of large language models (LLMs) into Workers AI, starting with Moonshot AI's Kimi K2.5. It details the technical advancements made to support these models, including custom kernels for the Infire inference engine, optimizations for GPU utilization, and the implementation of parallelization techniques. The post also introduces platform improvements for agentic workloads, such as prefix caching with surfaced metrics and discounts, a new `x-session-affinity` header for improved cache hit rates, and redesigned asynchronous APIs for durable inference processing. It highlights significant cost savings achieved by using Kimi K2.5 for internal security review agents.

Read the original post ↗