
6/30/2026
What this post added
Introduced GLM 5.2 Fast, a new serving path for GLM 5.2 that achieves 2-3x higher inference throughput than the Standard path on shared serverless infrastructure. This is enabled by architectural optimizations including advanced Mixture-of-Experts (MoE) sharding strategies, sparse attention serving, and speculative decoding. The post details the trade-offs in MoE and attention sharding, emphasizing that different parallelism strategies are chosen independently based on workload characteristics. It also highlights the continued support for the full 1M-token context window, production-ready features like structured output modes and tool calling, and aggressive prompt-caching pricing to ensure speed does not compromise reliability or cost-effectiveness for agentic workloads.