Baseten AI Model Deployment and Serving Platform
Introducing GLM-5.2 Fast

Introducing GLM-5.2 Fast

7/23/2026

What this post added

Introduces GLM-5.2 Fast, a new Model API tier for GLM-5.2 weights, optimized for per-user throughput for real-time agentic applications. This tier is designed to handle variable, bursty workloads with consistent throughput and low latency. It maintains ease-of-use with OpenAI-compatible API endpoints and a pay-per-token model. The post provides a Python code snippet demonstrating how to switch to the Fast endpoint by changing the model slug.

Read the original post ↗