
Open, convenient and predictable: Introducing Provisioned Throughput
7/8/2026
Introduced Provisioned Throughput, a new inference capacity model for open LLMs. This feature provides guaranteed token capacity with a 99% uptime SLA, offering predictable pricing and simplifying inference management for production workloads. It defines Provisioned Throughput Units (PTUs) and their consumption rates for input, cached input, and output tokens, enabling users to estimate costs based on traffic shape. The initial offering supports MiniMax M3 and GLM-5.2 models.
