
Improved Batch Inference API: Enhanced UI, Expanded Model Support, and 3000× Rate Limit Increase
9/15/2025
This post details significant enhancements to the Batch Inference API, including a new UI, expanded model support to include all serverless and private deployments, and a 3000x increase in rate limits (from 10M to 30B tokens). It also highlights a cost reduction of approximately 50% compared to the real-time API for most serverless models. The post outlines ideal use cases for high-throughput, non-real-time inference tasks.
