BlogsTogether AIBatch Inference API

Batch Inference API

Batch Inference API

2
posts
2025

Together AI now offers native deployment of elite proprietary Text-to-Speech (TTS) models, including MiniMax Speech 2.6 Turbo, on dedicated infrastructure. This enables high-quality, low-latency voice generation with advanced features like multilingual streaming, emotional awareness, and rapid voice cloning, integrated alongside LLM and STT workloads for a unified real-time voice agent pipeline. The Batch Inference API has been significantly improved with a streamlined UI for easier job creation and tracking. It now supports all serverless models and private deployments, offering universal model access. Rate limits have been increased by 3000x, from 10M to 30B enqueued tokens per model per user. The API also offers lower costs, typically 50% of the real-time API for most serverless models, making it an economical choice for high-throughput workloads. Ideal use cases include large-scale batch inference on models like DeepSeek-R1-0528, leveraging NVIDIA Blackwell GPUs and optimized inference engines for top speeds.

2025

Improved Batch Inference API: Enhanced UI, Expanded Model Support, and 3000× Rate Limit Increase

9/15/2025

This post details significant enhancements to the Batch Inference API, including a new UI, expanded model support to include all serverless and private deployments, and a 3000x increase in rate limits (from 10M to 30B tokens). It also highlights a cost reduction of approximately 50% compared to the real-time API for most serverless models. The post outlines ideal use cases for high-throughput, non-real-time inference tasks.

Together AI Delivers Top Speeds for DeepSeek-R1-0528 Inference on NVIDIA Blackwell

7/17/2025

This post details the significant performance improvements for DeepSeek-R1-0528 inference on NVIDIA Blackwell GPUs, achieved through a combination of bespoke GPU kernels, a proprietary inference engine, speculative decoding methods (Together Turbo Speculator), and calibrated/quantized model optimization. It highlights the fastest serverless inference performance for DeepSeek-R1-0528 to date, with up to 334 tokens/sec on HGX B200. Dedicated Endpoints are also discussed as a further optimization layer for production workloads, offering up to an 84 tokens/sec speedup. The post elaborates on the technical components: the inference engine's use of FlashAttention-3 and CUDA graphs, the development of Blackwell GPU kernels leveraging 5th-gen Tensor Cores, the Turbo Speculator's improved acceptance rate over other methods, and a lossless quantization technique (NVFP4/MXFP4/6/8) preserving model accuracy.