
7/10/2025
What this post added
This post introduces Together AI's new Speech-to-Text (STT) APIs, focusing on the technical optimizations that enable 15x faster transcription than OpenAI's offering. Key technical details include the use of Silero for smart voice activity detection, intelligent chunking and batching strategies for long audio files, and engine improvements to maximize GPU utilization for the Whisper V3 Large model. The post also highlights production-ready API features like enterprise-scale file handling (over 1GB), superior word-level alignment, comprehensive language support (50+ languages), dedicated endpoints for low latency, and batch processing capabilities. It positions STT as a foundational element for voice AI applications, integrating with existing LLM services.