.png)
12/18/2025
What this post added
This post introduces the integration of Rime's Arcana v2 and Mist v2 Text-to-Speech (TTS) models onto the Together AI platform. Arcana v2 offers expressive, conversational voices with multilingual support and advanced features like natural breathing and backchannel cues, trained on real customer service interactions. Mist v2 provides deterministic pronunciation control, ensuring consistent pronunciation of specific terms across calls and channels, with a p50 time-to-first-audio of approximately 225ms on dedicated endpoints. Both models are deployed on dedicated GPU endpoints, co-located with LLM and STT workloads, enabling a unified API, authentication, and observability surface for end-to-end voice agent pipelines. The post highlights use cases in global contact centers, real-time customer service, healthcare voice agents, and voice banking, emphasizing the production-grade infrastructure and developer experience provided by Together AI.