BlogsTogether AIText-to-Speech (TTS) Model Deployment

Text-to-Speech (TTS) Model Deployment

Text-to-Speech (TTS) Model Deployment

4
posts
2025–2026

Together AI now offers native deployment of elite proprietary Text-to-Speech (TTS) models, including Rime's Arcana V3 Turbo (English-Spanish code-switching, ~120ms TTFA) and Arcana V3 (11-language switching, ~160ms TTFA), on dedicated infrastructure. These models feature natural code-switching that maintains cadence and prosody across language boundaries, enabling more conversational and trustworthy voice agents. They are co-located with LLM and STT workloads for unified API, authentication, and billing. A new 'Voice Finder' tool has been introduced to simplify the selection of over 600+ voices across multiple TTS models by allowing search via prompt or audio sample, and by providing model-aware metadata across 15+ dimensions (pitch, accent, language, age, emotion, speaking style). This tool aims to accelerate the process of finding the right voice for specific use cases, improving the overall voice agent development experience.

2026

Introducing voice finder — a new tool to quickly find the right voice for your app from over 600+ voices

5/12/2026

Introduced 'Voice Finder', a new tool that indexes over 600 voices from 10 TTS models on Together AI. This tool enables searching for voices by prompt or audio sample and provides structured metadata across 15+ dimensions (pitch, gender, accent, language, age, emotion, speaking style) powered by an omni-model. This enhances the developer experience for selecting appropriate TTS voices for voice agents.

Rime Arcana V3 Turbo and Rime Arcana V3 now available on Together AI

2/4/2026

Introduces Rime Arcana V3 Turbo and Rime Arcana V3 to the Together AI platform. V3 Turbo offers English-Spanish code-switching with ~120ms time-to-first-audio on dedicated endpoints, trained on bilingual speech patterns. V3 supports 11-language switching with ~160ms time-to-first-audio. Both models are designed for natural code-switching, maintaining cadence and prosody, and are co-located with LLM and STT workloads for unified pipeline performance.

2025

MiniMax Speech 2.6 Turbo now available natively on Together AI

12/23/2025

This post introduces the native availability of MiniMax Speech 2.6 Turbo, a top-ranked TTS model, on Together AI's dedicated infrastructure. It details the model's technical capabilities including sub-250ms latency, 40+ language support with inline switching, 10-second voice cloning, and automatic emotional awareness. The post emphasizes the benefits of running TTS alongside LLM and STT workloads on unified infrastructure for reduced latency and improved developer experience, including unified observability and API access.

Rime voice models now available on Together AI

12/18/2025

This post introduces the integration of Rime's Arcana v2 and Mist v2 Text-to-Speech (TTS) models onto the Together AI platform. Arcana v2 offers expressive, conversational voices with multilingual support and advanced features like natural breathing and backchannel cues, trained on real customer service interactions. Mist v2 provides deterministic pronunciation control, ensuring consistent pronunciation of specific terms across calls and channels, with a p50 time-to-first-audio of approximately 225ms on dedicated endpoints. Both models are deployed on dedicated GPU endpoints, co-located with LLM and STT workloads, enabling a unified API, authentication, and observability surface for end-to-end voice agent pipelines. The post highlights use cases in global contact centers, real-time customer service, healthcare voice agents, and voice banking, emphasizing the production-grade infrastructure and developer experience provided by Together AI.