Speaking of Voxtral | Mistral AI
3/23/2026
This post introduces Voxtral TTS, a new text-to-speech model. It details the model's architecture (transformer-based, autoregressive, flow-matching, built on Ministral 3B, with specific parameter counts for backbone, acoustic transformer, and codec), its performance metrics (low latency, RTF, multilingual capabilities, voice adaptation), and its use cases in enterprise voice workflows. It also includes comparative human evaluation data against ElevenLabs.